Skip to content

Example export observations

Scope

The supplied Knowledge Discovery exports are an illustrative document sample. They are useful for understanding field structure, format variety, extraction behavior, and the shape of the proposed AI enrichment. They are not a complete or statistically representative inventory of the SharePoint corpus.

Do not extrapolate sample counts or percentages to the production collection. Run the baseline audit against the complete index before setting priorities or success targets.

What the sample demonstrates

The exports show three conditions that the full-corpus audit should test:

  • extracted text can reach a configured character limit, so the exported value may be a prefix rather than the complete document;
  • some formats or records have no extracted DRECONTENT and need a specialized extractor or an explicit metadata-only state;
  • AI enrichment can be based on incomplete content, paths, filenames, or metadata and must therefore retain evidence and review state.

These are design signals, not estimates of how common each condition is.

Formats observed in the sample

Family Observed MIME types Examples Extraction approach to validate
Office documents application/x-ms-word07, application/msword, application/x-ms-powerpoint07, application/x-ms-powerpoint DOCX, DOC, PPTX, PPT Logical sections, headings, pages, and slides
Tables application/x-ms-excel07, application/x-ms-excel, text/csv XLSX, XLS, CSV Sheets or row groups with repeated headers; numeric facets stay structured
PDF application/pdf Native and scanned PDF Page-aware text extraction with OCR fallback
Text and links text/plain, text/xml TXT, URL, XML Distinguish files from shortcuts; crawl only permitted targets
CAD and diagrams image/x-dwg, application/octet-stream, application/vnd.visio DWG, DGN, VSDX Title blocks, drawing metadata, layers, layouts, and visible text
Images image/png, image/jpeg PNG, JPG OCR and evidence-backed captions
Video video/mp4, video/quicktime MP4, MOV Timestamped speech transcription
Email application/vnd.ms-outlook MSG Subject, sender, body, and separately processed attachments
Archives application/zip ZIP Process permitted members with inherited security and provenance

MIME type is not sufficient by itself. The sample includes text/plain records representing both text documents and URL shortcuts, while a CAD file can appear as application/octet-stream. Route extraction using a validated combination of extension, detected format, and MIME type.

Extraction patterns to investigate

Sample observation Corpus-wide check Desired handling
Text reaches a configured field limit Plot extracted length and limit hits by format and connector Export/index complete logical content and flag partial extraction
Images have no searchable body text Measure OCR coverage and quality Store OCR text, location, confidence, and failure reason
Video has no transcript Measure transcript coverage and language accuracy Store timestamped transcript chunks
Email or archive records lack body content Test body and permitted-member processing Preserve parent/child provenance and inherited ACLs
CAD content is absent or sparse Compare extracted text with title blocks and visible drawing text Index drawing number, revision, layout/model, layers, and text
PDF or spreadsheet extraction varies Sample native, scanned, formula-heavy, and large files Route failures and partial results to format-specific handling

Useful existing metadata

The sample contains fields worth normalizing instead of regenerating:

  • stable lookup and provenance fields such as AUTN_IDENTIFIER, DREREFERENCE, AUTN_SOURCE, CONNECTOR_GROUP, and DREDBNAME;
  • file type, application, size, author, last author, created, and modified metadata;
  • existing document type, site, department, and DMS fields where present;
  • detected language when content is available;
  • file-specific metadata such as slide count, image dimensions, video duration, CAD layouts, and embedded DMS properties.

The raw field space is noisy, with repeated application/date/size aliases and low-frequency embedded metadata. Do not expose every raw field as a search facet. Map only fields that support a known search or governance use case.

Audit of the example enrichment

Review and clustering

The provided enriched sample keeps results in a proposed review state and adds cluster labels, coordinates, distances, outlier flags, confidence, and review reasons. These fields can support offline quality review, but they are not authoritative business metadata.

semanticClusterX and semanticClusterY are projection coordinates for a particular visualization run. They are unstable across models, datasets, and reruns, so they must not become search facets. Cluster distance and outlier flags are useful for corpus exploration and review prioritization, not normal user ranking.

Generated document metadata

  • Generated author and last-save values can duplicate existing extractor or SharePoint metadata.
  • Generated owner and status values must not replace authoritative SharePoint or DMS fields.
  • The contentless DGN example is labelled Technical Drawing with low confidence. That is a path or filename inference, not content analysis.
  • Any inferred field should carry its evidence, source, confidence, model version, and review status.

Security labels

The sample proposes labels such as Internal, Confidential, Restricted, and Highly Restricted. These can prioritize human review but cannot grant or deny access. SharePoint ACLs remain authoritative, including when the proposed label looks plausible.

Full-corpus audit required

Before implementation, calculate across the complete index:

  1. document and version totals by connector, site, format, and MIME type;
  2. extraction success, empty content, partial content, and length-limit hits by format;
  3. duplicate, stale, deleted, and orphaned records;
  4. ACL coverage and permission-update lag;
  5. metadata completeness and conflicting values by authoritative source;
  6. enrichment coverage, evidence quality, confidence calibration, and review outcomes.

Consequences for the design

  1. Verify extraction quality across the full corpus before embedding.
  2. Add format-specific extraction where measured gaps justify it.
  3. Normalize authoritative fields before asking AI to recreate them.
  4. Keep extraction quality and enrichment review status visible in the index.
  5. Never represent missing or capped content as a complete semantic description of a document.