Example export observations¶
Scope¶
The supplied Knowledge Discovery exports are an illustrative document sample. They are useful for understanding field structure, format variety, extraction behavior, and the shape of the proposed AI enrichment. They are not a complete or statistically representative inventory of the SharePoint corpus.
Do not extrapolate sample counts or percentages to the production collection. Run the baseline audit against the complete index before setting priorities or success targets.
What the sample demonstrates¶
The exports show three conditions that the full-corpus audit should test:
- extracted text can reach a configured character limit, so the exported value may be a prefix rather than the complete document;
- some formats or records have no extracted
DRECONTENTand need a specialized extractor or an explicit metadata-only state; - AI enrichment can be based on incomplete content, paths, filenames, or metadata and must therefore retain evidence and review state.
These are design signals, not estimates of how common each condition is.
Formats observed in the sample¶
| Family | Observed MIME types | Examples | Extraction approach to validate |
|---|---|---|---|
| Office documents | application/x-ms-word07, application/msword, application/x-ms-powerpoint07, application/x-ms-powerpoint |
DOCX, DOC, PPTX, PPT | Logical sections, headings, pages, and slides |
| Tables | application/x-ms-excel07, application/x-ms-excel, text/csv |
XLSX, XLS, CSV | Sheets or row groups with repeated headers; numeric facets stay structured |
application/pdf |
Native and scanned PDF | Page-aware text extraction with OCR fallback | |
| Text and links | text/plain, text/xml |
TXT, URL, XML | Distinguish files from shortcuts; crawl only permitted targets |
| CAD and diagrams | image/x-dwg, application/octet-stream, application/vnd.visio |
DWG, DGN, VSDX | Title blocks, drawing metadata, layers, layouts, and visible text |
| Images | image/png, image/jpeg |
PNG, JPG | OCR and evidence-backed captions |
| Video | video/mp4, video/quicktime |
MP4, MOV | Timestamped speech transcription |
application/vnd.ms-outlook |
MSG | Subject, sender, body, and separately processed attachments | |
| Archives | application/zip |
ZIP | Process permitted members with inherited security and provenance |
MIME type is not sufficient by itself. The sample includes text/plain
records representing both text documents and URL shortcuts, while a CAD file
can appear as application/octet-stream. Route extraction using a validated
combination of extension, detected format, and MIME type.
Extraction patterns to investigate¶
| Sample observation | Corpus-wide check | Desired handling |
|---|---|---|
| Text reaches a configured field limit | Plot extracted length and limit hits by format and connector | Export/index complete logical content and flag partial extraction |
| Images have no searchable body text | Measure OCR coverage and quality | Store OCR text, location, confidence, and failure reason |
| Video has no transcript | Measure transcript coverage and language accuracy | Store timestamped transcript chunks |
| Email or archive records lack body content | Test body and permitted-member processing | Preserve parent/child provenance and inherited ACLs |
| CAD content is absent or sparse | Compare extracted text with title blocks and visible drawing text | Index drawing number, revision, layout/model, layers, and text |
| PDF or spreadsheet extraction varies | Sample native, scanned, formula-heavy, and large files | Route failures and partial results to format-specific handling |
Useful existing metadata¶
The sample contains fields worth normalizing instead of regenerating:
- stable lookup and provenance fields such as
AUTN_IDENTIFIER,DREREFERENCE,AUTN_SOURCE,CONNECTOR_GROUP, andDREDBNAME; - file type, application, size, author, last author, created, and modified metadata;
- existing document type, site, department, and DMS fields where present;
- detected language when content is available;
- file-specific metadata such as slide count, image dimensions, video duration, CAD layouts, and embedded DMS properties.
The raw field space is noisy, with repeated application/date/size aliases and low-frequency embedded metadata. Do not expose every raw field as a search facet. Map only fields that support a known search or governance use case.
Audit of the example enrichment¶
Review and clustering¶
The provided enriched sample keeps results in a proposed review state and adds cluster labels, coordinates, distances, outlier flags, confidence, and review reasons. These fields can support offline quality review, but they are not authoritative business metadata.
semanticClusterX and semanticClusterY are projection coordinates for a
particular visualization run. They are unstable across models, datasets, and
reruns, so they must not become search facets. Cluster distance and outlier
flags are useful for corpus exploration and review prioritization, not normal
user ranking.
Generated document metadata¶
- Generated author and last-save values can duplicate existing extractor or SharePoint metadata.
- Generated owner and status values must not replace authoritative SharePoint or DMS fields.
- The contentless DGN example is labelled
Technical Drawingwith low confidence. That is a path or filename inference, not content analysis. - Any inferred field should carry its evidence, source, confidence, model version, and review status.
Security labels¶
The sample proposes labels such as Internal, Confidential, Restricted,
and Highly Restricted. These can prioritize human review but cannot grant or
deny access. SharePoint ACLs remain authoritative, including when the proposed
label looks plausible.
Full-corpus audit required¶
Before implementation, calculate across the complete index:
- document and version totals by connector, site, format, and MIME type;
- extraction success, empty content, partial content, and length-limit hits by format;
- duplicate, stale, deleted, and orphaned records;
- ACL coverage and permission-update lag;
- metadata completeness and conflicting values by authoritative source;
- enrichment coverage, evidence quality, confidence calibration, and review outcomes.
Consequences for the design¶
- Verify extraction quality across the full corpus before embedding.
- Add format-specific extraction where measured gaps justify it.
- Normalize authoritative fields before asking AI to recreate them.
- Keep extraction quality and enrichment review status visible in the index.
- Never represent missing or capped content as a complete semantic description of a document.