Evaluation and rollout¶
Baseline before change¶
Capture the current Knowledge Discovery behavior before altering extraction or ranking:
- indexed document and version count;
- extraction success and content length by format;
- zero-result rate and p50/p95 query latency;
- Recall@10, nDCG@10, and MRR on judged queries;
- current facet coverage and accuracy;
- authorization failures and any security leakage;
- enrichment coverage, confidence, review rate, and reviewer agreement.
Keep the current lexical/conceptual configuration as the control. Evaluate each phase against it so extraction gains are not incorrectly attributed to vectors.
Judged query set¶
Build 100 representative queries with one or more relevant documents and, where applicable, explicit non-relevant documents.
| Query group | Count | Examples of intent |
|---|---|---|
| Exact and known-item | 25 | Document number, filename, revision, title phrase |
| Semantic and topical | 25 | Natural-language description, terminology mismatch |
| Faceted | 20 | Project/site/type/date/author combinations |
| Permission-sensitive | 20 | Same query under identities with different ACLs |
| Content-limited | 10 | Image OCR, transcript, email, CAD, URL, archive member |
| Total | 100 |
Include all supported production languages in proportion to actual use. Include short ambiguous queries, misspellings, document-number punctuation variants, and queries where no result is the correct result.
Judgments must be made from authorized source documents, not from the current rank order. Store query, identity/role, filters, relevant document IDs, graded relevance, and rationale.
Metrics and release gates¶
| Area | Metric | Gate |
|---|---|---|
| Security | Unauthorized documents, snippets, facets, summaries, or previews | Exactly zero |
| Relevance | nDCG@10 | At least 10% above normalized lexical baseline for hybrid adoption |
| Known item | MRR and Recall@10 | No regression on exact/identifier queries |
| Coverage | Extraction success by supported format | Improves in every targeted format; configured content limits do not silently truncate indexed text |
| Latency | p95 end-to-end latency | No more than 20% above baseline unless explicitly accepted |
| Operations | Failed/stale documents and chunks | Observable and retryable; no silently stale vectors |
| Enrichment | Precision and reviewer agreement by field | Threshold agreed by the field owner before use as a facet |
Report metrics by query group and format. A good average can hide a failure in exact drawing-number search or a security-sensitive role.
Representative acceptance scenarios¶
- Exact drawing number:
DRAWING-EXAMPLE-001ranks the matching file first through exact fields without requiring a vector. - Natural-language concept: a user describes a technical topic using terms absent from the title; hybrid retrieval improves ranking over lexical alone and shows the matching location.
- Faceted search: project, site, document type, and modified-date filters reduce the result set without semantic reinterpretation.
- Image: OCR text is searchable and the result points to the matching
image region; a failed OCR record remains visibly
metadata_only. - Video: transcript text is searchable and links to a timestamp.
- Contentless CAD: filename/path metadata finds the drawing with a low semantic weight and an extraction warning.
- Version update: a new SharePoint version replaces old chunks and vectors; stale content is not returned.
- Move: the SharePoint stable identity remains unchanged while path-derived facets and display URL update.
- Delete: document, chunks, vectors, previews, and enrichments disappear.
- Permission change: an ACL-only update changes visibility without waiting for re-embedding.
- Security trimming: an unauthorized identity receives no result, facet contribution, snippet, related-document suggestion, or preview.
- Model migration: old and new vectors can be compared, and rollback does not require re-extracting source content.
Phased rollout¶
Phase 1 — audit, normalize, and remove harmful limits¶
- Measure extraction lengths and configured limits across the full corpus, then export/index complete logical content instead of capped prefixes.
- Create the canonical field mapping and typed Knowledge Discovery fields.
- Configure useful parametric facets and authoritative dates.
- Verify SharePoint ACL ingestion and security trimming.
- Establish the judged-query lexical baseline.
Exit when counts reconcile, truncation is gone, supported text formats extract reliably, and all security tests pass.
Phase 2 — close format gaps¶
- Add OCR for images and scanned PDF pages.
- Add transcription for audio/video.
- Parse email bodies and attachments.
- Extract CAD title blocks, layouts/models, layers, and visible text.
- Define safe URL and archive-member processing policies.
Exit when each targeted format improves extraction coverage and its acceptance scenarios return evidence with a useful location.
Phase 3 — vector pilot¶
- Select a trusted-boundary multilingual embedding model using the judged set.
- Generate vectors for semantic chunks and query text with identical settings.
- Store vectors in Knowledge Discovery
VectorTypefields. - Compare lexical, vector-only, and hybrid retrieval by query group.
- Tune query-class weights and metadata-only down-weighting.
Adopt hybrid retrieval only when relevance and latency gates pass with zero security leakage. Keep lexical-only fallback available during rollout.
Phase 4 — governed enrichment¶
- Add evidence-backed summaries, controlled taxonomy labels, and useful entities.
- Calibrate confidence per field and introduce targeted review queues.
- Promote only approved or demonstrably precise values to normal facets.
- Monitor drift by content type, language, taxonomy version, and model version.
Operational monitoring¶
Track at minimum:
- ingest lag and source/index version mismatch;
- extraction success, warnings, and content length by format;
- documents and chunks per source version;
- embedding failures, model versions, and stale vectors;
- query latency and zero-result rate by query class;
- relevance metrics on a scheduled regression run;
- ACL update lag and security-test results;
- proposal, approval, rejection, and reviewer-agreement rates.
Alert on permission-processing failures, missing security context, orphaned chunks, stale versions, sudden extraction drops, and mixed embedding models in the active vector field.
Rollback¶
Each phase must be independently reversible:
- retain raw source metadata and extracted content;
- version canonical mappings, chunking, taxonomy, and embedding models;
- keep old and new fields side by side during migration;
- switch query weights/configuration without re-ingesting source documents;
- never couple an AI label rollout to ACL enforcement.