Skip to content

Evaluation and rollout

Baseline before change

Capture the current Knowledge Discovery behavior before altering extraction or ranking:

  • indexed document and version count;
  • extraction success and content length by format;
  • zero-result rate and p50/p95 query latency;
  • Recall@10, nDCG@10, and MRR on judged queries;
  • current facet coverage and accuracy;
  • authorization failures and any security leakage;
  • enrichment coverage, confidence, review rate, and reviewer agreement.

Keep the current lexical/conceptual configuration as the control. Evaluate each phase against it so extraction gains are not incorrectly attributed to vectors.

Judged query set

Build 100 representative queries with one or more relevant documents and, where applicable, explicit non-relevant documents.

Query group Count Examples of intent
Exact and known-item 25 Document number, filename, revision, title phrase
Semantic and topical 25 Natural-language description, terminology mismatch
Faceted 20 Project/site/type/date/author combinations
Permission-sensitive 20 Same query under identities with different ACLs
Content-limited 10 Image OCR, transcript, email, CAD, URL, archive member
Total 100

Include all supported production languages in proportion to actual use. Include short ambiguous queries, misspellings, document-number punctuation variants, and queries where no result is the correct result.

Judgments must be made from authorized source documents, not from the current rank order. Store query, identity/role, filters, relevant document IDs, graded relevance, and rationale.

Metrics and release gates

Area Metric Gate
Security Unauthorized documents, snippets, facets, summaries, or previews Exactly zero
Relevance nDCG@10 At least 10% above normalized lexical baseline for hybrid adoption
Known item MRR and Recall@10 No regression on exact/identifier queries
Coverage Extraction success by supported format Improves in every targeted format; configured content limits do not silently truncate indexed text
Latency p95 end-to-end latency No more than 20% above baseline unless explicitly accepted
Operations Failed/stale documents and chunks Observable and retryable; no silently stale vectors
Enrichment Precision and reviewer agreement by field Threshold agreed by the field owner before use as a facet

Report metrics by query group and format. A good average can hide a failure in exact drawing-number search or a security-sensitive role.

Representative acceptance scenarios

  1. Exact drawing number: DRAWING-EXAMPLE-001 ranks the matching file first through exact fields without requiring a vector.
  2. Natural-language concept: a user describes a technical topic using terms absent from the title; hybrid retrieval improves ranking over lexical alone and shows the matching location.
  3. Faceted search: project, site, document type, and modified-date filters reduce the result set without semantic reinterpretation.
  4. Image: OCR text is searchable and the result points to the matching image region; a failed OCR record remains visibly metadata_only.
  5. Video: transcript text is searchable and links to a timestamp.
  6. Contentless CAD: filename/path metadata finds the drawing with a low semantic weight and an extraction warning.
  7. Version update: a new SharePoint version replaces old chunks and vectors; stale content is not returned.
  8. Move: the SharePoint stable identity remains unchanged while path-derived facets and display URL update.
  9. Delete: document, chunks, vectors, previews, and enrichments disappear.
  10. Permission change: an ACL-only update changes visibility without waiting for re-embedding.
  11. Security trimming: an unauthorized identity receives no result, facet contribution, snippet, related-document suggestion, or preview.
  12. Model migration: old and new vectors can be compared, and rollback does not require re-extracting source content.

Phased rollout

Phase 1 — audit, normalize, and remove harmful limits

  • Measure extraction lengths and configured limits across the full corpus, then export/index complete logical content instead of capped prefixes.
  • Create the canonical field mapping and typed Knowledge Discovery fields.
  • Configure useful parametric facets and authoritative dates.
  • Verify SharePoint ACL ingestion and security trimming.
  • Establish the judged-query lexical baseline.

Exit when counts reconcile, truncation is gone, supported text formats extract reliably, and all security tests pass.

Phase 2 — close format gaps

  • Add OCR for images and scanned PDF pages.
  • Add transcription for audio/video.
  • Parse email bodies and attachments.
  • Extract CAD title blocks, layouts/models, layers, and visible text.
  • Define safe URL and archive-member processing policies.

Exit when each targeted format improves extraction coverage and its acceptance scenarios return evidence with a useful location.

Phase 3 — vector pilot

  • Select a trusted-boundary multilingual embedding model using the judged set.
  • Generate vectors for semantic chunks and query text with identical settings.
  • Store vectors in Knowledge Discovery VectorType fields.
  • Compare lexical, vector-only, and hybrid retrieval by query group.
  • Tune query-class weights and metadata-only down-weighting.

Adopt hybrid retrieval only when relevance and latency gates pass with zero security leakage. Keep lexical-only fallback available during rollout.

Phase 4 — governed enrichment

  • Add evidence-backed summaries, controlled taxonomy labels, and useful entities.
  • Calibrate confidence per field and introduce targeted review queues.
  • Promote only approved or demonstrably precise values to normal facets.
  • Monitor drift by content type, language, taxonomy version, and model version.

Operational monitoring

Track at minimum:

  • ingest lag and source/index version mismatch;
  • extraction success, warnings, and content length by format;
  • documents and chunks per source version;
  • embedding failures, model versions, and stale vectors;
  • query latency and zero-result rate by query class;
  • relevance metrics on a scheduled regression run;
  • ACL update lag and security-test results;
  • proposal, approval, rejection, and reviewer-agreement rates.

Alert on permission-processing failures, missing security context, orphaned chunks, stale versions, sudden extraction drops, and mixed embedding models in the active vector field.

Rollback

Each phase must be independently reversible:

  • retain raw source metadata and extracted content;
  • version canonical mappings, chunking, taxonomy, and embedding models;
  • keep old and new fields side by side during migration;
  • switch query weights/configuration without re-ingesting source documents;
  • never couple an AI label rollout to ACL enforcement.