Skip to content

SharePoint search and enrichment blueprint

Recommendation

Keep OpenText Knowledge Discovery as the primary search index. Improve the quality of extracted content and normalized metadata first, then add vectors to the same index for the searches that benefit from semantic matching.

The order matters. The supplied exports are an illustrative sample, not an inventory of the full SharePoint corpus. They show risks worth measuring across the full collection: content reaching a configured character limit, records without extracted content, and AI labels inferred from paths or filenames. An embedding cannot recover content that was never extracted, so vectorization should follow a corpus-wide extraction audit.

Decisions

Area Decision
Search platform Extend Knowledge Discovery; do not add a separate vector database.
Retrieval Combine exact, lexical/conceptual, parametric, and vector retrieval.
Vector storage Store chunk vectors in Knowledge Discovery VectorType fields.
First priority Audit extraction limits and add format-aware extraction.
Metadata Preserve raw metadata and create a small canonical search schema.
AI enrichment Store as a proposal with confidence, evidence, provenance, and model version.
Security SharePoint ACLs remain authoritative and constrain every search and preview.
AI hosting Run enrichment and embedding inside the trusted environment.

Delivery path

  1. Normalize and benchmark. Map authoritative metadata into stable search, facet, date, and security fields. Establish the current lexical baseline.
  2. Fix extraction. Identify and remove harmful export/indexing limits and route each format to an appropriate extractor.
  3. Pilot hybrid search. Embed semantic chunks, generate query vectors with the same model, and combine vector similarity with existing search.
  4. Add governed enrichment. Generate summaries, taxonomy labels, and entities only when they improve a measured search or review task.

Non-negotiable security invariant

Search, vectors, snippets, summaries, facets, related-document suggestions, and previews must never reveal a document unless the requesting identity is authorized by the source SharePoint ACL. An AI-generated classification such as Internal or Restricted is descriptive metadata, not authorization.

Deliberate omissions

This blueprint does not add a second search platform, a custom vector store, image-similarity search, or a generative answer layer. Add one only after a measured requirement shows that Knowledge Discovery and text-based hybrid retrieval cannot meet it.