SharePoint search and enrichment blueprint¶
Recommendation¶
Keep OpenText Knowledge Discovery as the primary search index. Improve the quality of extracted content and normalized metadata first, then add vectors to the same index for the searches that benefit from semantic matching.
The order matters. The supplied exports are an illustrative sample, not an inventory of the full SharePoint corpus. They show risks worth measuring across the full collection: content reaching a configured character limit, records without extracted content, and AI labels inferred from paths or filenames. An embedding cannot recover content that was never extracted, so vectorization should follow a corpus-wide extraction audit.
Decisions¶
| Area | Decision |
|---|---|
| Search platform | Extend Knowledge Discovery; do not add a separate vector database. |
| Retrieval | Combine exact, lexical/conceptual, parametric, and vector retrieval. |
| Vector storage | Store chunk vectors in Knowledge Discovery VectorType fields. |
| First priority | Audit extraction limits and add format-aware extraction. |
| Metadata | Preserve raw metadata and create a small canonical search schema. |
| AI enrichment | Store as a proposal with confidence, evidence, provenance, and model version. |
| Security | SharePoint ACLs remain authoritative and constrain every search and preview. |
| AI hosting | Run enrichment and embedding inside the trusted environment. |
Delivery path¶
- Normalize and benchmark. Map authoritative metadata into stable search, facet, date, and security fields. Establish the current lexical baseline.
- Fix extraction. Identify and remove harmful export/indexing limits and route each format to an appropriate extractor.
- Pilot hybrid search. Embed semantic chunks, generate query vectors with the same model, and combine vector similarity with existing search.
- Add governed enrichment. Generate summaries, taxonomy labels, and entities only when they improve a measured search or review task.
Non-negotiable security invariant¶
Search, vectors, snippets, summaries, facets, related-document suggestions,
and previews must never reveal a document unless the requesting identity is
authorized by the source SharePoint ACL. An AI-generated classification such
as Internal or Restricted is descriptive metadata, not authorization.
Read next¶
- Example export observations explains what the supplied sample can teach without treating it as representative of the full corpus.
- Target architecture defines the ingestion and query components and their trust boundaries.
- Search and vectorization explains where vectors help and where they do not.
- Enrichment and canonical schema defines field precedence, provenance, and the target records.
- Evaluation and rollout defines the test set, acceptance gates, and staged deployment.
Deliberate omissions¶
This blueprint does not add a second search platform, a custom vector store, image-similarity search, or a generative answer layer. Add one only after a measured requirement shows that Knowledge Discovery and text-based hybrid retrieval cannot meet it.