Target architecture¶
Design intent¶
The target keeps SharePoint as the source of documents and permissions and Knowledge Discovery as the searchable content index. A small processing layer normalizes metadata, routes formats, produces governed enrichments, and writes text chunks and vectors back to Knowledge Discovery.
OpenText documents the relevant native capabilities: typed field content, parametric investigation, vector search, and document previews. Confirm the installed Knowledge Discovery version and licensed features before applying version-specific configuration.
Logical architecture¶
flowchart LR
subgraph sources ["Authoritative sources"]
sharepoint[SharePoint documents]
permissions[SharePoint ACLs]
end
subgraph processing ["Trusted processing boundary"]
connector[Connector ingestion]
normalize[Metadata normalization]
extract[Format extraction]
enrich[Governed enrichment]
embed[Embedding model]
end
subgraph knowledge ["Knowledge Discovery"]
content[(Content index)]
vectors[(VectorType fields)]
view[View previews]
end
subgraph search ["Search experience"]
query[Hybrid query service]
ui[Search UI]
end
sharepoint -->|"Fetch versions"| connector
permissions -->|"Preserve ACL"| connector
connector --> normalize
normalize --> extract
extract --> enrich
enrich --> embed
normalize -->|"Typed fields"| content
extract -->|"Full text"| content
enrich -->|"Proposals"| content
embed -->|"Chunk vectors"| vectors
query -->|"Text, fields, security"| content
query -->|"Query vector"| vectors
content --> view
vectors --> query
view --> query
query --> ui
The diagram is logical, not a requirement for separate deployable services. Normalization, extraction routing, enrichment, and embedding can live in the existing ingestion flow if that is the shortest operational path.
Ingestion responsibilities¶
Connector ingestion¶
- Read the SharePoint stable item identity, version/ETag, URL, library, folder, source fields, and ACL.
- Detect creates, updates, moves, permission changes, and deletes.
- Keep a source version on every canonical document and chunk so stale content and vectors can be replaced together.
- Preserve the original connector metadata for audit and troubleshooting.
Metadata normalization¶
- Map duplicate raw aliases to one canonical field without deleting raw data.
- Parse epoch dates and numeric sizes into typed fields.
- Derive filename, extension, readable path segments, and document-number candidates deterministically.
- Resolve precedence using the rules in Enrichment and canonical schema.
- Mark invalid or contradictory values instead of silently correcting them.
Format-specific extraction¶
- Extract complete logical units, not a prefix capped by a fixed field limit.
- Record extractor, version, content hash, success state, warnings, and unit location such as page, slide, sheet, model, or timestamp.
- Route unsupported or failed formats to a review queue; do not fabricate content to make the record appear complete.
Enrichment and embedding¶
- Run only inside the trusted boundary.
- Enrich extracted content and reliable metadata, never the raw metadata dump.
- Attach evidence and provenance to every inferred value.
- Generate vectors only after chunking and language/quality checks.
- Use the same embedding model and version for indexed and query vectors.
Processing decision flow¶
flowchart TD
received([Document version received])
preserve[Preserve raw metadata and ACL]
route{Format family?}
text[Extract prose and structure]
table[Extract sheets and tables]
visual[Render and OCR]
media[Transcribe with timestamps]
container[Resolve target or members]
quality{Usable content?}
canonical[Build canonical record]
chunk[Create logical chunks]
infer[Generate proposals]
vector[Generate vectors]
fallback[Create metadata fallback]
review[Queue for review]
index([Index secured document])
received --> preserve
preserve --> route
route -->|"PDF and Office"| text
route -->|"Sheets and CSV"| table
route -->|"Image and CAD"| visual
route -->|"Audio and video"| media
route -->|"URL, email, archive"| container
text --> quality
table --> quality
visual --> quality
media --> quality
container --> quality
quality -->|"Yes"| canonical
canonical --> chunk
chunk --> infer
infer --> vector
vector --> index
quality -->|"No"| fallback
fallback --> review
review --> index
style quality fill:#FFECBD,stroke:#FFC943
style review fill:#FFE0C2,stroke:#FF9E42
style index fill:#CDF4D3,stroke:#66D575
Security invariants¶
- The SharePoint ACL is copied by the connector and stored in the configured Knowledge Discovery security field.
- Every query carries the authenticated user's security context.
- Facet counts, snippets, summaries, related results, and previews are derived only from security-trimmed results.
- Chunks inherit the parent document security reference; no independently searchable chunk can have broader visibility.
- Enrichment workers may read only documents their service identity is authorized to process, and their output inherits the source ACL.
- Permission-only changes update security fields even when content and vectors are unchanged.
- Generated sensitivity labels never modify authorization automatically.
Update and delete behavior¶
Use (documentId, sourceVersion) as the version boundary. On a new version,
write the canonical document and all new chunks, then atomically make the new
version searchable and remove the old chunks. On deletion, remove the document,
chunks, vectors, cached previews, and generated enrichments. On a move, retain
the stable SharePoint item identity and update path-derived metadata.