Skip to content

Target architecture

Design intent

The target keeps SharePoint as the source of documents and permissions and Knowledge Discovery as the searchable content index. A small processing layer normalizes metadata, routes formats, produces governed enrichments, and writes text chunks and vectors back to Knowledge Discovery.

OpenText documents the relevant native capabilities: typed field content, parametric investigation, vector search, and document previews. Confirm the installed Knowledge Discovery version and licensed features before applying version-specific configuration.

Logical architecture

flowchart LR
    subgraph sources ["Authoritative sources"]
        sharepoint[SharePoint documents]
        permissions[SharePoint ACLs]
    end

    subgraph processing ["Trusted processing boundary"]
        connector[Connector ingestion]
        normalize[Metadata normalization]
        extract[Format extraction]
        enrich[Governed enrichment]
        embed[Embedding model]
    end

    subgraph knowledge ["Knowledge Discovery"]
        content[(Content index)]
        vectors[(VectorType fields)]
        view[View previews]
    end

    subgraph search ["Search experience"]
        query[Hybrid query service]
        ui[Search UI]
    end

    sharepoint -->|"Fetch versions"| connector
    permissions -->|"Preserve ACL"| connector
    connector --> normalize
    normalize --> extract
    extract --> enrich
    enrich --> embed
    normalize -->|"Typed fields"| content
    extract -->|"Full text"| content
    enrich -->|"Proposals"| content
    embed -->|"Chunk vectors"| vectors
    query -->|"Text, fields, security"| content
    query -->|"Query vector"| vectors
    content --> view
    vectors --> query
    view --> query
    query --> ui

The diagram is logical, not a requirement for separate deployable services. Normalization, extraction routing, enrichment, and embedding can live in the existing ingestion flow if that is the shortest operational path.

Ingestion responsibilities

Connector ingestion

  • Read the SharePoint stable item identity, version/ETag, URL, library, folder, source fields, and ACL.
  • Detect creates, updates, moves, permission changes, and deletes.
  • Keep a source version on every canonical document and chunk so stale content and vectors can be replaced together.
  • Preserve the original connector metadata for audit and troubleshooting.

Metadata normalization

  • Map duplicate raw aliases to one canonical field without deleting raw data.
  • Parse epoch dates and numeric sizes into typed fields.
  • Derive filename, extension, readable path segments, and document-number candidates deterministically.
  • Resolve precedence using the rules in Enrichment and canonical schema.
  • Mark invalid or contradictory values instead of silently correcting them.

Format-specific extraction

  • Extract complete logical units, not a prefix capped by a fixed field limit.
  • Record extractor, version, content hash, success state, warnings, and unit location such as page, slide, sheet, model, or timestamp.
  • Route unsupported or failed formats to a review queue; do not fabricate content to make the record appear complete.

Enrichment and embedding

  • Run only inside the trusted boundary.
  • Enrich extracted content and reliable metadata, never the raw metadata dump.
  • Attach evidence and provenance to every inferred value.
  • Generate vectors only after chunking and language/quality checks.
  • Use the same embedding model and version for indexed and query vectors.

Processing decision flow

flowchart TD
    received([Document version received])
    preserve[Preserve raw metadata and ACL]
    route{Format family?}
    text[Extract prose and structure]
    table[Extract sheets and tables]
    visual[Render and OCR]
    media[Transcribe with timestamps]
    container[Resolve target or members]
    quality{Usable content?}
    canonical[Build canonical record]
    chunk[Create logical chunks]
    infer[Generate proposals]
    vector[Generate vectors]
    fallback[Create metadata fallback]
    review[Queue for review]
    index([Index secured document])

    received --> preserve
    preserve --> route
    route -->|"PDF and Office"| text
    route -->|"Sheets and CSV"| table
    route -->|"Image and CAD"| visual
    route -->|"Audio and video"| media
    route -->|"URL, email, archive"| container
    text --> quality
    table --> quality
    visual --> quality
    media --> quality
    container --> quality
    quality -->|"Yes"| canonical
    canonical --> chunk
    chunk --> infer
    infer --> vector
    vector --> index
    quality -->|"No"| fallback
    fallback --> review
    review --> index

    style quality fill:#FFECBD,stroke:#FFC943
    style review fill:#FFE0C2,stroke:#FF9E42
    style index fill:#CDF4D3,stroke:#66D575

Security invariants

  1. The SharePoint ACL is copied by the connector and stored in the configured Knowledge Discovery security field.
  2. Every query carries the authenticated user's security context.
  3. Facet counts, snippets, summaries, related results, and previews are derived only from security-trimmed results.
  4. Chunks inherit the parent document security reference; no independently searchable chunk can have broader visibility.
  5. Enrichment workers may read only documents their service identity is authorized to process, and their output inherits the source ACL.
  6. Permission-only changes update security fields even when content and vectors are unchanged.
  7. Generated sensitivity labels never modify authorization automatically.

Update and delete behavior

Use (documentId, sourceVersion) as the version boundary. On a new version, write the canonical document and all new chunks, then atomically make the new version searchable and remove the old chunks. On deletion, remove the document, chunks, vectors, cached previews, and generated enrichments. On a move, retain the stable SharePoint item identity and update path-derived metadata.