Skip to content

Search and vectorization

Retrieval strategy

Vectors are one retrieval signal, not the search system. A production query should apply security and structured constraints, retrieve lexical and semantic candidates, combine their scores, group chunks back to documents, and return evidence.

Query shape Primary behavior Vector behavior
Identifier or drawing number Exact match in title, filename, reference, and document-number fields Skip or use negligible weight
Short keyword query Keyword/conceptual search with field boosts Add a smaller semantic contribution
Natural-language question Hybrid lexical and vector retrieval Use normal semantic contribution
Filter-only navigation Parametric/date/numeric fields Do not generate a vector

Start with simple rules. A query containing a configured document-number pattern or an exact filename should not require an LLM intent classifier.

Query flow

sequenceDiagram
    title Secured hybrid document search
    participant User
    participant SearchUI
    participant QueryService
    participant EmbeddingModel
    participant KDContent
    participant KDView

    User->>SearchUI: Enter query and filters
    SearchUI->>QueryService: Query with security context
    QueryService->>EmbeddingModel: Create query vector
    EmbeddingModel-->>QueryService: Matching model vector
    QueryService->>KDContent: Hybrid query with SecurityInfo
    KDContent-->>QueryService: Security-trimmed hits
    QueryService->>KDView: Request highlighted preview
    KDView-->>QueryService: Authorized preview
    QueryService-->>SearchUI: Grouped results and evidence
    SearchUI-->>User: Authorized results

For an identifier or filter-only query, the query service omits the embedding step. The rest of the security and result-processing path is unchanged.

Search field roles

Knowledge Discovery role Canonical data
Reference Document ID, source version, canonical URL, current reference
Index text Title, headings, extracted body, OCR, transcript, vetted entity text
Parametric MIME, file type, document type, site, project, department, language, review state
Numeric File size, page/slide/sheet counts where meaningful
Date Created, modified, last indexed, enrichment timestamp
Vector Repeated semantic chunk vectors with source/location metadata
Security Connector-produced ACL/security fields
Print/display Title, URL, type, modified date, summary, extraction status, matched location

Do not make high-cardinality opaque identifiers into facets. Do not place ACL principals, raw metadata, query URLs, or vector values in display fields.

What to vectorize

Vectorize text that carries meaning a user might express differently in a query:

  • complete paragraphs and sections;
  • slide text with its title;
  • table row groups with column headers;
  • drawing title blocks, annotations, and visible text;
  • OCR text and evidence-backed captions;
  • transcript segments;
  • email subject and body;
  • short, controlled taxonomy labels when they add context.

Each embedding input may include a small context prefix containing the document title, stable human-readable folder labels, document type, section heading, and chunk location. Keep those same structured values as facets too.

What not to vectorize

  • the raw JSON record or arbitrary concatenation of every metadata value;
  • ACLs, user/group names, security tokens, or generated sensitivity labels;
  • timestamps, file sizes, status codes, hashes, opaque IDs, and request URLs;
  • semanticClusterX, semanticClusterY, distances, or outlier flags;
  • binary PDF, Office, CAD, image, video, email, or ZIP bytes;
  • repeated boilerplate, navigation, headers, and footers;
  • content capped at an export or indexing limit and presented as complete.

These values belong in exact fields, filters, security fields, diagnostics, or nowhere in the user-facing index.

Chunking defaults

Use logical boundaries before token windows. The values below are starting defaults to be tuned with the judged query set, not universal truths.

Format Search unit Initial rule
PDF and Word Heading-aware prose About 600 tokens with 100-token overlap; retain page and heading
PowerPoint Slide One slide per chunk; merge only very small adjacent slides
Excel and CSV Table or row group Repeat headers; cap around 800 tokens; keep numeric columns structured
DWG and DGN Layout, model, or title block Extract number, revision, layers, annotations, and visible text; never embed binary CAD
Images OCR region or caption Retain page/region coordinates and extraction confidence
Video and audio Transcript interval 60–90 seconds with timestamps and speaker when available
MSG/email Subject and body section Keep sender/date structured; process attachments as child documents
URL Permitted target content Store the shortcut as a relation; index the authorized target
ZIP/archive Permitted member Index safe supported members as child documents; keep an inventory

Short documents below one useful chunk remain one chunk. Long pages or sheets are split without discarding headings, column names, page references, or timestamps needed to explain the match.

Contentless documents

When extraction fails, create a separate metadata fallback representation from filename, readable folder labels, file type, and reliable source facets. It may have a vector, but it must:

  • be marked metadata_only;
  • use a distinct vector field or source marker so it can be down-weighted;
  • never claim to summarize unseen content;
  • surface extraction failure to operators and, where useful, to users;
  • be replaced when successful extraction becomes available.

This allows a query such as DRAWING-EXAMPLE-001 or “CAD drawings in the technical folder” to find a contentless drawing without pretending that its design was semantically analysed.

Hybrid ranking

Use Knowledge Discovery's normal query fields and VECTOR operator against the same security-trimmed index. Retrieve enough candidates from both lexical and vector signals, then combine using configured weights or reciprocal-rank fusion in the query layer. Group repeated chunk matches by document and retain the best matching chunk plus a small number of distinct supporting locations.

Tune weights by query class against the evaluation set. Do not hard-code a single vector-heavy score for all searches. Exact identifier matches, current versions, title matches, and authoritative facets should be allowed to outrank semantic similarity.

Model lifecycle

  • Select a multilingual model using the production languages plus identifier-heavy and technical queries from the full corpus.
  • Host it inside the trusted boundary.
  • Store model name, immutable model version, vector dimensions, normalization choice, chunking version, and creation time.
  • Use the identical model and preprocessing for index and query vectors.
  • Re-embed when content, source version, model version, or chunking version changes.
  • Maintain old and new vector fields during a model migration; switch only after evaluation, then remove the old field.

Visual embeddings and video frame embeddings are deferred. Add them only for a validated “find visually similar” use case that text/OCR retrieval cannot meet.