Search and vectorization¶
Retrieval strategy¶
Vectors are one retrieval signal, not the search system. A production query should apply security and structured constraints, retrieve lexical and semantic candidates, combine their scores, group chunks back to documents, and return evidence.
| Query shape | Primary behavior | Vector behavior |
|---|---|---|
| Identifier or drawing number | Exact match in title, filename, reference, and document-number fields | Skip or use negligible weight |
| Short keyword query | Keyword/conceptual search with field boosts | Add a smaller semantic contribution |
| Natural-language question | Hybrid lexical and vector retrieval | Use normal semantic contribution |
| Filter-only navigation | Parametric/date/numeric fields | Do not generate a vector |
Start with simple rules. A query containing a configured document-number pattern or an exact filename should not require an LLM intent classifier.
Query flow¶
sequenceDiagram
title Secured hybrid document search
participant User
participant SearchUI
participant QueryService
participant EmbeddingModel
participant KDContent
participant KDView
User->>SearchUI: Enter query and filters
SearchUI->>QueryService: Query with security context
QueryService->>EmbeddingModel: Create query vector
EmbeddingModel-->>QueryService: Matching model vector
QueryService->>KDContent: Hybrid query with SecurityInfo
KDContent-->>QueryService: Security-trimmed hits
QueryService->>KDView: Request highlighted preview
KDView-->>QueryService: Authorized preview
QueryService-->>SearchUI: Grouped results and evidence
SearchUI-->>User: Authorized results
For an identifier or filter-only query, the query service omits the embedding step. The rest of the security and result-processing path is unchanged.
Search field roles¶
| Knowledge Discovery role | Canonical data |
|---|---|
| Reference | Document ID, source version, canonical URL, current reference |
| Index text | Title, headings, extracted body, OCR, transcript, vetted entity text |
| Parametric | MIME, file type, document type, site, project, department, language, review state |
| Numeric | File size, page/slide/sheet counts where meaningful |
| Date | Created, modified, last indexed, enrichment timestamp |
| Vector | Repeated semantic chunk vectors with source/location metadata |
| Security | Connector-produced ACL/security fields |
| Print/display | Title, URL, type, modified date, summary, extraction status, matched location |
Do not make high-cardinality opaque identifiers into facets. Do not place ACL principals, raw metadata, query URLs, or vector values in display fields.
What to vectorize¶
Vectorize text that carries meaning a user might express differently in a query:
- complete paragraphs and sections;
- slide text with its title;
- table row groups with column headers;
- drawing title blocks, annotations, and visible text;
- OCR text and evidence-backed captions;
- transcript segments;
- email subject and body;
- short, controlled taxonomy labels when they add context.
Each embedding input may include a small context prefix containing the document title, stable human-readable folder labels, document type, section heading, and chunk location. Keep those same structured values as facets too.
What not to vectorize¶
- the raw JSON record or arbitrary concatenation of every metadata value;
- ACLs, user/group names, security tokens, or generated sensitivity labels;
- timestamps, file sizes, status codes, hashes, opaque IDs, and request URLs;
semanticClusterX,semanticClusterY, distances, or outlier flags;- binary PDF, Office, CAD, image, video, email, or ZIP bytes;
- repeated boilerplate, navigation, headers, and footers;
- content capped at an export or indexing limit and presented as complete.
These values belong in exact fields, filters, security fields, diagnostics, or nowhere in the user-facing index.
Chunking defaults¶
Use logical boundaries before token windows. The values below are starting defaults to be tuned with the judged query set, not universal truths.
| Format | Search unit | Initial rule |
|---|---|---|
| PDF and Word | Heading-aware prose | About 600 tokens with 100-token overlap; retain page and heading |
| PowerPoint | Slide | One slide per chunk; merge only very small adjacent slides |
| Excel and CSV | Table or row group | Repeat headers; cap around 800 tokens; keep numeric columns structured |
| DWG and DGN | Layout, model, or title block | Extract number, revision, layers, annotations, and visible text; never embed binary CAD |
| Images | OCR region or caption | Retain page/region coordinates and extraction confidence |
| Video and audio | Transcript interval | 60–90 seconds with timestamps and speaker when available |
| MSG/email | Subject and body section | Keep sender/date structured; process attachments as child documents |
| URL | Permitted target content | Store the shortcut as a relation; index the authorized target |
| ZIP/archive | Permitted member | Index safe supported members as child documents; keep an inventory |
Short documents below one useful chunk remain one chunk. Long pages or sheets are split without discarding headings, column names, page references, or timestamps needed to explain the match.
Contentless documents¶
When extraction fails, create a separate metadata fallback representation from filename, readable folder labels, file type, and reliable source facets. It may have a vector, but it must:
- be marked
metadata_only; - use a distinct vector field or source marker so it can be down-weighted;
- never claim to summarize unseen content;
- surface extraction failure to operators and, where useful, to users;
- be replaced when successful extraction becomes available.
This allows a query such as DRAWING-EXAMPLE-001 or “CAD drawings in the technical
folder” to find a contentless drawing without pretending that its design was
semantically analysed.
Hybrid ranking¶
Use Knowledge Discovery's normal query fields and VECTOR operator against
the same security-trimmed index. Retrieve enough candidates from both lexical
and vector signals, then combine using configured weights or reciprocal-rank
fusion in the query layer. Group repeated chunk matches by document and retain
the best matching chunk plus a small number of distinct supporting locations.
Tune weights by query class against the evaluation set. Do not hard-code a single vector-heavy score for all searches. Exact identifier matches, current versions, title matches, and authoritative facets should be allowed to outrank semantic similarity.
Model lifecycle¶
- Select a multilingual model using the production languages plus identifier-heavy and technical queries from the full corpus.
- Host it inside the trusted boundary.
- Store model name, immutable model version, vector dimensions, normalization choice, chunking version, and creation time.
- Use the identical model and preprocessing for index and query vectors.
- Re-embed when content, source version, model version, or chunking version changes.
- Maintain old and new vector fields during a model migration; switch only after evaluation, then remove the old field.
Visual embeddings and video frame embeddings are deferred. Add them only for a validated “find visually similar” use case that text/OCR retrieval cannot meet.