Using Embeddings for Conceptual Search
Documentation status: guide — see Maturity and evidence.
This page describes the engine-level workflow without exposing private implementation classes. Binding-specific APIs should wrap the same lifecycle when those contracts are published.
1. Choose the embedding space
Define a stable name and policy for the embedding space:
registry/index: support-ticket-v3
model: <your embedding model and version>
dimension: <model dimension>
normalization: <policy>
source snapshot: <conceptual model/data version>
Do not mix vectors produced by incompatible models in the same search space.
2. Produce a vector for a conceptual value
The vector may come from a local model, an external embedding provider, a graph projection or another validated numerical pipeline.
The embedding computation and the conceptual registration are separate operations.
3. Register the conceptual identity
Register the pair:
conceptual value <-> vector
This gives the engine a reversible mapping. Treat conflicting registration as a data-integrity problem rather than silently overwriting it.
4. Insert the vector into the HNSW index
The registry resolves identity; HNSW provides approximate numerical retrieval. A conceptual search space normally needs both.
registry: identity mapping
HNSW: numerical neighborhood
5. Search
For an unconstrained semantic search:
query vector
-> HNSW KNN
-> registry resolution
-> conceptual candidates
For a type-constrained search:
query vector
-> wider HNSW candidate set
-> registry resolution
-> hard Kind/instance filter
-> first K compatible candidates
6. Build composite queries when one vector is not enough
Several terms can be combined with weights. Choose between:
- one normalized weighted query vector;
- per-term HNSW searches followed by fusion.
Then choose the scoring strategy independently from candidate generation.
Example intent:
Term 1: embedding(Document) weight 0.3
Term 2: embedding(Alice) weight 0.7
Hard constraint: Kind = Document
The weights are numerical preferences. The hard constraint remains symbolic.
7. Interpret the result correctly
A result should be consumed as:
conceptual identity + numerical evidence
Use the conceptual object for permissions, rules, relations, actions and H-Logic. Use the distance/score for ranking, confidence presentation or downstream numerical processing.
Operational checklist
- keep registry and HNSW membership synchronized;
- version embedding spaces explicitly;
- rebuild or migrate rather than silently mixing model versions;
- over-fetch candidates when hard filters may reject some hits;
- do not compare uncalibrated scores across different embedding spaces;
- log registry/index identity with search diagnostics;
- preserve the rule that vector similarity is not logical proof.