Skip to content
EN FR

Using Embeddings for Conceptual Search

Documentation status: guide — see Maturity and evidence.

This page describes the engine-level workflow without exposing private implementation classes. Binding-specific APIs should wrap the same lifecycle when those contracts are published.

1. Choose the embedding space

Define a stable name and policy for the embedding space:

registry/index: support-ticket-v3
model: <your embedding model and version>
dimension: <model dimension>
normalization: <policy>
source snapshot: <conceptual model/data version>

Do not mix vectors produced by incompatible models in the same search space.

2. Produce a vector for a conceptual value

The vector may come from a local model, an external embedding provider, a graph projection or another validated numerical pipeline.

The embedding computation and the conceptual registration are separate operations.

3. Register the conceptual identity

Register the pair:

conceptual value <-> vector

This gives the engine a reversible mapping. Treat conflicting registration as a data-integrity problem rather than silently overwriting it.

4. Insert the vector into the HNSW index

The registry resolves identity; HNSW provides approximate numerical retrieval. A conceptual search space normally needs both.

registry: identity mapping
HNSW: numerical neighborhood

For an unconstrained semantic search:

query vector
 -> HNSW KNN
 -> registry resolution
 -> conceptual candidates

For a type-constrained search:

query vector
 -> wider HNSW candidate set
 -> registry resolution
 -> hard Kind/instance filter
 -> first K compatible candidates

6. Build composite queries when one vector is not enough

Several terms can be combined with weights. Choose between:

  • one normalized weighted query vector;
  • per-term HNSW searches followed by fusion.

Then choose the scoring strategy independently from candidate generation.

Example intent:

Term 1: embedding(Document)    weight 0.3
Term 2: embedding(Alice)       weight 0.7
Hard constraint: Kind = Document

The weights are numerical preferences. The hard constraint remains symbolic.

7. Interpret the result correctly

A result should be consumed as:

conceptual identity + numerical evidence

Use the conceptual object for permissions, rules, relations, actions and H-Logic. Use the distance/score for ranking, confidence presentation or downstream numerical processing.

Operational checklist

  • keep registry and HNSW membership synchronized;
  • version embedding spaces explicitly;
  • rebuild or migrate rather than silently mixing model versions;
  • over-fetch candidates when hard filters may reject some hits;
  • do not compare uncalibrated scores across different embedding spaces;
  • log registry/index identity with search diagnostics;
  • preserve the rule that vector similarity is not logical proof.