HNSW Vector Index
Documentation status: architecture — see Maturity and evidence.
The current engine contains a native HNSW-style approximate nearest-neighbor index over float vectors. It is integrated with hypergraph storage rather than being only an external search service.
Search model
HNSW organizes vectors in several graph layers. A query starts from an entry point, moves greedily through upper layers, then performs a wider search at the base layer.
Conceptually:
query vector
-> entry point
-> greedy descent through upper layers
-> candidate expansion at layer 0
-> K nearest scored vectors
The index supports insertion, layer-aware neighborhood management and K-nearest-neighbor search.
Scored results
The engine exposes scored search results containing:
- vector identity;
- vector reference;
- squared Euclidean distance to the query.
Results are ordered by increasing distance. The unscored KNN form delegates to the same scored result order.
K and search breadth
Two controls have different roles:
Kis the maximum number of neighbors requested;- the search-breadth parameter controls how many candidates are explored while searching.
A larger search breadth generally trades more work for better recall. Current engine tests explicitly check that increasing this breadth does not reduce recall on the tested workloads. That is an implementation-quality property, not a guarantee that every dataset will have identical behavior.
Persistence and index identity
The current implementation persists HNSW structural state in the hypergraph. Tests verify that:
- an index can be reopened by name and recover its state;
- named indexes remain isolated from each other;
- the entry point exists after insertion;
- stored layer/neighborhood invariants survive reopening.
This allows the vector index to participate in the same persistent conceptual environment as the values it indexes.
Construction parameters
HNSW construction uses familiar controls such as connection counts, construction search breadth and level distribution. These are index-building choices, not conceptual semantics.
Do not encode them into business models. Keep them in search/index configuration so they can evolve without changing conceptual meaning.
Index membership versus conceptual membership
An HNSW index can technically contain a vector that is not registered as an embedding. Such a vector can participate in numerical KNN search, but the conceptual-search bridge will skip it because no symbolic identity can be recovered.
For conceptual search, keep these two structures consistent:
embedding registry
HNSW index
Maturity
Insertion, scored KNN, persistence/reopen behavior, named-index isolation, multi-dimensional tests and recall-oriented tests are present in the current engine test suite. The concrete low-level storage representation remains an internal implementation detail.