Hypergraph Embedding Architecture
Documentation status: architecture — see Maturity and evidence.
The current engine integrates vector representations with conceptual memory through an explicit registry and a vector index. This is more precise than treating embeddings as detached arrays in an external database.
Four distinct layers
1. conceptual identity
|
2. embedding registry
|
3. vector index / HNSW retrieval
|
4. symbolic filtering and numerical reranking
1. Conceptual identity
The hypergraph value remains the authoritative identity. It carries conceptual type, relations, context and other symbolic structure independently of its embedding.
2. Embedding registry
A registry materializes the relationship between one conceptual value and one embedding vector. The engine maintains lookup paths in both directions:
conceptual value -> vector
vector -> conceptual value
The current implementation treats a conflicting remapping as an error: a symbolic value cannot silently acquire a different vector in the same registry, and one vector cannot silently be assigned to another symbolic value. Replacing an embedding is therefore an explicit lifecycle operation rather than an accidental overwrite.
Registries are named. Reopening the same named registry resolves the same stored associations, while different registry names remain isolated. This makes it possible to keep several embedding spaces without confusing their identities.
3. HNSW vector index
Registered vectors can be inserted into an HNSW index. The index is layered and approximate: upper layers locate a promising region, then the base layer expands candidates. Search returns vectors ordered by numerical distance together with a score.
The current engine stores index state in the hypergraph and supports named indexes, so index identity can survive reopening and multiple indexes can coexist.
4. Conceptual search boundary
An HNSW hit is only useful to the conceptual layer if the vector is registered. The search bridge therefore resolves each vector hit back to its symbolic value. Unregistered vectors are skipped rather than being fabricated into conceptual results.
A second search form can enforce an expected conceptual Kind. The engine intentionally retrieves a wider numerical candidate set when necessary, rejects candidates that are not instances of the expected kind, then keeps the first K compatible results.
This gives the architecture a clear separation:
HNSW proximity proposes
symbolic compatibility admits or rejects
reranking orders admitted candidates
Distance semantics
The HNSW implementation uses Euclidean geometry and exposes squared distance in scored results. Public application logic should treat the score as an ordering signal unless a particular embedding model defines a stronger interpretation.
Do not compare raw distances across unrelated embedding spaces or model versions without an explicit calibration policy.
Lifecycle and versioning
The engine can persist registry and index structures, but application-level embedding governance still matters. A production design should identify at least:
- embedding model/version;
- registry/index name;
- vector dimension and normalization convention;
- source concept version or snapshot;
- refresh/rebuild policy;
- compatibility policy when the embedding model changes.
The current registry prevents conflicting one-to-one mappings inside one registry; it does not by itself define your model-version migration policy.
Public boundary
This page describes engine behavior validated in the current source. Internal record layouts, private class names, hash-table objects and storage implementation details are deliberately not part of the public contract.