LLM Runtime Integration
Documentation status: architecture — see Maturity and evidence.
logiCells separates the LLM service contract from the inference backend. Applications should depend on the generic LLM capability, while a provider owns model loading, scheduling, token generation and backend-specific resources.
The current architecture supports two distinct local paths that should not be confused:
- the native logiCells GGUF/Transformer runtime, used by the tensor/Transformer stack;
- a llama.cpp provider, which loads a compatible local model through llama.cpp behind the generic LLM service boundary.
The canonical provider name for the llama.cpp integration is logicells.llm.llama.cpp. llama.cpp is the human-readable backend name. The previous MiniOllama terminology is being retired because the provider does not implement the Ollama server/API.
Public boundary
application / agent / process
|
v
LLM service contract
|
+-- native logiCells LLM runtime
|
+-- llama.cpp provider
|
+-- remote or future providers
Application code should not depend on native llama.cpp handles, queue objects, worker objects or session implementations. Those remain backend implementation details.
Where to continue
- LLM service contract
- Run a local model with the llama.cpp provider
- llama.cpp provider lifecycle
- LLM usage patterns
The provider naming described here is the target canonical naming for the current source rename. Until the corresponding source change lands, older internal identifiers may still appear in the implementation tree.