LLM Service Contract
Documentation status: architecture — see Maturity and evidence.
The LLM service contract is the stable application-facing boundary. A provider may implement it with llama.cpp, the native logiCells runtime, a remote endpoint or another inference engine.
Core lifecycle
A service exposes the following semantic operations:
configure model
-> start service
-> submit request(s)
-> stream generated text
-> complete / cancel
-> stop service
-> optionally unload model
The current contract includes model configuration, Start, Stop, explicit model unload, cancellation of current work, request submission, and basic state/diagnostic queries.
Request contract
A request carries at least:
- a caller-provided or generated request identifier;
- the prompt;
- a maximum generated-token budget;
- a token callback for streaming output;
- a completion callback;
- a callback-dispatch policy.
An empty prompt and a non-positive token budget are rejected before the request enters the backend queue.
Provider independence
Provider-specific objects must not cross this boundary. In particular, public bindings should not expose llama.cpp model/context/sampler handles or implementation scheduler objects.
The provider registry should identify the llama.cpp implementation as:
logicells.llm.llama.cpp
This lets higher-level code select the backend without coupling the application model to its implementation.
Contract evolution
The current contract is intentionally small. The code-evolution proposal supplied with this documentation recommends adding per-request cancellation/state, structured generation options and a structured completion result while keeping this provider-neutral boundary.