Run a Local Model with the llama.cpp Provider
Documentation status: roadmap — see Maturity and evidence.
This page documents the target public naming for the source rename. The underlying implementation already loads and generates through llama.cpp; the logiCells.LLM.LlamaCpp.* names are the canonical names to use once the rename is applied.
Prerequisites
You need:
- a llama.cpp native library compatible with the bindings shipped for the target platform;
- a local model file accepted by that llama.cpp build;
- a model context size greater than zero;
- the llama.cpp provider registered with the LLM service registry.
Provider identity
logicells.llm.llama.cpp
Target provider-specific factory
The provider-specific API should offer a convenience factory equivalent to:
service = CreateLlamaCppService(
modelPath,
contextSize = 8192,
useMMap = true)
Higher-level code may instead resolve the same capability through the generic provider registry.
First request
The application flow is:
service.Start()
request = {
prompt: "Explain the current conceptual context.",
maxTokens: 128,
onToken: stream token text,
onDone: observe completion
}
requestId = service.Submit(request)
The provider tokenizes the prompt, decodes it into the llama.cpp context, samples generated tokens, converts token pieces to valid UTF-8 text, and streams text fragments to the caller.
Stop versus unload
Stop stops request scheduling and cancels current/pending work, but deliberately keeps the successfully loaded model/session resident. A later Start can therefore reuse the resident model.
UnloadModel is the explicit resource-release boundary and is only valid while the service is stopped. Reconfiguring a stopped service to a different model/context/mmap configuration also releases the previous resident session.
See llama.cpp provider lifecycle for the detailed state model.
Current limitations of the public contract
The current request contract exposes a prompt and maximum-token budget but not yet the complete llama.cpp sampling surface. Temperature, top-k/top-p, seed, stop sequences, chat templates and per-request sampler policies should be added through provider-neutral generation options rather than llama.cpp-specific parameters.