Skip to content
EN FR

llama.cpp Provider Lifecycle

Documentation status: architecture — see Maturity and evidence.

The llama.cpp provider separates model/session residency from request scheduler lifetime.

State model

configured, stopped, model not loaded
          |
        Start
          v
started, model loaded, scheduler running
          |
        Stop
          v
stopped, model still resident
          |
   +------+------+
   |             |
 Start       UnloadModel
   |             |
   v             v
reuse       stopped, model unloaded

On the first Start, the provider loads the configured model if no resident session exists. Subsequent Stop/Start cycles reuse that session. This avoids an unnecessary model reload for temporary service stops.

A changed model configuration while stopped invalidates the resident session and unloads it before the new configuration is accepted.

Queue and cancellation

Submitted requests enter a synchronized queue. Stopping the service:

  1. requests cancellation of the currently executing request;
  2. shuts down the queue;
  3. completes pending requests as cancelled rather than silently dropping them;
  4. waits for the scheduler worker to terminate;
  5. leaves the model/session resident.

The current public cancellation operation targets the current request. Per-request cancellation is a recommended contract evolution.

Backend thread affinity

The current macOS/FPC + Metal implementation deliberately executes llama.cpp generation on the same control thread that owns/loaded the session, while a worker handles queue scheduling. This is an implementation constraint, not an application-level LLM semantic.

Public bindings should therefore avoid promising a particular callback or native-execution thread. A future implementation should encapsulate thread affinity behind a backend execution dispatcher and validate whether a dedicated backend-owner thread can replace main-thread synchronization on each supported accelerator.

Diagnostics

The implementation currently exposes whether the service is started, whether a model is loaded, pending request count, and model load/unload counters. These counters are useful for lifecycle tests, but long-term they belong more naturally to a diagnostics surface than to the minimal cross-provider service contract.