Entries of only +name/-name modify the inherited tool selection, so "defaultTools": ["+codemode"] enables codemode without repeating the defaults. Project entries layer on top of user settings. Document how to enable codemode without MCP and how to use classifier models from it.
6.4 KiB
Local Models with llama.cpp
Pi supports the llama.cpp router server. The router discovers multiple GGUF models and loads or unloads them on demand.
Use a current llama.cpp build with router support. Follow the build instructions or install a prebuilt release for your platform.
Start the router
Start llama-server without --model or -m. Passing a model starts single-model mode instead of router mode.
llama-server \
--models-dir ~/models \
--no-models-autoload \
--jinja \
--host 127.0.0.1 \
--port 8080 \
-ngl 999 \
-c 32768
Important options:
--models-dir ~/modelsdiscovers local GGUF files.--no-models-autoloadkeeps loading explicit through/llama.--jinjaenables compatible chat templates and tool calling.-ngl 999offloads as many layers as possible to the GPU.-c 32768sets the context window for each loaded model. Omit it to use the model's native context, which may require substantially more memory.
A single-file model can sit directly in the model directory. Put multimodal and multi-shard models in separate subdirectories:
~/models/
├── llama-3.2-1b-Q4_K_M.gguf
├── gemma-3-4b-it-Q4_K_M/
│ ├── gemma-3-4b-it-Q4_K_M.gguf
│ └── mmproj-F16.gguf
└── large-model-Q4_K_M/
├── large-model-Q4_K_M-00001-of-00003.gguf
├── large-model-Q4_K_M-00002-of-00003.gguf
└── large-model-Q4_K_M-00003-of-00003.gguf
Restart the router after manually adding files. For per-model context sizes and other options, use llama.cpp model presets.
Configure Pi
Start Pi and configure the provider:
/login llama.cpp
Enter the router URL and optional API key. The default URL is http://127.0.0.1:8080.
If you start the router with --no-models-autoload, /login llama.cpp only stores the connection. Run /llama to load a model, then /model to select the loaded model for the current session.
Environment variables can configure the same values without /login:
export LLAMA_BASE_URL=http://127.0.0.1:8080
export LLAMA_API_KEY=optional-secret
pi
If the server uses an API key, start llama-server with the matching --api-key value. Keep --host 127.0.0.1 for local-only access.
Manage models
Run:
/llama
- Select an unloaded model to load it.
- Select a loaded model to unload it.
- Select Download model…, search Hugging Face, then choose a repository and quantization. Exact
owner/repository[:quant]values also work. - Press Escape during a load or download to confirm cancellation.
Hugging Face search uses HF_TOKEN when set, then checks $HF_TOKEN_PATH, $HF_HOME/token, $XDG_CACHE_HOME/huggingface/token, and ~/.cache/huggingface/token. Search also works without authentication, subject to lower rate limits. Pi warns before downloading gated repositories and links to their access page. The llama.cpp server performs the download, so its process must also have HF_TOKEN when the selected repository requires access.
If other models are loaded, Pi asks whether to unload them first or keep them loaded. Pi does not silently unload models and never deletes model files. The router may be shared with other clients, so /llama always displays the router's current state.
Loaded and sleeping models appear in /model. Sleeping models wake automatically when selected. With router autoload enabled, unloaded preset models also appear and load when selected. With --no-models-autoload, load a model through /llama before selecting it.
If the router disconnects, /llama shows Retry and Close. Retry reconnects and refreshes model state without replaying the interrupted operation.
Classification
Every model listed for chat is also listed as a classifier model with the same ID and the llama-cpp-classify API. Classifier models answer typed choice, bool, and score questions about JSON state, like TypeSafe's Jev models. The model reaches them from codemode scripts, and extensions through ctx.modelRegistry.classify(); see Classifier models.
The model does not generate an answer. Each question becomes one chat prompt: the state, every question of the request, the state again, and then the question with its answers under single-token labels. Labels are letters for a choice (up to 62 options), Yes/No for a bool, and digits for a score (up to 10 levels). The second copy of the state is read with the questions in view, which improved accuracy on JevBench with small models. Pi reads the probabilities of the labels as the next token and normalizes them. A choice returns every option's probability and a confidence of (n * peak - 1) / (n - 1); a score returns the expected level.
- Raw label probabilities are usually overconfident. The per-request
temperatureoption divides the label logits before normalizing; values above 1 soften the distribution. It changes no answer. - Questions run one after another. Everything before the final question is the same for all questions of a request, so the server's prompt cache evaluates it once. The state appears twice, so it needs twice its size in context.
- Small models may follow instructions written inside the state. The prompt tells the model to judge the state as data, but that is not a guarantee.
- Hybrid models such as Qwen3.5 cannot rewind a partially cached prompt without context checkpoints. If each question reprocesses the whole state, start the router with
--ctx-checkpoints 32 --checkpoint-min-step 0.
Troubleshooting
Check that the router is reachable:
curl http://127.0.0.1:8080/health
curl http://127.0.0.1:8080/models
- No models in
/llama: Check--models-dir, the directory layout, and restart the router. - Model missing from
/modelwith--no-models-autoload: Load it with/llamafirst. - Load fails or uses too much memory: Lower
-cor unload another model. - Server is not in router mode: Start it without
--model,-m, or-hf.
To remove the llama.cpp provider and /llama, disable llama.cpp under Built-in in pi config, or set "extensions": ["-builtin:llama.cpp"] in settings.