--config. Every section is optional; omit it to use defaults.
Minimal example
Top-level sections
[server]
[runtime]
Global defaults shared by every generator. Per-generator values under [generators.model] win; otherwise these defaults apply.
[runtime.session_cache]
For ggml-llm generators, the server can persist KV cache state between requests so that a follow-up completion sharing a prompt prefix skips prompt processing.
Cache files are stored under
{cache_dir}/.session-state-cache/.
mlx-llm keeps a separate session cache under {cache_dir}/mlx-session-cache/, configured independently per generator.
[[generators]]
Each [[generators]] block declares one model the server can host. Repeat the block to host multiple. Every block has a type, an optional [generators.backend] table, and a [generators.model] table. type is one of ggml-llm, ggml-stt, ggml-tts, ggml-decision, mlx-llm, onnx-stt, or onnx-tts.
Common [generators.model] keys
Shared by all generator types (ggml-llm, ggml-stt, ggml-tts, ggml-decision, mlx-llm, onnx-stt, onnx-tts):
Honored by ggml-llm, ggml-stt, ggml-tts, and ggml-decision only.
mlx-llm derives quantization from the repo itself, and the ONNX backends (onnx-stt, onnx-tts) select weights with dtype instead — so all three ignore these:
ggml-llm (llama.cpp / GGUF)
[generators.backend] only controls backend selection and resource planning. Runtime overrides (n_ctx, n_gpu_layers, flash_attn_type, etc.) go under [generators.model].
[generators.backend]
[generators.model] — in addition to the common ggml keys above, every [runtime] key can be overridden per-generator: n_ctx, n_gpu_layers, n_batch, n_ubatch, n_threads, n_parallel, n_cpu_moe, flash_attn_type, cache_type_k, cache_type_v, kv_unified, swa_full, ctx_shift, use_mmap, use_mlock, no_extra_bufts, cpu_mask, cpu_strict, devices.
Embedding models
A dedicated embedding model needs
embedding = true. context.buttress.embedding in a local function rejects any generator that does not set it, so give the embedding model its own [[generators]] block rather than reusing the chat model’s. An embedding context runs a single native sequence, so parallel decoding slots do not apply to it. A vocabulary-only context has no compute backend at all, so it serves tokenizer calls only.
The same repo can be hosted more than once in different modes — chat, embedding, vocabulary-only. Each mode is a separate entry in the generator registry, so one consumer’s context never inherits another’s configuration. A Vector Store brick drives both of these keys remotely, without any of them appearing in this file.
Multimodal (mtmd) — auto-downloads the matching mmproj-*.gguf from the same repo:
Speculative decoding
ggml-stt (whisper.cpp)
[generators.backend]
The
snapdragon variant runs Whisper on the Hexagon NPU. It is available on Linux arm64 hosts with a working Qualcomm HTP/FastRPC runtime, and it does not bundle OpenCL — install it with --ggml-variant=snapdragon, which keeps the default CPU build alongside it as a fallback. The variant ships its own HTP libraries, and the loader points ADSP_LIBRARY_PATH at them and sets GGML_HEXAGON_NDEV to 16 unless you set either variable yourself. Hexagon always runs with flash attention, even when use_flash_attn is "off".
Hexagon acceleration for speech-to-text is experimental.
[generators.model] — in addition to the common ggml keys above:
Runtime extras — under
[runtime] for ggml-stt only:
ggml-tts (llama.cpp + codec.cpp)
No phonemizer is wired into this backend, so NeuTTS receives raw text rather than phonemes — the same limitation the on-device generator has.
[generators.backend]
[generators.model] — in addition to the common ggml keys above:
A repo that ships the backbone and the codec together is split by filename (
codec, vocoder, wavtokenizer, dac, mimi). A repo hosting several unrelated codecs cannot be resolved by quantization preference alone, so set vocoder_filename explicitly for those.
Runtime extras — under [generators.runtime] for ggml-tts:
Output cache — synthesized WAVs are cached on disk under
{cache_dir}/.tts-cache/, keyed by text + model + options, so repeated phrases return instantly. Configure it under [generators.runtime.output_cache]:
ggml-decision (llama.cpp typed decisions)
choice, score, and noul (yes/no) questions about a state with calibrated probabilities. Each answer takes one forward pass, and no tokens are generated. Devices reach it through the Typed Decision (GGML) Generator, and local functions through context.buttress.decide.
The model is checked when its context loads. A regular chat model, or a decision model whose decision type this build cannot serve, fails the load with the reason. Requests on one generator run one at a time.
[generators.backend] — same keys as ggml-tts: variant, variant_preference, gpu_memory_fraction, and cpu_memory_fraction.
[generators.model] — in addition to the common ggml keys above:
Setting any
mmproj_* key loads a projector, and requests can then pass images as file paths or data URLs. If the model’s prompt has no place for images, the server ignores the projector and logs a warning.
Runtime extras — under [generators.runtime] for ggml-decision:
mlx-llm (Apple Silicon)
[generators.backend] section for mlx-llm.
On first use, the backend creates a Python virtualenv at {cache_dir}/mlx-env and installs mlx_lm_package, mlx_vlm_package, plus torch and torchvision (required by some VLM processors). If an existing venv already has mlx_vlm and torch importable, the install step is skipped.
[generators.model] — common repo_id / revision / download plus:
quantization, filename, and preferred_quantizations are not used — the MLX repo itself determines the quantization.
Runtime extras — under [runtime] for mlx-llm:
onnx-stt (ONNX Runtime / Whisper)
onnx/ subfolder on Hugging Face into {cache_dir}/{owner}/{repo}/, and the backend streams partial transcription results before returning the final text.
[generators.backend]
If none of the configured providers initialize, the generator fails to start.
[generators.model]
Runtime extras — under
[generators.runtime]:
The global
[runtime] keys cache_dir, huggingface_token, and http_headers also apply.
onnx-tts (ONNX Runtime / Kokoro, VITS, SpeechT5)
onnx/ subfolder on Hugging Face into {cache_dir}/{owner}/{repo}/.
[generators.backend] — same keys as onnx-stt above: provider, provider_preference, gpu_memory_fraction, cpu_memory_fraction.
[generators.model]
Output cache — synthesized audio is cached on disk under
{cache_dir}/.tts-cache/, keyed by text + model + options, so repeated phrases return instantly. Configure it under [generators.runtime.output_cache]:
The
[generators.runtime] prefer_providers key and the global [runtime] keys (cache_dir, huggingface_token, http_headers) apply as with onnx-stt.
[autodiscover]
The server announces itself on UDP 8089 so Foundation devices on the same LAN can find it. Auto-discovery is on by default.
autodiscover = false to disable discovery entirely. See the autodiscovery reference for protocol details.
[env]
Environment variables applied at startup, but only if they are not already set in the system environment. System variables and command-line exports take precedence.
HUGGINGFACE_TOKEN (not HF_TOKEN). For a single token that applies to every backend regardless of variable name, set [runtime] huggingface_token instead.
Compatibility endpoints
These endpoints are experimental. The schemas, error shapes, and CORS defaults may change.
You can also enable each endpoint via env var:
ENABLE_OPENAI_COMPAT_ENDPOINT=1 or ENABLE_ANTHROPIC_MESSAGES_ENDPOINT=1.
[functions]
Local functions are experimental. The endpoints, the function file contract, and these config keys may change between releases.
.ts / .js files as MCP tools and HTTP endpoints. Off by default.
Env-var equivalents:
ENABLE_FUNCTIONS_ENDPOINT=1, BUTTRESS_FUNCTIONS_DIR=<dir>, BUTTRESS_FUNCTIONS_HOT_RELOAD=1, and BUTTRESS_FUNCTIONS_ALLOW_UNAUTHENTICATED=1.
See Local functions for the file contract, the injected context, the endpoints, and the security model.
[[agents]]
Agents are experimental. These config keys, the
/agents endpoints, and the context.agents API may change between releases.[[agents]] table exists.
Unrecognized keys pass through to generation, so
temperature and top_p work as written.
Agent definitions are validated at startup — a bad model reference, tool name, or prompt file stops the server.
See Agents for the models, tools, sessions, endpoints, CLI, and security model.
Next steps
Workspace binding
Pair the server with a BRICKS workspace and enable auth.
LAN auto-discovery
How Foundation devices find your server on the LAN.