# oc-prefix-cache — opencode KV prefix cache Keeps opencode's **stable harness prefix** (its 18 tool schemas + the stable head of the system prompt) resident in a **llama-server** KV-cache slot across sessions, and can persist that prefix to disk so a restarted server resumes in **0.25 s instead of 27 s** on a slow GPU. - Fail-open: any error is logged and the normal request proceeds. - It only talks to llama-server's HTTP API; all server details live in `warmup.ts`. Measured on opencode 1.18.34 + Qwen3.8-27B (ROCm): a real agent turn is 8,619 rendered tokens; **7,666 of them (89%) are restored from the cached prefix**, warm reuse costs 0.8 s and a disk restore 0.25 s. ## How it works The plugin's `config` hook wraps the `fetch` of the provider whose `baseURL` matches `OC_PREFIX_CACHE_BASE_URL`, so it sees the exact bytes opencode sends (including the `tools` array) and can hold the request while the slot is prepared. Per agent request: 1. Skip anything that is not the agent turn (no `tools`, no system message) — e.g. opencode's title generator. 2. Cut the system message at the first volatile marker: ``, `Instructions from:`, `` or the skills block. The `` block mutates every turn, so it must never be part of the fingerprint. 3. Fingerprint: model id, GGUF identity, server build id, opencode version, extension version, `tools[]` verbatim, stable head. 4. Already warm in this process? Check the slot is still alive (a server restart zeroes its prompt counter), then forward. 5. Miss → with persistence on, restore `oc-prefix-.bin` and verify it in two stages (≥ 90 % of the identical prompt from cache, then a divergent probe that must resume from a context checkpoint). Invalid → erase and rebuild. No checkpoint → cold warm-up (`n_predict: 0`, prefix + empty user turn), then save. 6. Forward the request untouched and report the observed reuse. ## Requirements - opencode ≥ 1.0 with a provider whose `options.baseURL` points at your local `llama-server` (the default `@ai-sdk/openai-compatible` path). - One slot for the shared prefix; with `-np 1` leave `OC_PREFIX_CACHE_SLOT` at `0`. - Server flags for hybrid/recurrent models (Qwen3.5/3.8 …): cross-prompt reuse resumes from context checkpoints at ubatch boundaries, so `--ctx-checkpoints N` must be large enough for the prefix chain and `--checkpoint-min-step` must be **smaller than the prefix**. The stock default (8192) is too coarse for a ~7.7k prefix — use e.g. `4096`, or `0` for maximum reuse at the cost of more checkpoint RAM. `-ub` sets the reuse granularity. See **Recommended llama-server flags** below for a full baseline and the reasoning. - For `OC_PREFIX_CACHE_PERSIST=1` on such models: a llama-server built with the **checkpoint sidecar patch** (below). Without it a restored slot re-prefills the whole prompt. ## Recommended llama-server flags A working baseline for a Qwen3.x hybrid model served to a code harness (this is the shape of the launcher in `~/run/scripts/one_shot.sh_ctx_long`): ``` --slot-save-path ~/llama/slot_caches/ # where save/restore writes .bin + .ckpt --ctx-checkpoints 32 # checkpoint chain depth per slot --checkpoint-min-step 4096 # spacing between checkpoints --cache-ram 6000 # parked contexts, MiB -np 1 # one context = the whole ctx pool -ub 128 # reuse granularity -b 1024 --jinja --chat-template-file chat_template_3.8.jinja ``` | flag | why | |---|---| | `--slot-save-path` | required for `OC_PREFIX_CACHE_PERSIST=1`; the extension saves and restores there (it cannot discover the path over HTTP, hence `OC_PREFIX_CACHE_SLOT_DIR`). | | `--ctx-checkpoints N` | how many context checkpoints the slot may keep. Needed at all for cross-prompt reuse on hybrid models. `N` must be at least a few more than `prefix / checkpoint-min-step`, so the chain can hold the prefix *and* a few points inside the conversation. | | `--checkpoint-min-step N` | minimum spacing between checkpoints. See below — this is the flag that decides how much of the prefix is reusable. | | `--cache-ram N` | MiB of RAM the server may use to keep *other* prompts resident. See below — this is what lets several sessions coexist on one slot. | | `-np N` | number of parallel contexts. See below — the VRAM/context trade-off. | | `-ub N` | ubatch size; sets how precisely a resume can land. See below. | | `--jinja` + `--chat-template-file` | render the prompt with a fixed external template instead of the model's built-in one. See below — this is what keeps the rendered prefix stable. | ### How `--checkpoint-min-step` and `-ub` decide how much is cacheable A resume does not need a perfect match; it needs a **checkpoint at or before the point where the new request diverges** from the cached one, plus the tokens from there to the divergence point. So: ``` reusable ≈ (position of the last checkpoint ≤ divergence point) rounded down to a ubatch boundary ``` - `--checkpoint-min-step` sets how far apart checkpoints may be. A resume can only land on the last checkpoint before the divergence point, so **up to `min_step` tokens of the prefix are re-read on every start**. Smaller ⇒ finer landings, less re-reading, more checkpoint VRAM. The stock `8192` is coarser than a typical harness prefix (this one is 7,666 tokens): there would be no checkpoint inside the prefix at all, and a restored slot would re-read all of it. `4096` puts one in the middle and lets later checkpoints track the conversation; `0` (no minimum) maximises reuse at the cost of VRAM. An intermediate value is the usual choice. - `-ub` (ubatch size) is the *finishing* granularity: checkpoints land on ubatch boundaries, so the resume position is the last boundary at or before the divergence point. Smaller `-ub` = finer landing = fewer wasted tokens, slightly more overhead. With `-ub 128` the real request resumed at token **7,534 of a 7,666-token prefix** — 132 tokens re-read; with `-ub 512` (the default) it would land on a coarser boundary. Checkpoint memory is not free: each one holds the KV/recurrent state up to its own position. On the 27B IQ4_XS at `q8_0/q5_0` the server reported `size = 158.2 MiB` for a checkpoint at 7,662 tokens, i.e. **≈21 KiB per token**, so the chain's cost scales with the positions it spans, not with `N`. Budget for the largest prefixes you expect, and let the server's fit logic (`--fit-target`) keep it inside VRAM. ### Pin one chat template across models Serve every Qwen 3.x model through the **same external jinja template** (`--jinja --chat-template-file …`) rather than letting each GGUF select its built-in one. It is sound because the cache is a property of the *rendered token stream*, and the template is what produces it: - the template decides where the tools block lands relative to the system content, the special tokens, and the tool-call syntax. This one renders tools **before** the system message, which is why the 18 tool schemas are 5,806 of the 8,619 tokens of a real request; - if a model upgrade or a GGUF swap silently changed the renderer, every cached prefix would be built against a layout the server no longer produces — and nothing in the fingerprint would notice (the fingerprint pins the model id, the GGUF's size+mtime, the tools, and the plugin and opencode versions, but **not** the template or the server flags). So: one template for the family, and **bump `OC_PREFIX_CACHE_SERVER_ID` whenever you change the template, the flags, or the server build** — that is the single knob that invalidates every checkpoint exactly once. ### `--cache-ram` with `-np 1`: many sessions, one at a time With `-np 1` there is a single context, but the RAM prompt cache (needs the best-match/park part of the patch below) lets the server keep **other** prompts around: when a request would evict the resident context, the server parks it in RAM (up to `--cache-ram` MiB), and when a later request matches a parked prompt better than what is resident, it swaps that one back in. So a handful of non-concurrent sessions — one per project, say — can each keep a warm context without re-reading, as long as they are not in flight at the same moment. Watch for `parked displaced context in prompt cache` in the server log; the plugin's log will then show a fast `warm-up done: … from_cache=…` even though the slot itself changed contents. Budget it against your prefixes: ~21 KiB/token means a 7.7k-token context is roughly 160 MiB of KV plus the recurrent state, so `6000` MiB holds comfortably more than a dozen; `0` disables the cache, `-1` removes the limit (not advisable — it is host RAM, and it grows silently). ### `-np` is a VRAM-versus-concurrency dial The KV/recurrent pool is sized by the **total** context the server allocates; `-np` only decides how it is divided: - `-np 1`: the whole pool is one context. Maximum context length per agent, minimal per-slot overhead, but only one agent can generate at a time — a second concurrent request queues behind the first. - `-np 2` or more: the same pool is split, so each agent gets roughly `total / np` context (this server reports `n_ctx ≈ 95,744` on `-np 1`; halve it per slot on `-np 2`) in exchange for real parallelism — two agents, or an agent and a subagent, generating at once. Asking for the same *per-slot* context as `-np 1` means asking for `np ×` the VRAM, which is usually what actually breaks the fit. For this plugin, more slots is strictly better if two harnesses share the server: point each at its own slot with `OC_PREFIX_CACHE_SLOT` (`0` for pi, `1` for opencode) and neither evicts the other. ## Install 1. Download and unpack (the tarball is the `oc-prefix-cache/` directory exactly as it lives in the source repo — copy it as-is): ``` wget https://store.piffa.net/lm/ocache/oc-prefix-cache.tgz tar xzf oc-prefix-cache.tgz mkdir -p ~/.config/opencode/plugin cp -r oc-prefix-cache ~/.config/opencode/ ``` That gives `~/.config/opencode/oc-prefix-cache/` with the five modules (`index.ts`, `prefix.ts`, `fingerprint.ts`, `warmup.ts`, `cache-store.ts`), this `README.md`, and the llama.cpp patch used in the build section below. 2. Write the shim. opencode auto-discovers any `*.ts` in `plugin/`, so no `opencode.json` edit is needed: ```bash printf 'export { default } from "../oc-prefix-cache/index.ts"\n' \ > ~/.config/opencode/plugin/oc-prefix-cache.ts ``` 3. Export the variables below in your shell profile — they are read **once, at opencode start**. 4. Restart opencode. `opencode --pure` disables plugins, so the cache is inert in that mode. ## Configuration All configuration is environment variables. Nothing in `opencode.json` has to change. ```bash # required: enables the plugin and acts as the provider gate export OC_PREFIX_CACHE_BASE_URL=http://127.0.0.1:8080/v1 # recommended export OC_PREFIX_CACHE_PERSIST=1 export OC_PREFIX_CACHE_SLOT_DIR=/home/eaman/llama/slot_caches # bump when you deploy a new llama-server build export OC_PREFIX_CACHE_SERVER_ID=9adc7f42-slckv2-ckpt4096 # lets the fingerprint notice a replaced GGUF export OC_PREFIX_CACHE_MODEL_PATH=/path/to/model.gguf ``` | variable | default | what it does | |---|---|---| | `OC_PREFIX_CACHE_BASE_URL` | *(unset)* | llama-server API base. **Unset → the plugin does nothing at all** (not even interception). Must equal a provider's `options.baseURL`; that match *is* the provider gate, so the plugin can only ever warm its own server. | | `OC_PREFIX_CACHE_SLOT` | `0` | slot that holds the shared prefix. Warm-up, restore, save, erase and the real request all target it. With `-np 1` keep `0`. | | `OC_PREFIX_CACHE_PERSIST` | off | `1` enables disk save/restore (needs the patched llama-server below). | | `OC_PREFIX_CACHE_SLOT_DIR` | *(unset)* | the server's `--slot-save-path`. Only used for stale-checkpoint cleanup; it cannot be discovered over HTTP, so it must be told. | | `OC_PREFIX_CACHE_CLEANUP` | `1` | `0` disables the once-per-process cleanup. | | `OC_PREFIX_CACHE_TTL_DAYS` | `30` | cleanup keeps checkpoints used within N days even if their fingerprint is no longer current. | | `OC_PREFIX_CACHE_KEEP_RECENT` | `3` | cleanup always keeps the N most recently used, so a config change does not force an immediate rebuild. | | `OC_PREFIX_CACHE_MAX_GB` | `0` (off) | hard cap on total checkpoint storage (`.bin` + `.ckpt`); oldest kept files go first, never the current one. | | `OC_PREFIX_CACHE_SERVER_ID` | *(empty)* | llama-server build identity, part of the fingerprint. **Bump it on every server rebuild** — checkpoint state is coupled to the server implementation. | | `OC_PREFIX_CACHE_BOUNDARY` | `env` | which marker to cut the system prompt at: `env` \| `instructions` \| `mcp` \| `skills`. `env` (default) gives one cached prefix for every project and day; `mcp` caches ~340 tokens more but rebuilds per project and per day. | | `OC_PREFIX_CACHE_MODEL_PATH` | *(unset)* | GGUF path; its size+mtime join the fingerprint so replacing the model under the same id forces a clean rebuild. | | `OC_PREFIX_CACHE_DIR` | `~/.cache/opencode/llama-prefix/` | metadata index (`index.json`). Data only — the policy is in the variables above. | | `OC_PREFIX_CACHE_LOG` | `1` | also write the log lines to `/tmp/oc-prefix-cache.log` (opencode's TUI owns stdout, so plugin output is otherwise invisible). | Numeric variables fall back to their defaults when unset or invalid. ## Expected logs ``` [prefix-cache] plugin loaded ext=1.0.0 opencode=1.18.34 base=http://127.0.0.1:8080/v1 slot=0 persist=true [prefix-cache] interception armed on provider "local" (baseURL http://127.0.0.1:8080/v1) [prefix-cache] request: provider=local model=qwen boundary=env@8676 stable-chars=8676 tools=18 fingerprint=7e3272f8… [prefix-cache] cache=MISS → warming slot=0 (cold) [prefix-cache] warm-up done: HTTP 200 total=7666 new=7666 from_cache=0 (26.7s, 287 tok/s) [prefix-cache] SAVED oc-prefix-7e3272f88e0dfe4f.bin: n_saved=7666 (151.1 ms) [prefix-cache] observed real request: prompt_n=8493 cache_n=7534 (89% from cache) ``` After a server restart (the disk-restore path): ``` [prefix-cache] RESTORE oc-prefix-7e3272f88e0dfe4f.bin: n_restored=7666 (254.4 ms) [prefix-cache] RESTORE verified: identical 7662/7666, divergent probe 7656/7672 ``` Server-side confirmation to look for in the llama-server log: ``` slot 0 | restored context checkpoint (pos_min = 7661, n_past = 7666, size = 158.219 MiB) ``` ## Disk persistence — the llama.cpp patch `/slots save|restore` alone is not enough for hybrid models: the state file carries the tokens and the KV/recurrent state, but the server-side `slot.prompt.checkpoints` chain is not serialized, so a restored slot re-prefills the whole prompt. One self-contained patch fixes that: **[`kv_prefix_checkpoint_sidecar_9adc7f42.patch`](https://store.piffa.net/lm/picache/kv_prefix_checkpoint_sidecar_9adc7f42.patch)** — applies to stock llama.cpp commit `9adc7f42` (upstream `latest`) and builds on its own. A copy sits next to this README. It does two things: - **Sidecar** (needed for `OC_PREFIX_CACHE_PERSIST=1`): persists the checkpoint chain next to the slot file as `.ckpt` (format `SLCK` v2). - **RAM prompt cache** (optional, only matters when several sessions share one slot): consults the cache when a cached prompt beats the resident slot, and parks the displaced context on a slot restore. Needs a server built with `--cache-ram` and `--ctx-checkpoints`. With the sidecar patch, `action=save` writes `.ckpt` next to the slot file and `action=restore` rebuilds the checkpoint chain from it, validating every entry (including a token-list hash that rejects stale sidecars). The sidecar is transparent to this plugin. Sidecars written by a v1 build are rejected on restore and fail open to a full re-prefill. ### How to build ```bash git clone https://github.com/ggml-org/llama.cpp cd llama.cpp git checkout 9adc7f42 wget https://store.piffa.net/lm/picache/kv_prefix_checkpoint_sidecar_9adc7f42.patch patch -p1 < kv_prefix_checkpoint_sidecar_9adc7f42.patch cmake -B build cmake --build build --config Release -j$(nproc) ``` Notes: - use `patch -p1` (it tolerates the `#` comment header at the top of the file; older `git apply` versions reject it). - build with the backend you actually run (ROCm: `-DGGML_HIP=ON`, Vulkan: `-DGGML_VULKAN=ON`) — the patch touches server code only and is backend-agnostic. - **deploy both** `llama-server` *and* `libllama-server-impl.so` to your server's bin directory: the code lives in the shared object, so replacing only the executable is a silent no-op. Replace with an atomic `mv` while other servers run from that directory (a plain write gives `ETXTBSY`). - then set `OC_PREFIX_CACHE_SERVER_ID` to something new so the extension invalidates checkpoints written by the previous build. ## Uninstall ```bash rm -rf ~/.config/opencode/oc-prefix-cache rm -f ~/.config/opencode/plugin/oc-prefix-cache.ts rm -f ~/.cache/opencode/llama-prefix/index.json rm -f "$OC_PREFIX_CACHE_SLOT_DIR"/oc-prefix-*.bin "$OC_PREFIX_CACHE_SLOT_DIR"/oc-prefix-*.bin.ckpt # and unset the OC_PREFIX_CACHE_* variables ``` ## Notes & limitations - The boundary markers are opencode's own text. If it renames `` / `` / `Instructions from:`, extraction fails open (logged) and the cache is simply unused — it never serves a wrong prefix. - The `tools` array is part of the fingerprint: adding, removing or editing a tool (including MCP tool schemas) forces one cold rebuild. - The tools are rendered *before* the system content by the Qwen3.8 chat template, and they are the bulk of the prefix (5,806 of 8,619 tokens). The warm request therefore carries the tools verbatim, in the same order — that is why the plugin intercepts the request instead of re-serializing schemas from hooks. - Slot counters reflect the **last** prompt, not a running total, so the `observed real request` line is only meaningful for the request that follows it. - No locking between concurrent opencode *processes*; within one process the warm/restore work is serialized. If another harness (e.g. pi) warms the same slot, the two evict each other — give each its own slot (`-np 2`) if that matters. - Per cached prefix the server writes ~377 MB `.bin` + ~498 MB `.ckpt`; set `OC_PREFIX_CACHE_MAX_GB` if that matters. - A corrupt `index.json` disables cleanup for that run (logged); checkpoints are never deleted against an index that cannot be trusted. A missing index is fine. - Changing the extension version changes the fingerprint → one cold rebuild after an upgrade.