oc-prefix-cache · opencode KV prefix cache
generated 2026-10-09 12:18

oc-prefix-cache: caching initial prompt for OpenCode

This extension, plus a llama-server of b11500 or newer (around version 0.6.0-dev ), allows caching the initial prompt ingestion for new OpenCode sessions, no matter the path. The idea is that you launch the harness and the initial prompt, tools, append stay cached so the first iteration is almost instant with little PP. Llama-server introduced preserving context checkpoints across /slots save/restore with PR #26004, October 8, 2026, previous version did require a sidecar patch.

1. What it does

Every time you start an opencode session against a local llama-server, the model has to read ~8,600 tokens before it can write a word:

part of the prompt tokens changes when?
18 tool schemas (bash, read, edit, MCP tools…) 5,806 you add/edit a tool, or the MCP server changes
the agent prompt ("You are opencode…") ~1,900 opencode is upgraded
<env> (cwd, workspace root, git, today's date) ~60 you switch project, or the date rolls over
Instructions from: (your AGENTS.md) ~300 you edit AGENTS.md
<mcp_instructions> (MCP server text + live session state) ~1,100 every turn (`0 calls
the skills catalogue ~250 you add/remove a skill
the conversation itself variable obviously, every message

Reading = prefill: the server runs the model over those tokens and stores the resulting KV state. On this machine that costs ~27 seconds (287 tok/s on a 27B IQ4_XS at ~7.7k tokens).

The plugin's whole job: do that reading once, keep the result, and hand it back on the next session start.

And the payoff is bigger than the clock: the first thing you ask the model no longer waits behind a quarter-minute of prefill.


2. The two hard problems

2.1 llama-server cannot save what actually matters (hybrid models)

llama-server has a slot API — POST /slots/0?action=save writes the slot's state to a file, action=restore reads it back. That sounds like everything needed for "keep the prefix on disk".

It is not enough, and the reason is specific to hybrid/recurrent models (Qwen3.5/3.8 and relatives: a regular KV cache plus a recurrent state that is rewritten as the sequence grows).

For those models, cross-prompt reuse does not work by comparing token strings. It works by rolling the model back to a context checkpoint: the server keeps a chain of checkpoints taken at ubatch boundaries while it processes a prompt, and when a new request shares a prefix with the resident context it restores the checkpoint nearest before the divergence point and re-runs only what follows.

/slots save serialises the tokens and the KV/recurrent state. For a long time it did not serialise slot.prompt.checkpoints, so a restored slot had the weights' memory but not the rollback chain — and the server fell back to re-reading the whole prompt. The file is large, the restore is fast, and it buys you nothing.

Upstream llama.cpp fixed that in #26004, merged into mainline after tag b11499 and first released as b11500: action=save now appends the checkpoint chain to the slot file itself, and action=restore reads it back. So the plugin needs no patch — just a recent enough llama-server (llama-server --version must report b11500 or newer). A file written by an older build still restores its KV but comes back chain-less, which the plugin's probes below detect.

Three upstream changes, and only the last one concerns us:

feature PR merged
SWA context checkpoints — the --ctx-checkpoints / --swa-checkpoints flag this project sets #15293 2025-08-14
checkpoints for hybrid/recurrent models, i.e. Qwen Next and relatives #16382 2025-10-03
preserving the checkpoint chain across /slots save/restore #26004 2026-10-08

The first two gave the server the checkpoints. For the year in between, save threw the chain away, which is the gap this project carried a local patch to bridge. The third closed that gap upstream — merged 2026-10-08 — and the local patch became unnecessary the same day.

There is a second, optional server behaviour that is still local to ~/llama/llama.cpp (commits 0455de52, d6eb82d7): when several sessions share one slot, the server can swap a parked prompt back in if it matches a new request better than the resident context, and parks the context a slot restore displaces instead of destroying it. It needs --cache-ram and --ctx-checkpoints. The plugin works without it — you just lose the RAM swap between harnesses.

2.2 The harness builds the prompt, so the prefix moves under you

This is the part that makes naive prefix caching fragile. The plugin does not control the prompt — opencode does — and opencode's system message is a single blob with volatile blocks inside it, not a stable head with a variable tail:

[agent prompt] [model prompt] <env> "Instructions from:" <mcp_instructions> [skills]

Two consequences:

  1. You cannot just hash the system prompt. The MCP block contains live session state (0 calls | 0 tok saved), so it changes on every turn. A fingerprint over the whole system string would miss on every single request — the cache would never be used. So the plugin cuts the message at the first volatile marker (<env> by default) and fingerprints only what is left. The cut point is deliberately the earliest volatile marker: everything before it is identical for every project and every day, which is what makes one cached prefix serve everything.

  2. Anything that changes the rendered token stream invalidates the cache. Not just the system text — the chat template renders the tools before the system content, so for this model the tool schemas are 5,806 of the 8,619 tokens. Send a warm-up request with a slightly different tools array and the two renders diverge at token ~3: the warm KV is useless and you have paid 27 s for nothing. (This is not hypothetical — it was the first real bug in the pi version of this project.)

So the plugin fingerprints everything that can move the boundary: model id, the GGUF's size+mtime, the llama-server build id, the opencode version, its own version, the tools array verbatim, and the stable head. Any change ⇒ a new fingerprint ⇒ one deliberate cold rebuild instead of a silent mis-cache.


3. How the plugin talks to the server

All of it is llama-server's HTTP API, timeout-bounded so a hung server fails open (the request proceeds normally, the cache is simply unused):

call purpose
POST /v1/chat/completions with n_predict: 0, slot: N, cache_prompt: true evaluate the prefix and leave its KV in the slot. Carries the truncated system message, an empty user turn (the template refuses to render without one) and the tools verbatim.
GET /slots is the slot alive, and how much of the last prompt came from cache
POST /slots/{id}?action=save|restore|erase the disk checkpoint: the slot's state and its context-checkpoint chain, in one file

Interception, not re-implementation

opencode's plugin API offers hooks, but none of them can hold a request or show the final wire body:

So the plugin's config hook injects a fetch wrapper into the one provider whose baseURL matches OC_PREFIX_CACHE_BASE_URL. It sees the exact bytes opencode sends, warms the slot, then forwards the request untouched. That also makes the provider gate free: we can only ever see requests to our own server. And it makes the "byte-identical body" property testable, which the offline gate test asserts.

Trusting a restored file

A restore can silently produce a slot that looks fine and re-prefills everything (that was the whole point of §2.1). So a restore is never trusted:

  1. identical-prompt gate — re-send the prefix; a valid restore must serve ≥ 90 % of it from cache. Below that: erase and rebuild.
  2. divergent probe — same prefix but with a non-empty user turn, so the request cannot be satisfied by a plain token match; it must resume from a context checkpoint before the divergence point (cache_n > 0). A zero here is logged loudly — almost always a server-flag problem, e.g. --checkpoint-min-step larger than the prefix, which is exactly why the launcher now passes 4096 instead of the stock 8192.

4. Slots, checkpoints and sessions

A llama-server "slot" is one independent context: -np 1 means one slot with n_ctx ≈ 95,744 tokens. Inside a slot the server keeps a chain of context checkpoints taken while processing a prompt:

The plugin targets one slot per harness (OC_PREFIX_CACHE_SLOT, default 0), and warm-up, restore, save and erase all address that slot — as does the real request, which lands there because there is only one. Two harnesses on one slot (pi and opencode today) evict each other; the optional best-match/park commits mitigate it, and -np 2 with one slot each is the clean answer.

5. Disk: meaningful storage, and cleaning it up relatively

Per cached prefix, on this 27B:

artifact size written by
oc-prefix-<fp16>.bin 881 MB llama-server, action=save — state payload plus the context-checkpoint chain (upstream #26004)
~/.cache/opencode/llama-prefix/index.json ~0.5 KB the plugin: fingerprint → filename, model, slot, token count, created/last-used

≈ 880 MB per fingerprint. The filename is the fingerprint, which is why the cleanup can be relative rather than heuristic — the plugin never has to guess whether a file is still wanted:

Because the default cut is <env>, there is normally exactly one fingerprint per (model, opencode version, tools, server build) — not one per project — so the steady-state footprint is ~875 MB, not a multiple.


6. What it took to get here

Not a weekend of typing, but most of the work was measurement, not code:

7. Where the state lives

thing path
source (shipped files) ~/dev/harness/opencode/cache/oc-prefix-cache/
installed plugin ~/.config/opencode/oc-prefix-cache/ + ~/.config/opencode/plugin/oc-prefix-cache.ts
metadata index ~/.cache/opencode/llama-prefix/index.json
checkpoints <--slot-save-path>/oc-prefix-*.bin
plugin log /tmp/oc-prefix-cache.log
config OC_PREFIX_CACHE_* in ~/.bashrc
server launcher ~/run/scripts/one_shot.sh_ctx_long
published https://store.piffa.net/lm/ocache/