oc-prefix-cache · opencode KV prefix cache
generated 2026-10-02 15:20

oc-prefix-cache: caching initial prompt for OpenCode

This extension + llama.cpp patch (required for persistent local storage of the cache, otherwise it dies with the server) allows caching the initial prompt ingestion for new OpenCode sessions, no matter the path. The idea is that you launch the harness and the initial prompt, tools, append stay cached so the first iteration is almost instant with little PP.

1. What it does

Every time you start an opencode session against a local llama-server, the model has to read ~8,600 tokens before it can write a word:

part of the prompt tokens changes when?
18 tool schemas (bash, read, edit, MCP tools…) 5,806 you add/edit a tool, or the MCP server changes
the agent prompt ("You are opencode…") ~1,900 opencode is upgraded
<env> (cwd, workspace root, git, today's date) ~60 you switch project, or the date rolls over
Instructions from: (your AGENTS.md) ~300 you edit AGENTS.md
<mcp_instructions> (MCP server text + live session state) ~1,100 every turn (`0 calls
the skills catalogue ~250 you add/remove a skill
the conversation itself variable obviously, every message

Reading = prefill: the server runs the model over those tokens and stores the resulting KV state. On this machine that costs ~27 seconds (287 tok/s on a 27B IQ4_XS at ~7.7k tokens).

The plugin's whole job: do that reading once, keep the result, and hand it back on the next session start.

And the payoff is bigger than the clock: the first thing you ask the model no longer waits behind a quarter-minute of prefill.


2. The two hard problems

2.1 llama-server cannot save what actually matters (hybrid models)

llama-server has a slot API — POST /slots/0?action=save writes the slot's state to a file, action=restore reads it back. That sounds like everything needed for "keep the prefix on disk".

It is not enough, and the reason is specific to hybrid/recurrent models (Qwen3.5/3.8 and relatives: a regular KV cache plus a recurrent state that is rewritten as the sequence grows).

For those models, cross-prompt reuse does not work by comparing token strings. It works by rolling the model back to a context checkpoint: the server keeps a chain of checkpoints taken at ubatch boundaries while it processes a prompt, and when a new request shares a prefix with the resident context it restores the checkpoint nearest before the divergence point and re-runs only what follows.

/slots save serialises the tokens and the KV/recurrent state. It does not serialise slot.prompt.checkpoints. So a restored slot has the weights' memory but not the rollback chain — and the server falls back to re-reading the whole prompt. The file is large, the restore is fast, and it buys you nothing.

Our fix is a single self-contained llama.cpp patch (kv_prefix_checkpoint_sidecar_9adc7f42.patch) with two halves.

The sidecar half persists the checkpoint chain next to the slot file:

The RAM prompt cache half adds the prompt-cache behaviour: when a cached prompt matches a new request better than the resident slot does, the server swaps it in; and a slot restore parks the context it displaces instead of destroying it. It needs --cache-ram and --ctx-checkpoints, and only matters when several sessions share one slot. The sidecar half is what OC_PREFIX_CACHE_PERSIST=1 needs; the prompt-cache half is optional.

2.2 The harness builds the prompt, so the prefix moves under you

This is the part that makes naive prefix caching fragile. The plugin does not control the prompt — opencode does — and opencode's system message is a single blob with volatile blocks inside it, not a stable head with a variable tail:

[agent prompt] [model prompt] <env> "Instructions from:" <mcp_instructions> [skills]

Two consequences:

  1. You cannot just hash the system prompt. The MCP block contains live session state (0 calls | 0 tok saved), so it changes on every turn. A fingerprint over the whole system string would miss on every single request — the cache would never be used. So the plugin cuts the message at the first volatile marker (<env> by default) and fingerprints only what is left. The cut point is deliberately the earliest volatile marker: everything before it is identical for every project and every day, which is what makes one cached prefix serve everything.

  2. Anything that changes the rendered token stream invalidates the cache. Not just the system text — the chat template renders the tools before the system content, so for this model the tool schemas are 5,806 of the 8,619 tokens. Send a warm-up request with a slightly different tools array and the two renders diverge at token ~3: the warm KV is useless and you have paid 27 s for nothing. (This is not hypothetical — it was the first real bug in the pi version of this project.)

So the plugin fingerprints everything that can move the boundary: model id, the GGUF's size+mtime, the llama-server build id, the opencode version, its own version, the tools array verbatim, and the stable head. Any change ⇒ a new fingerprint ⇒ one deliberate cold rebuild instead of a silent mis-cache.


3. How the plugin talks to the server

All of it is llama-server's HTTP API, timeout-bounded so a hung server fails open (the request proceeds normally, the cache is simply unused):

call purpose
POST /v1/chat/completions with n_predict: 0, slot: N, cache_prompt: true evaluate the prefix and leave its KV in the slot. Carries the truncated system message, an empty user turn (the template refuses to render without one) and the tools verbatim.
GET /slots is the slot alive, and how much of the last prompt came from cache
POST /slots/{id}?action=save|restore|erase the disk checkpoint, enabled by the patch above

Interception, not re-implementation

opencode's plugin API offers hooks, but none of them can hold a request or show the final wire body:

So the plugin's config hook injects a fetch wrapper into the one provider whose baseURL matches OC_PREFIX_CACHE_BASE_URL. It sees the exact bytes opencode sends, warms the slot, then forwards the request untouched. That also makes the provider gate free: we can only ever see requests to our own server. And it makes the "byte-identical body" property testable, which the offline gate test asserts.

Trusting a restored file

A restore can silently produce a slot that looks fine and re-prefills everything (that was the whole point of §2.1). So a restore is never trusted:

  1. identical-prompt gate — re-send the prefix; a valid restore must serve ≥ 90 % of it from cache. Below that: erase and rebuild.
  2. divergent probe — same prefix but with a non-empty user turn, so the request cannot be satisfied by a plain token match; it must resume from a context checkpoint before the divergence point (cache_n > 0). A zero here is logged loudly — almost always a server-flag problem, e.g. --checkpoint-min-step larger than the prefix, which is exactly why the launcher now passes 4096 instead of the stock 8192.

4. Slots, checkpoints and sessions

A llama-server "slot" is one independent context: -np 1 means one slot with n_ctx ≈ 95,744 tokens. Inside a slot the server keeps a chain of context checkpoints taken while processing a prompt:

The plugin targets one slot per harness (OC_PREFIX_CACHE_SLOT, default 0), and warm-up, restore, save and erase all address that slot — as does the real request, which lands there because there is only one. Two harnesses on one slot (pi and opencode today) evict each other; the park/best-match patch mitigates it, and -np 2 with one slot each is the clean answer.

5. Disk: meaningful storage, and cleaning it up relatively

Per cached prefix, on this 27B:

artifact size written by
oc-prefix-<fp16>.bin 377 MB llama-server, action=save
oc-prefix-<fp16>.bin.ckpt 498 MB the sidecar patch (the checkpoint chain)
~/.cache/opencode/llama-prefix/index.json ~0.5 KB the plugin: fingerprint → filename, model, slot, token count, created/last-used

≈ 875 MB per fingerprint. The filename is the fingerprint, which is why the cleanup can be relative rather than heuristic — the plugin never has to guess whether a file is still wanted:

Because the default cut is <env>, there is normally exactly one fingerprint per (model, opencode version, tools, server build) — not one per project — so the steady-state footprint is ~875 MB, not a multiple.


6. What it took to get here

Not a weekend of typing, but most of the work was measurement, not code:

7. Where the state lives

thing path
source (shipped files) ~/dev/harness/opencode/cache/oc-prefix-cache/
installed plugin ~/.config/opencode/oc-prefix-cache/ + ~/.config/opencode/plugin/oc-prefix-cache.ts
metadata index ~/.cache/opencode/llama-prefix/index.json
checkpoints <--slot-save-path>/oc-prefix-*.bin(+.ckpt)
plugin log /tmp/oc-prefix-cache.log
config OC_PREFIX_CACHE_* in ~/.bashrc
server launcher ~/run/scripts/one_shot.sh_ctx_long
published https://store.piffa.net/lm/ocache/