oc-prefix-cache: caching initial prompt for OpenCode
This extension + llama.cpp patch (required for persistent local storage of the cache, otherwise it dies with the server) allows caching the initial prompt ingestion for new OpenCode sessions, no matter the path. The idea is that you launch the harness and the initial prompt, tools, append stay cached so the first iteration is almost instant with little PP.
1. What it does
Every time you start an opencode session against a local llama-server, the model has to read ~8,600 tokens before it can write a word:
| part of the prompt | tokens | changes when? |
|---|---|---|
| 18 tool schemas (bash, read, edit, MCP tools…) | 5,806 | you add/edit a tool, or the MCP server changes |
| the agent prompt ("You are opencode…") | ~1,900 | opencode is upgraded |
<env> (cwd, workspace root, git, today's date) |
~60 | you switch project, or the date rolls over |
Instructions from: (your AGENTS.md) |
~300 | you edit AGENTS.md |
<mcp_instructions> (MCP server text + live session state) |
~1,100 | every turn (`0 calls |
| the skills catalogue | ~250 | you add/remove a skill |
| the conversation itself | variable | obviously, every message |
Reading = prefill: the server runs the model over those tokens and stores the resulting KV state. On this machine that costs ~27 seconds (287 tok/s on a 27B IQ4_XS at ~7.7k tokens).
The plugin's whole job: do that reading once, keep the result, and hand it back on the next session start.
- Same server, same process, next session → 0.8 s (only the new tail is read).
- Server restarted → 0.25 s, restored from a file on disk.
- First ever run for a given configuration → 27 s, once.
And the payoff is bigger than the clock: the first thing you ask the model no longer waits behind a quarter-minute of prefill.
2. The two hard problems
2.1 llama-server cannot save what actually matters (hybrid models)
llama-server has a slot API — POST /slots/0?action=save writes the
slot's state to a file, action=restore reads it back. That sounds like
everything needed for "keep the prefix on disk".
It is not enough, and the reason is specific to hybrid/recurrent models (Qwen3.5/3.8 and relatives: a regular KV cache plus a recurrent state that is rewritten as the sequence grows).
For those models, cross-prompt reuse does not work by comparing token strings. It works by rolling the model back to a context checkpoint: the server keeps a chain of checkpoints taken at ubatch boundaries while it processes a prompt, and when a new request shares a prefix with the resident context it restores the checkpoint nearest before the divergence point and re-runs only what follows.
/slots save serialises the tokens and the KV/recurrent state. It does
not serialise slot.prompt.checkpoints. So a restored slot has the
weights' memory but not the rollback chain — and the server falls back to
re-reading the whole prompt. The file is large, the restore is fast, and
it buys you nothing.
Our fix is a single self-contained llama.cpp patch
(kv_prefix_checkpoint_sidecar_9adc7f42.patch) with two halves.
The sidecar half persists the checkpoint chain next to the slot file:
action=savealso writes<file>.ckpt(formatSLCKv2)action=restorerebuilds the chain from it, validating every entry- the header carries a hash of the token list, so a sidecar that does not belong to the restored slot is rejected instead of loading foreign recurrent state (which would silently corrupt the model's memory)
- the format is version-checked, native byte order, host-local
The RAM prompt cache half adds the prompt-cache behaviour: when a cached
prompt matches a new request better than the resident slot does, the server
swaps it in; and a slot restore parks the context it displaces instead of
destroying it. It needs --cache-ram and --ctx-checkpoints, and only
matters when several sessions share one slot. The sidecar half is what
OC_PREFIX_CACHE_PERSIST=1 needs; the prompt-cache half is optional.
2.2 The harness builds the prompt, so the prefix moves under you
This is the part that makes naive prefix caching fragile. The plugin does not control the prompt — opencode does — and opencode's system message is a single blob with volatile blocks inside it, not a stable head with a variable tail:
[agent prompt] [model prompt] <env> "Instructions from:" <mcp_instructions> [skills]
Two consequences:
You cannot just hash the system prompt. The MCP block contains live session state (
0 calls | 0 tok saved), so it changes on every turn. A fingerprint over the whole system string would miss on every single request — the cache would never be used. So the plugin cuts the message at the first volatile marker (<env>by default) and fingerprints only what is left. The cut point is deliberately the earliest volatile marker: everything before it is identical for every project and every day, which is what makes one cached prefix serve everything.Anything that changes the rendered token stream invalidates the cache. Not just the system text — the chat template renders the tools before the system content, so for this model the tool schemas are 5,806 of the 8,619 tokens. Send a warm-up request with a slightly different tools array and the two renders diverge at token ~3: the warm KV is useless and you have paid 27 s for nothing. (This is not hypothetical — it was the first real bug in the pi version of this project.)
So the plugin fingerprints everything that can move the boundary:
model id, the GGUF's size+mtime, the llama-server build id, the opencode
version, its own version, the tools array verbatim, and the stable
head. Any change ⇒ a new fingerprint ⇒ one deliberate cold rebuild
instead of a silent mis-cache.
3. How the plugin talks to the server
All of it is llama-server's HTTP API, timeout-bounded so a hung server fails open (the request proceeds normally, the cache is simply unused):
| call | purpose |
|---|---|
POST /v1/chat/completions with n_predict: 0, slot: N, cache_prompt: true |
evaluate the prefix and leave its KV in the slot. Carries the truncated system message, an empty user turn (the template refuses to render without one) and the tools verbatim. |
GET /slots |
is the slot alive, and how much of the last prompt came from cache |
POST /slots/{id}?action=save|restore|erase |
the disk checkpoint, enabled by the patch above |
Interception, not re-implementation
opencode's plugin API offers hooks, but none of them can hold a request or show the final wire body:
chat.paramsnever sees the messages or the tools;experimental.chat.system.transformgives the system string, but the tools would have to be re-serialised from zod schemas per tool — the exact fragility described above;- none of them can delay the outgoing request while the slot is prepared.
So the plugin's config hook injects a fetch wrapper into the one
provider whose baseURL matches OC_PREFIX_CACHE_BASE_URL. It sees the
exact bytes opencode sends, warms the slot, then forwards the request
untouched. That also makes the provider gate free: we can only ever
see requests to our own server. And it makes the "byte-identical body"
property testable, which the offline gate test asserts.
Trusting a restored file
A restore can silently produce a slot that looks fine and re-prefills everything (that was the whole point of §2.1). So a restore is never trusted:
- identical-prompt gate — re-send the prefix; a valid restore must serve ≥ 90 % of it from cache. Below that: erase and rebuild.
- divergent probe — same prefix but with a non-empty user turn, so
the request cannot be satisfied by a plain token match; it must resume
from a context checkpoint before the divergence point
(
cache_n > 0). A zero here is logged loudly — almost always a server-flag problem, e.g.--checkpoint-min-steplarger than the prefix, which is exactly why the launcher now passes4096instead of the stock8192.
4. Slots, checkpoints and sessions
A llama-server "slot" is one independent context: -np 1 means one slot
with n_ctx ≈ 95,744 tokens. Inside a slot the server keeps a chain of
context checkpoints taken while processing a prompt:
- how many:
--ctx-checkpoints N(78 in our launcher) - how far apart:
--checkpoint-min-step(4096 here) — larger means less checkpoint RAM but a resume can only land on the last checkpoint before the divergence point, so up tomin_stepprefix tokens are re-read - how precisely a resume can land:
-ub(128 here) — the last ubatch boundary at or before the divergence point. In practice the real request resumed at token 7,534 of a 7,666-token prefix. - each checkpoint costs real memory: the server log reports
size = 158.219 MiBfor ours, so--ctx-checkpoints 78is on the order of 12 GB of VRAM budget for the chain.
The plugin targets one slot per harness (OC_PREFIX_CACHE_SLOT,
default 0), and warm-up, restore, save and erase all address that slot —
as does the real request, which lands there because there is only one.
Two harnesses on one slot (pi and opencode today) evict each other; the
park/best-match patch mitigates it, and -np 2 with one slot each is the
clean answer.
5. Disk: meaningful storage, and cleaning it up relatively
Per cached prefix, on this 27B:
| artifact | size | written by |
|---|---|---|
oc-prefix-<fp16>.bin |
377 MB | llama-server, action=save |
oc-prefix-<fp16>.bin.ckpt |
498 MB | the sidecar patch (the checkpoint chain) |
~/.cache/opencode/llama-prefix/index.json |
~0.5 KB | the plugin: fingerprint → filename, model, slot, token count, created/last-used |
≈ 875 MB per fingerprint. The filename is the fingerprint, which is why the cleanup can be relative rather than heuristic — the plugin never has to guess whether a file is still wanted:
- only files matching
oc-prefix-[0-9a-f]{16}.bin(and their.ckpt) are ever touched; anything you named yourself is never deleted; - the current fingerprint is never deleted;
- plus everything in the index used within
OC_PREFIX_CACHE_TTL_DAYS(default 30) — so an edit to AGENTS.md or a tool change does not immediately throw away the old prefix you might switch back to; - plus the
OC_PREFIX_CACHE_KEEP_RECENTmost recent (default 3); - then
OC_PREFIX_CACHE_MAX_GB(default off) trims oldest-first until under the cap; - a corrupt or unreadable
index.jsondisables cleanup entirely for that run — cleaning against an index you cannot trust would delete valid checkpoints. A missing index is fine (treated as empty).
Because the default cut is <env>, there is normally exactly one
fingerprint per (model, opencode version, tools, server build) — not one
per project — so the steady-state footprint is ~875 MB, not a multiple.
6. What it took to get here
Not a weekend of typing, but most of the work was measurement, not code:
- Reading the harness, not guessing at it. The prompt layout was
recovered by capturing a real request body and by reading the
opencode binary's request-assembly code; the 8,619 / 5,806 / 7,666
numbers come from the server's own
/apply-template+/tokenize, and the final proof from--log-prompts-dir: the warm render is a byte-prefix of the real render, diverging exactly at the cut marker. - Two counter-intuitive facts about llama-server's numbers.
timings.prompt_ncounts only the tokens newly processed by that request (a fully cached verification reportsprompt_n = 0), and/slots.n_prompt_tokens_cachedoes not count checkpoint-resume reuse at all. Trusting either one produces gates that pass when they should fail. The authoritative signal is the server log linerestored context checkpoint (pos_min = …). - Inherited scar tissue. The pi version of this project
(
~/dev/harness/starter) lost a day to a warm-up request that omitted the tools array, and to a sidecar bug that stored bytes where the restore expected tokens. Both failure modes are now structurally impossible here — the warm request is the real request minus its tail, and the sidecar validates a token hash.
7. Where the state lives
| thing | path |
|---|---|
| source (shipped files) | ~/dev/harness/opencode/cache/oc-prefix-cache/ |
| installed plugin | ~/.config/opencode/oc-prefix-cache/ + ~/.config/opencode/plugin/oc-prefix-cache.ts |
| metadata index | ~/.cache/opencode/llama-prefix/index.json |
| checkpoints | <--slot-save-path>/oc-prefix-*.bin(+.ckpt) |
| plugin log | /tmp/oc-prefix-cache.log |
| config | OC_PREFIX_CACHE_* in ~/.bashrc |
| server launcher | ~/run/scripts/one_shot.sh_ctx_long |
| published | https://store.piffa.net/lm/ocache/ |