oc-prefix-cache — opencode KV prefix cache
Keeps opencode's stable harness prefix (its 18 tool schemas + the stable head of the system prompt) resident in a llama-server KV-cache slot across sessions, and can persist that prefix to disk so a restarted server resumes in 0.25 s instead of 27 s on a slow GPU.
- Fail-open: any error is logged and the normal request proceeds.
- It only talks to llama-server's HTTP API; all server details live in
warmup.ts.
Measured on opencode 1.18.34 + Qwen3.8-27B (ROCm): a real agent turn is 8,619 rendered tokens; 7,666 of them (89%) are restored from the cached prefix, warm reuse costs 0.8 s and a disk restore 0.25 s.
How it works
The plugin's config hook wraps the fetch of the provider whose
baseURL matches OC_PREFIX_CACHE_BASE_URL, so it sees the exact bytes
opencode sends (including the tools array) and can hold the request
while the slot is prepared.
Per agent request:
- Skip anything that is not the agent turn (no
tools, no system message) — e.g. opencode's title generator. - Cut the system message at the first volatile marker:
<env>,Instructions from:,<mcp_instructions>or the skills block. The<mcp_instructions>block mutates every turn, so it must never be part of the fingerprint. - Fingerprint: model id, GGUF identity, server build id, opencode
version, extension version,
tools[]verbatim, stable head. - Already warm in this process? Check the slot is still alive (a server restart zeroes its prompt counter), then forward.
- Miss → with persistence on, restore
oc-prefix-<fp16>.binand verify it in two stages (≥ 90 % of the identical prompt from cache, then a divergent probe that must resume from a context checkpoint). Invalid → erase and rebuild. No checkpoint → cold warm-up (n_predict: 0, prefix + empty user turn), then save. - Forward the request untouched and report the observed reuse.
Requirements
- opencode ≥ 1.0 with a provider whose
options.baseURLpoints at your localllama-server(the default@ai-sdk/openai-compatiblepath). - One slot for the shared prefix; with
-np 1leaveOC_PREFIX_CACHE_SLOTat0. - Server flags for hybrid/recurrent models (Qwen3.5/3.8 …): cross-prompt
reuse resumes from context checkpoints at ubatch boundaries, so
--ctx-checkpoints Nmust be large enough for the prefix chain and--checkpoint-min-stepmust be smaller than the prefix. The stock default (8192) is too coarse for a ~7.7k prefix — use e.g.4096, or0for maximum reuse at the cost of more checkpoint RAM.-ubsets the reuse granularity. See Recommended llama-server flags below for a full baseline and the reasoning. - For
OC_PREFIX_CACHE_PERSIST=1on such models: a llama-server built with the checkpoint sidecar patch (below). Without it a restored slot re-prefills the whole prompt.
Recommended llama-server flags
A working baseline for a Qwen3.x hybrid model served to a code harness
(this is the shape of the launcher in ~/run/scripts/one_shot.sh_ctx_long):
--slot-save-path ~/llama/slot_caches/ # where save/restore writes .bin + .ckpt
--ctx-checkpoints 32 # checkpoint chain depth per slot
--checkpoint-min-step 4096 # spacing between checkpoints
--cache-ram 6000 # parked contexts, MiB
-np 1 # one context = the whole ctx pool
-ub 128 # reuse granularity
-b 1024
--jinja --chat-template-file chat_template_3.8.jinja
| flag | why |
|---|---|
--slot-save-path |
required for OC_PREFIX_CACHE_PERSIST=1; the extension saves and restores there (it cannot discover the path over HTTP, hence OC_PREFIX_CACHE_SLOT_DIR). |
--ctx-checkpoints N |
how many context checkpoints the slot may keep. Needed at all for cross-prompt reuse on hybrid models. N must be at least a few more than prefix / checkpoint-min-step, so the chain can hold the prefix and a few points inside the conversation. |
--checkpoint-min-step N |
minimum spacing between checkpoints. See below — this is the flag that decides how much of the prefix is reusable. |
--cache-ram N |
MiB of RAM the server may use to keep other prompts resident. See below — this is what lets several sessions coexist on one slot. |
-np N |
number of parallel contexts. See below — the VRAM/context trade-off. |
-ub N |
ubatch size; sets how precisely a resume can land. See below. |
--jinja + --chat-template-file |
render the prompt with a fixed external template instead of the model's built-in one. See below — this is what keeps the rendered prefix stable. |
How --checkpoint-min-step and -ub decide how much is cacheable
A resume does not need a perfect match; it needs a checkpoint at or before the point where the new request diverges from the cached one, plus the tokens from there to the divergence point. So:
reusable ≈ (position of the last checkpoint ≤ divergence point)
rounded down to a ubatch boundary
--checkpoint-min-stepsets how far apart checkpoints may be. A resume can only land on the last checkpoint before the divergence point, so up tomin_steptokens of the prefix are re-read on every start. Smaller ⇒ finer landings, less re-reading, more checkpoint VRAM. The stock8192is coarser than a typical harness prefix (this one is 7,666 tokens): there would be no checkpoint inside the prefix at all, and a restored slot would re-read all of it.4096puts one in the middle and lets later checkpoints track the conversation;0(no minimum) maximises reuse at the cost of VRAM. An intermediate value is the usual choice.-ub(ubatch size) is the finishing granularity: checkpoints land on ubatch boundaries, so the resume position is the last boundary at or before the divergence point. Smaller-ub= finer landing = fewer wasted tokens, slightly more overhead. With-ub 128the real request resumed at token 7,534 of a 7,666-token prefix — 132 tokens re-read; with-ub 512(the default) it would land on a coarser boundary.
Checkpoint memory is not free: each one holds the KV/recurrent state up
to its own position. On the 27B IQ4_XS at q8_0/q5_0 the server reported
size = 158.2 MiB for a checkpoint at 7,662 tokens, i.e. ≈21 KiB per
token, so the chain's cost scales with the positions it spans, not with
N. Budget for the largest prefixes you expect, and let the server's
fit logic (--fit-target) keep it inside VRAM.
Pin one chat template across models
Serve every Qwen 3.x model through the same external jinja template
(--jinja --chat-template-file …) rather than letting each GGUF select
its built-in one. It is sound because the cache is a property of the
rendered token stream, and the template is what produces it:
- the template decides where the tools block lands relative to the system content, the special tokens, and the tool-call syntax. This one renders tools before the system message, which is why the 18 tool schemas are 5,806 of the 8,619 tokens of a real request;
- if a model upgrade or a GGUF swap silently changed the renderer, every cached prefix would be built against a layout the server no longer produces — and nothing in the fingerprint would notice (the fingerprint pins the model id, the GGUF's size+mtime, the tools, and the plugin and opencode versions, but not the template or the server flags).
So: one template for the family, and bump OC_PREFIX_CACHE_SERVER_ID
whenever you change the template, the flags, or the server build — that
is the single knob that invalidates every checkpoint exactly once.
--cache-ram with -np 1: many sessions, one at a time
With -np 1 there is a single context, but the RAM prompt cache (needs
the best-match/park patch) lets the server keep other prompts around:
when a request would evict the resident context, the server parks it in
RAM (up to --cache-ram MiB), and when a later request matches a parked
prompt better than what is resident, it swaps that one back in. So a
handful of non-concurrent sessions — one per project, say — can each keep
a warm context without re-reading, as long as they are not in flight at
the same moment. Watch for parked displaced context in prompt cache in
the server log; the plugin's log will then show a fast
warm-up done: … from_cache=… even though the slot itself changed
contents.
Budget it against your prefixes: ~21 KiB/token means a 7.7k-token context
is roughly 160 MiB of KV plus the recurrent state, so 6000 MiB holds
comfortably more than a dozen; 0 disables the cache, -1 removes the
limit (not advisable — it is host RAM, and it grows silently).
-np is a VRAM-versus-concurrency dial
The KV/recurrent pool is sized by the total context the server
allocates; -np only decides how it is divided:
-np 1: the whole pool is one context. Maximum context length per agent, minimal per-slot overhead, but only one agent can generate at a time — a second concurrent request queues behind the first.-np 2or more: the same pool is split, so each agent gets roughlytotal / npcontext (this server reportsn_ctx ≈ 95,744on-np 1; halve it per slot on-np 2) in exchange for real parallelism — two agents, or an agent and a subagent, generating at once. Asking for the same per-slot context as-np 1means asking fornp ×the VRAM, which is usually what actually breaks the fit.
For this plugin, more slots is strictly better if two harnesses share the
server: point each at its own slot with OC_PREFIX_CACHE_SLOT (0 for
pi, 1 for opencode) and neither evicts the other.
Install
Download and unpack (the tarball is the
oc-prefix-cache/directory exactly as it lives in the source repo — copy it as-is):wget https://store.piffa.net/lm/ocache/oc-prefix-cache.tgz tar xzf oc-prefix-cache.tgz mkdir -p ~/.config/opencode/plugin cp -r oc-prefix-cache ~/.config/opencode/That gives
~/.config/opencode/oc-prefix-cache/with the five modules (index.ts,prefix.ts,fingerprint.ts,warmup.ts,cache-store.ts), thisREADME.md, and the llama.cpp patch used in the build section below.Write the shim. opencode auto-discovers any
*.tsinplugin/, so noopencode.jsonedit is needed:printf 'export { default } from "../oc-prefix-cache/index.ts"\n' \ > ~/.config/opencode/plugin/oc-prefix-cache.tsExport the variables below in your shell profile — they are read once, at opencode start.
Restart opencode.
opencode --puredisables plugins, so the cache is inert in that mode.
Configuration
All configuration is environment variables. Nothing in opencode.json
has to change.
# required: enables the plugin and acts as the provider gate
export OC_PREFIX_CACHE_BASE_URL=http://127.0.0.1:8080/v1
# recommended
export OC_PREFIX_CACHE_PERSIST=1
export OC_PREFIX_CACHE_SLOT_DIR=/home/eaman/llama/slot_caches
# bump when you deploy a new llama-server build
export OC_PREFIX_CACHE_SERVER_ID=9adc7f42-slckv2-ckpt4096
# lets the fingerprint notice a replaced GGUF
export OC_PREFIX_CACHE_MODEL_PATH=/path/to/model.gguf
| variable | default | what it does |
|---|---|---|
OC_PREFIX_CACHE_BASE_URL |
(unset) | llama-server API base. Unset → the plugin does nothing at all (not even interception). Must equal a provider's options.baseURL; that match is the provider gate, so the plugin can only ever warm its own server. |
OC_PREFIX_CACHE_SLOT |
0 |
slot that holds the shared prefix. Warm-up, restore, save, erase and the real request all target it. With -np 1 keep 0. |
OC_PREFIX_CACHE_PERSIST |
off | 1 enables disk save/restore (needs the patched llama-server below). |
OC_PREFIX_CACHE_SLOT_DIR |
(unset) | the server's --slot-save-path. Only used for stale-checkpoint cleanup; it cannot be discovered over HTTP, so it must be told. |
OC_PREFIX_CACHE_CLEANUP |
1 |
0 disables the once-per-process cleanup. |
OC_PREFIX_CACHE_TTL_DAYS |
30 |
cleanup keeps checkpoints used within N days even if their fingerprint is no longer current. |
OC_PREFIX_CACHE_KEEP_RECENT |
3 |
cleanup always keeps the N most recently used, so a config change does not force an immediate rebuild. |
OC_PREFIX_CACHE_MAX_GB |
0 (off) |
hard cap on total checkpoint storage (.bin + .ckpt); oldest kept files go first, never the current one. |
OC_PREFIX_CACHE_SERVER_ID |
(empty) | llama-server build identity, part of the fingerprint. Bump it on every server rebuild — checkpoint state is coupled to the server implementation. |
OC_PREFIX_CACHE_BOUNDARY |
env |
which marker to cut the system prompt at: env | instructions | mcp | skills. env (default) gives one cached prefix for every project and day; mcp caches ~340 tokens more but rebuilds per project and per day. |
OC_PREFIX_CACHE_MODEL_PATH |
(unset) | GGUF path; its size+mtime join the fingerprint so replacing the model under the same id forces a clean rebuild. |
OC_PREFIX_CACHE_DIR |
~/.cache/opencode/llama-prefix/ |
metadata index (index.json). Data only — the policy is in the variables above. |
OC_PREFIX_CACHE_LOG |
1 |
also write the log lines to /tmp/oc-prefix-cache.log (opencode's TUI owns stdout, so plugin output is otherwise invisible). |
Numeric variables fall back to their defaults when unset or invalid.
Expected logs
[prefix-cache] plugin loaded ext=1.0.0 opencode=1.18.34 base=http://127.0.0.1:8080/v1 slot=0 persist=true
[prefix-cache] interception armed on provider "local" (baseURL http://127.0.0.1:8080/v1)
[prefix-cache] request: provider=local model=qwen boundary=env@8676 stable-chars=8676 tools=18 fingerprint=7e3272f8…
[prefix-cache] cache=MISS → warming slot=0 (cold)
[prefix-cache] warm-up done: HTTP 200 total=7666 new=7666 from_cache=0 (26.7s, 287 tok/s)
[prefix-cache] SAVED oc-prefix-7e3272f88e0dfe4f.bin: n_saved=7666 (151.1 ms)
[prefix-cache] observed real request: prompt_n=8493 cache_n=7534 (89% from cache)
After a server restart (the disk-restore path):
[prefix-cache] RESTORE oc-prefix-7e3272f88e0dfe4f.bin: n_restored=7666 (254.4 ms)
[prefix-cache] RESTORE verified: identical 7662/7666, divergent probe 7656/7672
Server-side confirmation to look for in the llama-server log:
slot 0 | restored context checkpoint (pos_min = 7661, n_past = 7666, size = 158.219 MiB)
Disk persistence — the llama.cpp patch
/slots save|restore alone is not enough for hybrid models: the state
file carries the tokens and the KV/recurrent state, but the server-side
slot.prompt.checkpoints chain is not serialized, so a restored slot
re-prefills the whole prompt. Two patches fix that; apply the first,
then the second:
| patch | what it does |
|---|---|
kv_prefix_checkpoint_sidecar_9adc7f42.patch |
persists the checkpoint chain next to the slot file as <file>.ckpt (format SLCK v2) |
prompt_cache_best_match_9adc7f42.patch |
consults the RAM prompt cache when a cached prompt beats the resident slot, and parks the displaced context on a slot restore |
Both are based on llama.cpp commit 9adc7f42 (upstream latest).
prompt_cache_best_match_9adc7f42.patch is always available as the
latest rebase from
https://store.piffa.net/lm/ocache/ — a copy also sits next to this
README. The prompt-cache part of it needs a server built with
--cache-ram and --ctx-checkpoints.
With the sidecar patch, action=save writes <file>.ckpt next to the
slot file and action=restore rebuilds the checkpoint chain from it,
validating every entry (including a token-list hash that rejects stale
sidecars). The sidecar is transparent to this plugin. Sidecars written by
a v1 build are rejected on restore and fail open to a full re-prefill.
How to build
git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp
git checkout 9adc7f42
wget https://store.piffa.net/lm/picache/kv_prefix_checkpoint_sidecar_9adc7f42.patch
wget https://store.piffa.net/lm/ocache/prompt_cache_best_match_9adc7f42.patch
patch -p1 < kv_prefix_checkpoint_sidecar_9adc7f42.patch
patch -p1 < prompt_cache_best_match_9adc7f42.patch
cmake -B build
cmake --build build --config Release -j$(nproc)
Notes:
- apply the patches in that order; the second one extends the first.
- use
patch -p1(it tolerates the#comment header at the top of the files; oldergit applyversions reject it). - build with the backend you actually run (ROCm:
-DGGML_HIP=ON, Vulkan:-DGGML_VULKAN=ON) — the patches touch server code only and are backend-agnostic. - deploy both
llama-serverandlibllama-server-impl.soto your server's bin directory: the code lives in the shared object, so replacing only the executable is a silent no-op. Replace with an atomicmvwhile other servers run from that directory (a plain write givesETXTBSY). - then set
OC_PREFIX_CACHE_SERVER_IDto something new so the extension invalidates checkpoints written by the previous build.
Uninstall
rm -rf ~/.config/opencode/oc-prefix-cache
rm -f ~/.config/opencode/plugin/oc-prefix-cache.ts
rm -f ~/.cache/opencode/llama-prefix/index.json
rm -f "$OC_PREFIX_CACHE_SLOT_DIR"/oc-prefix-*.bin "$OC_PREFIX_CACHE_SLOT_DIR"/oc-prefix-*.bin.ckpt
# and unset the OC_PREFIX_CACHE_* variables
Notes & limitations
- The boundary markers are opencode's own text. If it renames
<env>/<mcp_instructions>/Instructions from:, extraction fails open (logged) and the cache is simply unused — it never serves a wrong prefix. - The
toolsarray is part of the fingerprint: adding, removing or editing a tool (including MCP tool schemas) forces one cold rebuild. - The tools are rendered before the system content by the Qwen3.8 chat template, and they are the bulk of the prefix (5,806 of 8,619 tokens). The warm request therefore carries the tools verbatim, in the same order — that is why the plugin intercepts the request instead of re-serializing schemas from hooks.
- Slot counters reflect the last prompt, not a running total, so the
observed real requestline is only meaningful for the request that follows it. - No locking between concurrent opencode processes; within one process
the warm/restore work is serialized. If another harness (e.g. pi) warms
the same slot, the two evict each other — give each its own slot
(
-np 2) if that matters. - Per cached prefix the server writes ~377 MB
.bin+ ~498 MB.ckpt; setOC_PREFIX_CACHE_MAX_GBif that matters. - A corrupt
index.jsondisables cleanup for that run (logged); checkpoints are never deleted against an index that cannot be trusted. A missing index is fine. - Changing the extension version changes the fingerprint → one cold rebuild after an upgrade.