Index of /lm/ocache/oc-prefix-cache

[ICO]NameLast modifiedSizeDescription

[PARENTDIR]Parent Directory  -  
[TXT]cache-store.ts2026-10-02 15:13 8.2K 
[TXT]kv_prefix_checkpoint_sidecar_9adc7f42.patch2026-10-02 15:13 16K 
[TXT]prefix.ts2026-10-02 15:13 5.9K 
[TXT]index.ts2026-10-02 15:13 18K 
[TXT]fingerprint.ts2026-10-02 15:13 3.1K 
[TXT]README.md2026-10-02 15:13 19K 
[TXT]warmup.ts2026-10-02 15:13 8.1K 

Extension manual — oc-prefix-cache
oc-prefix-cache · opencode KV prefix cache
generated 2026-10-02 15:11

oc-prefix-cache — opencode KV prefix cache

Keeps opencode's stable harness prefix (its 18 tool schemas + the stable head of the system prompt) resident in a llama-server KV-cache slot across sessions, and can persist that prefix to disk so a restarted server resumes in 0.25 s instead of 27 s on a slow GPU.

Measured on opencode 1.18.34 + Qwen3.8-27B (ROCm): a real agent turn is 8,619 rendered tokens; 7,666 of them (89%) are restored from the cached prefix, warm reuse costs 0.8 s and a disk restore 0.25 s.

How it works

The plugin's config hook wraps the fetch of the provider whose baseURL matches OC_PREFIX_CACHE_BASE_URL, so it sees the exact bytes opencode sends (including the tools array) and can hold the request while the slot is prepared.

Per agent request:

  1. Skip anything that is not the agent turn (no tools, no system message) — e.g. opencode's title generator.
  2. Cut the system message at the first volatile marker: <env>, Instructions from:, <mcp_instructions> or the skills block. The <mcp_instructions> block mutates every turn, so it must never be part of the fingerprint.
  3. Fingerprint: model id, GGUF identity, server build id, opencode version, extension version, tools[] verbatim, stable head.
  4. Already warm in this process? Check the slot is still alive (a server restart zeroes its prompt counter), then forward.
  5. Miss → with persistence on, restore oc-prefix-<fp16>.bin and verify it in two stages (≥ 90 % of the identical prompt from cache, then a divergent probe that must resume from a context checkpoint). Invalid → erase and rebuild. No checkpoint → cold warm-up (n_predict: 0, prefix + empty user turn), then save.
  6. Forward the request untouched and report the observed reuse.

Requirements

Recommended llama-server flags

A working baseline for a Qwen3.x hybrid model served to a code harness (this is the shape of the launcher in ~/run/scripts/one_shot.sh_ctx_long):

--slot-save-path ~/llama/slot_caches/      # where save/restore writes .bin + .ckpt
--ctx-checkpoints 32                       # checkpoint chain depth per slot
--checkpoint-min-step 4096                 # spacing between checkpoints
--cache-ram 6000                           # parked contexts, MiB
-np 1                                      # one context = the whole ctx pool
-ub 128                                    # reuse granularity
-b 1024
--jinja --chat-template-file chat_template_3.8.jinja
flag why
--slot-save-path required for OC_PREFIX_CACHE_PERSIST=1; the extension saves and restores there (it cannot discover the path over HTTP, hence OC_PREFIX_CACHE_SLOT_DIR).
--ctx-checkpoints N how many context checkpoints the slot may keep. Needed at all for cross-prompt reuse on hybrid models. N must be at least a few more than prefix / checkpoint-min-step, so the chain can hold the prefix and a few points inside the conversation.
--checkpoint-min-step N minimum spacing between checkpoints. See below — this is the flag that decides how much of the prefix is reusable.
--cache-ram N MiB of RAM the server may use to keep other prompts resident. See below — this is what lets several sessions coexist on one slot.
-np N number of parallel contexts. See below — the VRAM/context trade-off.
-ub N ubatch size; sets how precisely a resume can land. See below.
--jinja + --chat-template-file render the prompt with a fixed external template instead of the model's built-in one. See below — this is what keeps the rendered prefix stable.

How --checkpoint-min-step and -ub decide how much is cacheable

A resume does not need a perfect match; it needs a checkpoint at or before the point where the new request diverges from the cached one, plus the tokens from there to the divergence point. So:

reusable ≈ (position of the last checkpoint ≤ divergence point)
           rounded down to a ubatch boundary

Checkpoint memory is not free: each one holds the KV/recurrent state up to its own position. On the 27B IQ4_XS at q8_0/q5_0 the server reported size = 158.2 MiB for a checkpoint at 7,662 tokens, i.e. ≈21 KiB per token, so the chain's cost scales with the positions it spans, not with N. Budget for the largest prefixes you expect, and let the server's fit logic (--fit-target) keep it inside VRAM.

Pin one chat template across models

Serve every Qwen 3.x model through the same external jinja template (--jinja --chat-template-file …) rather than letting each GGUF select its built-in one. It is sound because the cache is a property of the rendered token stream, and the template is what produces it:

So: one template for the family, and bump OC_PREFIX_CACHE_SERVER_ID whenever you change the template, the flags, or the server build — that is the single knob that invalidates every checkpoint exactly once.

--cache-ram with -np 1: many sessions, one at a time

With -np 1 there is a single context, but the RAM prompt cache (needs the best-match/park part of the patch below) lets the server keep other prompts around: when a request would evict the resident context, the server parks it in RAM (up to --cache-ram MiB), and when a later request matches a parked prompt better than what is resident, it swaps that one back in. So a handful of non-concurrent sessions — one per project, say — can each keep a warm context without re-reading, as long as they are not in flight at the same moment. Watch for parked displaced context in prompt cache in the server log; the plugin's log will then show a fast warm-up done: … from_cache=… even though the slot itself changed contents.

Budget it against your prefixes: ~21 KiB/token means a 7.7k-token context is roughly 160 MiB of KV plus the recurrent state, so 6000 MiB holds comfortably more than a dozen; 0 disables the cache, -1 removes the limit (not advisable — it is host RAM, and it grows silently).

-np is a VRAM-versus-concurrency dial

The KV/recurrent pool is sized by the total context the server allocates; -np only decides how it is divided:

For this plugin, more slots is strictly better if two harnesses share the server: point each at its own slot with OC_PREFIX_CACHE_SLOT (0 for pi, 1 for opencode) and neither evicts the other.

Install

  1. Download and unpack (the tarball is the oc-prefix-cache/ directory exactly as it lives in the source repo — copy it as-is):

    wget https://store.piffa.net/lm/ocache/oc-prefix-cache.tgz
    tar xzf oc-prefix-cache.tgz
    mkdir -p ~/.config/opencode/plugin
    cp -r oc-prefix-cache ~/.config/opencode/
    

    That gives ~/.config/opencode/oc-prefix-cache/ with the five modules (index.ts, prefix.ts, fingerprint.ts, warmup.ts, cache-store.ts), this README.md, and the llama.cpp patch used in the build section below.

  2. Write the shim. opencode auto-discovers any *.ts in plugin/, so no opencode.json edit is needed:

    printf 'export { default } from "../oc-prefix-cache/index.ts"\n' \
      > ~/.config/opencode/plugin/oc-prefix-cache.ts
    
  3. Export the variables below in your shell profile — they are read once, at opencode start.

  4. Restart opencode. opencode --pure disables plugins, so the cache is inert in that mode.

Configuration

All configuration is environment variables. Nothing in opencode.json has to change.

# required: enables the plugin and acts as the provider gate
export OC_PREFIX_CACHE_BASE_URL=http://127.0.0.1:8080/v1
# recommended
export OC_PREFIX_CACHE_PERSIST=1
export OC_PREFIX_CACHE_SLOT_DIR=/home/eaman/llama/slot_caches
# bump when you deploy a new llama-server build
export OC_PREFIX_CACHE_SERVER_ID=9adc7f42-slckv2-ckpt4096
# lets the fingerprint notice a replaced GGUF
export OC_PREFIX_CACHE_MODEL_PATH=/path/to/model.gguf
variable default what it does
OC_PREFIX_CACHE_BASE_URL (unset) llama-server API base. Unset → the plugin does nothing at all (not even interception). Must equal a provider's options.baseURL; that match is the provider gate, so the plugin can only ever warm its own server.
OC_PREFIX_CACHE_SLOT 0 slot that holds the shared prefix. Warm-up, restore, save, erase and the real request all target it. With -np 1 keep 0.
OC_PREFIX_CACHE_PERSIST off 1 enables disk save/restore (needs the patched llama-server below).
OC_PREFIX_CACHE_SLOT_DIR (unset) the server's --slot-save-path. Only used for stale-checkpoint cleanup; it cannot be discovered over HTTP, so it must be told.
OC_PREFIX_CACHE_CLEANUP 1 0 disables the once-per-process cleanup.
OC_PREFIX_CACHE_TTL_DAYS 30 cleanup keeps checkpoints used within N days even if their fingerprint is no longer current.
OC_PREFIX_CACHE_KEEP_RECENT 3 cleanup always keeps the N most recently used, so a config change does not force an immediate rebuild.
OC_PREFIX_CACHE_MAX_GB 0 (off) hard cap on total checkpoint storage (.bin + .ckpt); oldest kept files go first, never the current one.
OC_PREFIX_CACHE_SERVER_ID (empty) llama-server build identity, part of the fingerprint. Bump it on every server rebuild — checkpoint state is coupled to the server implementation.
OC_PREFIX_CACHE_BOUNDARY env which marker to cut the system prompt at: env | instructions | mcp | skills. env (default) gives one cached prefix for every project and day; mcp caches ~340 tokens more but rebuilds per project and per day.
OC_PREFIX_CACHE_MODEL_PATH (unset) GGUF path; its size+mtime join the fingerprint so replacing the model under the same id forces a clean rebuild.
OC_PREFIX_CACHE_DIR ~/.cache/opencode/llama-prefix/ metadata index (index.json). Data only — the policy is in the variables above.
OC_PREFIX_CACHE_LOG 1 also write the log lines to /tmp/oc-prefix-cache.log (opencode's TUI owns stdout, so plugin output is otherwise invisible).

Numeric variables fall back to their defaults when unset or invalid.

Expected logs

[prefix-cache] plugin loaded ext=1.0.0 opencode=1.18.34 base=http://127.0.0.1:8080/v1 slot=0 persist=true
[prefix-cache] interception armed on provider "local" (baseURL http://127.0.0.1:8080/v1)
[prefix-cache] request: provider=local model=qwen boundary=env@8676 stable-chars=8676 tools=18 fingerprint=7e3272f8…
[prefix-cache] cache=MISS → warming slot=0 (cold)
[prefix-cache] warm-up done: HTTP 200 total=7666 new=7666 from_cache=0 (26.7s, 287 tok/s)
[prefix-cache] SAVED oc-prefix-7e3272f88e0dfe4f.bin: n_saved=7666 (151.1 ms)
[prefix-cache] observed real request: prompt_n=8493 cache_n=7534 (89% from cache)

After a server restart (the disk-restore path):

[prefix-cache] RESTORE oc-prefix-7e3272f88e0dfe4f.bin: n_restored=7666 (254.4 ms)
[prefix-cache] RESTORE verified: identical 7662/7666, divergent probe 7656/7672

Server-side confirmation to look for in the llama-server log:

slot 0 | restored context checkpoint (pos_min = 7661, n_past = 7666, size = 158.219 MiB)

Disk persistence — the llama.cpp patch

/slots save|restore alone is not enough for hybrid models: the state file carries the tokens and the KV/recurrent state, but the server-side slot.prompt.checkpoints chain is not serialized, so a restored slot re-prefills the whole prompt. One self-contained patch fixes that:

kv_prefix_checkpoint_sidecar_9adc7f42.patch — applies to stock llama.cpp commit 9adc7f42 (upstream latest) and builds on its own. A copy sits next to this README. It does two things:

With the sidecar patch, action=save writes <file>.ckpt next to the slot file and action=restore rebuilds the checkpoint chain from it, validating every entry (including a token-list hash that rejects stale sidecars). The sidecar is transparent to this plugin. Sidecars written by a v1 build are rejected on restore and fail open to a full re-prefill.

How to build

git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp
git checkout 9adc7f42
wget https://store.piffa.net/lm/picache/kv_prefix_checkpoint_sidecar_9adc7f42.patch
patch -p1 < kv_prefix_checkpoint_sidecar_9adc7f42.patch
cmake -B build
cmake --build build --config Release -j$(nproc)

Notes:

Uninstall

rm -rf ~/.config/opencode/oc-prefix-cache
rm -f    ~/.config/opencode/plugin/oc-prefix-cache.ts
rm -f    ~/.cache/opencode/llama-prefix/index.json
rm -f    "$OC_PREFIX_CACHE_SLOT_DIR"/oc-prefix-*.bin "$OC_PREFIX_CACHE_SLOT_DIR"/oc-prefix-*.bin.ckpt
# and unset the OC_PREFIX_CACHE_* variables

Notes & limitations