# pi-prefix-cache — Pi KV prefix checkpoint extension Reuses Pi's stable harness prefix (system-prompt sections, tool schemas, `chat_template_kwargs`) in a llama-server KV-cache slot across sessions, and can persist that prefix to disk so a restarted server (or a new Pi install) resumes in seconds instead of re-ingesting the 5–20k-token prefix. - Pi is never modified; the provider payload is never touched. - Fail-open: any error is logged and the normal request proceeds. - The extension only talks to llama-server's HTTP API; all server details live in `warmup.ts`. ## How it works Per provider request: 1. Extract the stable prefix from the `developer` message (everything before the first `\n` / `\nCurrent working directory: ` marker). 2. Compute a fingerprint (SHA-256 over model id, model file identity (size+mtime), server id, pi version, extension version, tools[], `chat_template_kwargs`, stable prefix text). 3. If this process already warmed that fingerprint, check the slot is still alive (`GET /slots`, nonzero prompt counter). A server restart clears the slot — the extension detects it and falls back to restore/re-warm instead of a false in-process hit. 4. Otherwise (miss): - with persistence on: restore `pi-prefix-.bin` into the slot, then verify in two stages: a. identical-prompt gate — re-send the prefix; a valid restore must serve ≥ 90% of it from cache. Invalid → erase + cold rebuild. b. divergent probe — same prefix with a non-empty user message, so the request must resume from a context checkpoint before the divergence point (a plain KV token match proves less on hybrid models). `cache_n > 0` verifies the checkpoint chain. `cache_n = 0` is logged loudly (usually a server-flags issue, e.g. prefix shorter than `--checkpoint-min-step`) and the restore is kept — a rebuild would produce the same chain. - no file / restore failed → cold warm-up (`n_predict: 0` request carrying only the stable prefix + an empty user message). - with persistence on: save the freshly built prefix to disk. 5. The normal Pi request proceeds and reuses the slot's KV cache. After each response, the extension polls the slot's last-prompt counters and logs the observed reuse of the normal request. ## Requirements - Pi (OpenAI-compatible provider path) pointing at a local `llama-server`. - One slot for the shared prefix. With `-np 1` there is only one — leave `PI_PREFIX_CACHE_SLOT` at `0`. - For hybrid/recurrent models (Qwen3.5/3.8 etc.), cross-prompt reuse resumes only from context checkpoints created at ubatch boundaries. `--ctx-checkpoints N` should be large enough to hold the prefix chain. `--checkpoint-min-step` is the minimum spacing between consecutive checkpoints: a larger value stores fewer checkpoints (less RAM) but a resume can only land on the last checkpoint before the divergence point, so up to `min_step` prefix tokens are re-prefilled per session start. The default 8192 is too coarse for typical 5–20k prefixes (no intermediate checkpoint → full re-prefill); `0` gives maximum reuse, an intermediate value (e.g. 4096) trades up to ~4k tokens of re-prefill for less checkpoint RAM. `-ub` sets the reuse granularity (last ubatch boundary ≤ divergence point) — a smaller `-ub` is finer at higher overhead. These are per-deployment tradeoffs (e.g. `--ctx-checkpoints 32 -ub 384` on the 27B main server, `--ctx-checkpoints 8 -ub 128` on the 9B test server). - For `PI_PREFIX_CACHE_PERSIST=1` on such models: a llama-server build with the checkpoint-sidecar patch (below). Without it, restored slots re-prefill the full prompt. ## Install 1. Copy the directory so Pi auto-loads it: ``` cp -r pi-prefix-cache ~/.pi/agent/extensions/ ``` (or per-invocation: `pi -e /path/to/pi-prefix-cache/index.ts`) 2. Set the env vars in your shell profile (read once at Pi start). Minimal production set: ```bash export PI_PREFIX_CACHE_BASE_URL=http://localhost:8080/v1 export PI_PREFIX_CACHE_PERSIST=1 export PI_PREFIX_CACHE_SLOT_DIR=/home/eaman/llama/slot_caches ``` ### Environment variables | var | default | what it does | |---|---|---| | `PI_PREFIX_CACHE_BASE_URL` | *(unset)* | llama-server API base, e.g. `http://localhost:8080/v1`. **Unset → observe-only**: the extension logs the fingerprint and skips all warm-up/restore/save. Must point at the server that serves the model Pi is using. | | `PI_PREFIX_CACHE_SLOT` | `0` | Slot id that holds the shared prefix. Warm-up, restore, save and erase all target this slot, and the normal Pi request must land on the same one. With `-np 1` leave it `0`. | | `PI_PREFIX_CACHE_PERSIST` | off | `"1"` enables disk save/restore. Off → runtime-only reuse (prefix lost when the server restarts). On → cold misses restore `pi-prefix-.bin` (then verify) and save a new checkpoint after a cold build. Requires the patched llama.cpp for the restore to help on hybrid models. | | `PI_PREFIX_CACHE_SLOT_DIR` | *(unset)* | The server's `--slot-save-path` directory. The extension cannot discover it over HTTP, so it must be told. Only used for stale-checkpoint cleanup. Unset (or persistence off) → no cleanup. | | `PI_PREFIX_CACHE_CLEANUP` | `1` | `"0"` disables the one-time-per-process stale-checkpoint cleanup. Cleanup only ever deletes files matching `pi-prefix-[0-9a-f]{16}.bin` (+ their `.ckpt` sidecar) — manually named files are never touched. | | `PI_PREFIX_CACHE_TTL_DAYS` | `30` | Cleanup keeps any checkpoint whose `lastUsed` (from the metadata index) is within this many days, even if its fingerprint is no longer current. | | `PI_PREFIX_CACHE_KEEP_RECENT` | `3` | Cleanup always keeps this many most-recently-used checkpoints regardless of TTL — a margin so a Pi downgrade or A/B model switch doesn't force an immediate cold rebuild. | | `PI_PREFIX_CACHE_MAX_GB` | `0` (off) | Hard cap on total checkpoint storage (`.bin` + `.ckpt`). After the TTL/recent pass, oldest kept files are deleted (never the current session's checkpoint) until under the cap. | | `PI_PREFIX_CACHE_SERVER_ID` | *(empty)* | Optional llama-server build/runtime identity, included in the fingerprint (the server exposes no JSON build endpoint, so it is env-provided). Bump it when you deploy a new llama-server build, e.g. `export PI_PREFIX_CACHE_SERVER_ID=ddddf03d-slckv2`. Forces one clean rebuild after a server upgrade. | | `PI_PREFIX_CACHE_DIR` | `~/.cache/pi/llama-prefix/` | Where the metadata index (`index.json`) lives. Data only — policy comes from the env vars above. | All numeric env vars fall back to their defaults when unset or invalid. ### Provider gate (v1.2.0) When `PI_PREFIX_CACHE_BASE_URL` is set, the extension acts **only** for models whose provider `baseUrl` in `models.json` (agent dir: `$PI_CODING_AGENT_DIR` or `~/.pi/agent`) equals it — it can only warm its own server. Other models (remote providers, other local servers) are skipped with a log line: no warm-up, no restore/save, no slot polling. If `models.json` is missing or unreadable, all requests are skipped (safe direction: never warm a server the request cannot be attributed to). ## Expected logs All lines are prefixed `[prefix-cache]`. First run for a fingerprint: ``` [prefix-cache] fingerprint=16ba1104… [prefix-cache] no checkpoint (restore HTTP 400: …) → cold build [prefix-cache] cache=MISS → warming slot=0 (cold) [prefix-cache] warm-up done: HTTP 200 prompt_n=1228 cache_n=0 [prefix-cache] SAVED pi-prefix-16ba1104….bin: n_saved=… [prefix-cache] normal request proceeding (source=warm-up) [prefix-cache] observed normal request: prompt_n=1360 cache_n=1096 (reuse 81%) ``` Later sessions (same fingerprint): ``` [prefix-cache] RESTORE pi-prefix-16ba1104….bin: n_restored=… [prefix-cache] RESTORE verified: identical 1224/1228, divergent probe 1096/1240 [prefix-cache] normal request proceeding (source=restore) ``` Within one Pi process, subsequent requests log `cache=HIT (slot N already warmed for this fingerprint)`. Verify server-side reuse with `LLAMA_SERVER_SLOTS_DEBUG=1`: `restored context checkpoint (pos_min = …, … n_past = …)`. ## Files | file | role | |---|---| | `index.ts` | Extension entrypoint. `before_provider_request` → extract stable prefix, fingerprint, warm/restore slot on miss; `after_provider_response` → observed-reuse diagnostics. Fail-open; never mutates the payload. | | `prefix.ts` | Splits the `developer` message into stable content and variable tail at the earliest of `\n` / `\nCurrent working directory: ` (trailing whitespace stripped). | | `fingerprint.ts` | Canonical JSON (sorted keys, array order preserved) + SHA-256 over format tag, model id, pi version, extension version, tools[], `chat_template_kwargs`, stable prefix. | | `warmup.ts` | All llama-server HTTP details: `warmSlot()` (prefix-only request, `n_predict: 0`, `slot: N`), `getSlotStats()` (`GET /slots`), `saveSlotFile()` / `restoreSlotFile()` / `eraseSlotFile()` (`POST /slots/{id}?action=save\|restore\|erase`). All calls are timeout-bounded (120 s warm/save/restore/erase, 5 s stats). | | `cache-store.ts` | Metadata index (`~/.cache/pi/llama-prefix/index.json`) + `cleanupStale()` pruning `pi-prefix-[0-9a-f]{16}.bin(+.ckpt)` files whose fingerprint is no longer active. Keep set = current fingerprint ∪ index entries used within TTL ∪ N most recent. Manually named files are never touched. | ## Disk persistence — server patch `/slots save|restore` alone is not enough for hybrid models: the llama state file carries tokens + KV/recurrent state, but the server-side `slot.prompt.checkpoints` chain is not serialized, so a restored slot re-prefills the full prompt. The fix is a custom llama.cpp patch (checkpoint sidecar `.ckpt`, format `SLCK` v2): - patch (current, mainline): `kv_prefix_checkpoint_sidecar_9adc7f42.patch` (base commit `9adc7f42` = upstream `latest`/mainline; tested llama.cpp v1890) - patch (legacy, older base): `llama-cpp-kv-prefix-checkpoint-sidecar_ddddf03d.patch` (base commit `ddddf03d`) — **no longer shipped in this repo**; only for the `ddddf03d` tree (external copy: `/home/eaman/llama/bug/patches/`); do not apply to mainline - apply with `patch -p1 < ` — `git apply` does not accept the `#` header - deploy **both** `llama-server` and `libllama-server-impl.so` to the server's bin directory (the code lives in the .so; replacing only the executable is a silent no-op). Replace via atomic `mv` while other servers run from that directory. With the patch, `action=save` writes `.ckpt` next to the slot file and `action=restore` rebuilds the checkpoint chain from it, validating every entry (incl. a token-list hash that rejects stale sidecars). The sidecar is transparent to this extension. Note: sidecars written by the v1 build are rejected on restore (version check) and fail open to a full re-prefill; re-save each slot once with the v2 build to regenerate. ## Notes & limitations - Boundary markers are Pi-core text; if Pi renames `` or the cwd line, extraction fails open (logged) and the cache is simply unused. - The rendered system-open marker plus newline (19 bytes before the payload content) is template/model-specific; it is implicitly pinned by the model id in the fingerprint. - Reuse granularity = last ubatch boundary ≤ token-match point (1096 of 1221 with `-ub 128`). Tune `-ub` for tighter reuse vs prompt-processing speed; ~51.5 MiB VRAM per checkpoint. - Slot counters reflect the **last** prompt, not a cumulative total; the post-request diagnostics rely on that. With two Pi instances on one server the line may attribute the other process's prompt. - No locking between concurrent Pi instances (single active session is the assumed setup, `-np 1`). - A corrupt `index.json` disables cleanup for that run (logged); checkpoint files are never deleted against an index that cannot be trusted. A missing index is fine (treated as empty). - The model fingerprint includes the GGUF file size + mtime; replacing the file at the same path forces a clean rebuild. - `piVersion()` resolves via the pi CLI entrypoint realpath; falls back to `unknown` (still deterministic within a machine). - Changing the extension version changes the fingerprint → one cold rebuild after an upgrade.