← Browse

@onlyterp/prompt-cache-skills

B

Most popular OSS agent harnesses (Cline, Roo Code, Continue, OpenCode, Aider) leave 30-90% off your API bill on the table because their prompt-caching code is subtly wrong, off-by-default, or just…

instructionscodex

Install

agr install @onlyterp/prompt-cache-skills --target codex

Writes 1 file into AGENTS.md, pinned to git-273167ce.

  • AGENTS.md

Document

AGENTS.md

Read this if you are an AI coding agent and the user has pointed you at this repo. This file tells you what to do.

What this repo is

A collection of self-contained drop-in fixes for prompt-caching bugs in popular LLM agent harnesses (Cline, Roo Code, Continue, OpenCode, Aider). Each fix lives at skills/<slug>/SKILL.md and is everything you need to land the patch: target file, symptom, diff, verification.

What you should do

When the user asks you to apply this repo:

Step 1 — identify which harnesses are in their project

Quick checks:

HarnessDetection
Clinepackage.json contains "name": "claude-dev" or "cline"; or path contains src/core/api/providers/anthropic.ts
Roo Codepackage.json contains "roo-cline" or "roo-code"; path src/api/providers/
Aiderpyproject.toml or setup.py declares aider; path aider/coders/
OpenCodepath packages/opencode/src/provider/transform.ts
Continuepath packages/openai-adapters/src/apis/Anthropic.ts or core/llm/llms/Bedrock.ts

If none match, the user may want you to apply skills to their own agent harness using these as templates. Ask them.

Step 2 — for each matching harness, check each relevant skill

For every skills/<slug>/SKILL.md:

  1. Read the YAML frontmatter:
    • target_harness — matches user's project?
    • target_files — exist at those paths?
    • target_commit — is the user's checkout near that commit, or have things drifted?
  2. Read the "Symptom" section. Verify the symptom exists in the user's code before applying. If it doesn't (e.g. user is on a newer version where the bug is fixed), SKIP that skill and report.
  3. If symptom present, apply the diff in the "Fix" section. If the diff doesn't apply cleanly (line drift), do the equivalent edit manually preserving the intent.
  4. Run the "Verify" section. Most verifications require capturing the wire (mitmproxy or harness debug log) and inspecting response usage fields. If you can't run a live verify, at minimum confirm the syntax of the patched file is valid (compile/type-check).

Step 3 — report back

Tell the user:

  • Which skills applied cleanly
  • Which skipped and why (symptom not present / file moved / etc.)
  • Which failed verification and were reverted
  • Estimated savings based on the per-skill estimated_savings field

Step 4 — don't bundle skills into one commit

One skill = one commit (or PR). Easy to review and easy to revert individually if one regresses. Use the skill name as the commit message subject.

What you should NOT do

  • Don't modify the skills themselves unless the user explicitly asks. If you spot a bug in a skill, write it up as a comment or patch and stop there.
  • Don't apply skills speculatively. Every skill has a target file + symptom check; verify before applying. The fixes are surgical, not heuristic.
  • Don't replace cache_control with random UUIDs or session IDs thinking it'll "randomize the cache". That's the #1 footgun. See docs/gotchas.md #9b.
  • Don't combine skills into one mega-fix. Each one is atomic for a reason — separate bugs can have separate reviewers and separate upstream PRs.
  • Don't open upstream PRs without the user's explicit go-ahead. The skills are written to be applied locally; pushing them upstream is a separate decision the user owns.

Useful reference paths

If you need to understand WHY a fix matters before applying:

  • docs/concepts/anthropic.md — Anthropic prompt caching mechanics
  • docs/concepts/openai.md — OpenAI prompt caching + prompt_cache_key
  • docs/concepts/gemini.md — Gemini implicit + explicit caching
  • docs/concepts/bedrock.md — Bedrock cachePoint semantics
  • docs/gotchas.md — 16 numbered failure modes
  • docs/verification.md — how to confirm caching is working

For the audit evidence behind each skill:

  • audits/<harness>.md — full source audit + permalinks

Verifying your work

The repo ships tools/check_cache.py — a zero-dependency Python script that fires any request body twice and dumps the cache token diff. Use it as a smoke test:

# After applying a skill that targets Anthropic:
python3 tools/check_cache.py --provider anthropic --body /tmp/req.json
# Expect: warm.cache_read > 0 and significantly larger than warm.input

If cache_read is 0 on the warm call, the fix didn't land or there's a second upstream issue. Stop and report rather than retrying blindly.

Development setup

Prerequisites

  • Python 3.11+ (stdlib only — no pip install needed)
  • Node.js 18+ (for markdownlint-cli2 via npx)

Build / lint / test commands

# Python syntax check (CI job: python-syntax)
python3 -m py_compile tools/check_cache.py tools/check_docs_consistency.py

# Docs consistency guard (CI job: consistency)
python3 tools/check_docs_consistency.py

# Markdown lint (CI job: markdown-lint)
npx markdownlint-cli2 '**/*.md' '!**/node_modules/**'

# Python unit tests (CI job: python-tests)
python3 -m pytest tests/ -v

# Python type check (CI job: python-typecheck)
python3 -m mypy tools/ --strict

# Secrets scan (CI job: secrets-scan — requires gitleaks binary)
gitleaks dir . --no-banner

CI

GitHub Actions workflow at .github/workflows/ci.yml runs on every push and PR: gitleaks, markdownlint, python syntax, docs consistency, pytest, mypy, and link check.

When in doubt

Read the SKILL.md fully. They're written to be self-explanatory for agents. If a skill is ambiguous, treat that as a bug in the skill and ask the user how to proceed rather than guessing.

Repository README

Describes OnlyTerp/prompt-cache-skills as a whole, which may contain artifacts other than this one. Where this artifact had no useful description of its own, its summary was taken from here.


Most popular OSS agent harnesses (Cline, Roo Code, Continue, OpenCode, Aider) leave 30-90% off your API bill on the table because their prompt-caching code is subtly wrong, off-by-default, or just missing for some providers.

This repo is a set of drop-in skills that any AI coding agent (Claude Code, Codex, Cline, Cursor, Devin, Gemini CLI, OpenCode…) can read and apply on its own.

You don't read the diffs. You point your agent at this repo and say:

"Apply every skill in this repo that matches the harnesses I use."

The agent reads each SKILL.md, checks if it applies to your setup, lands the diff, and verifies the fix on the wire. You go from broken or partial caching to 80-99% cache hit rates without doing the research yourself.


What you actually save

One row per completed audit, so the coverage matches the scorecard:

HarnessFindingCost impact todayFix / status
Claude Desktop CodeDefault Desktop Code launches embedded Claude Code; clean Mac logs show non-zero cache read/create counters by defaultAlready gets Anthropic cache benefits; no prompt-caching fix neededNo skill; working baseline
Codex CLICorrect OpenAI cache design: stable thread_id cache keyAlready gets OpenAI cache benefitsNo skill; reference implementation
Aider--cache-prompts off by default; 5min TTL/keepalive overheadMany users get 0% cache reads unless they opt in; shorter cache windowSkills: default-on caching + 1h TTL
OpenCodeStrong Anthropic path, but proxy/Bedrock edge cases existSome OpenAI-compatible→Anthropic/Bedrock routes miss cacheSkills: proxy detection + Bedrock doc-block fix
Roo CodeAnthropic volatile-message bug; Bedrock custom ARN gapWastes breakpoints; custom ARNs can drop to 0% cache readsSkills: volatile-msg fix + Bedrock custom ARN fix
ClineAnthropic volatile-message bug; OpenAI lacks prompt_cache_keyWastes Anthropic breakpoint; OpenAI native can get 0% cache readsSkills: volatile-msg fix + OpenAI cache key + timestamp pin
ContinueCache opt-in default; Gemini explicit caching missing; volatile-message bugMany users get 0% cache reads; Gemini relies on implicit luckSkills: default-on + volatile-msg + Gemini explicit cache
Hermes / NousMulti-provider cache plumbing works; xAI wire showed cached tokensNo verified savings bug in this auditNo skill; working audit
Codex DesktopChatGPT Codex backend cache-scope headers observed/inferredNo verified savings bug in this auditNo skill; inferred working
Devin CLIRaw CLI model path is opaque Codeium/Devin protobufCache behavior not inspectable from public CLI captureNo skill; unverified managed backend
Windsurf / CascadeClosed desktop; model turn not captured from CLICache behavior unverifiedNo skill; needs desktop capture
AntigravityClosed desktop; no model turn capturedCache behavior unverifiedNo skill; needs desktop capture
Grok CLIDocumented CLI chat proxy returns non-zero prompt_tokens_details.cached_tokens with real CLI headersAlready gets xAI cache benefits through managed proxyNo skill; working managed proxy

13 skills total cover the verified patchable OSS bugs. See skills/README.md for the full index.


How to use it

Option A — point any AI coding agent at this repo

In your agent of choice (Claude Code, Codex, Cline, Cursor, Devin, etc.):

Read https://github.com/OnlyTerp/prompt-cache-skills

Apply every skill in skills/ that matches the harnesses I currently
use. For each one:
1. Confirm the target file exists in my project at the cited path.
2. Apply the diff.
3. Run the SKILL's Verify steps and confirm the assertion passes.
4. If verify fails, revert and tell me why.

That's it. The agent picks up the rest from each SKILL.md's machine-readable frontmatter and instructions.

Option B — install as a skill bundle in Claude Code / Devin / etc.

If you use one of the agents that supports a skills directory:

# Claude Code
git clone https://github.com/OnlyTerp/prompt-cache-skills ~/.claude/skills/prompt-cache-skills

# Devin
git clone https://github.com/OnlyTerp/prompt-cache-skills ~/.config/devin/skills/prompt-cache-skills

# OpenCode
git clone https://github.com/OnlyTerp/prompt-cache-skills ~/.config/opencode/skills/prompt-cache-skills

Then ask your agent:

Run the prompt-cache-skills bundle on this codebase.

Option C — read and apply by hand

Each skills/<name>/SKILL.md is a complete fix: target, symptom, diff, verification. Apply the relevant ones manually if you don't trust your agent to do it.


What's in here

prompt-cache-skills/
├── skills/                       ← the fixes (this is what your agent reads)
│   ├── cline-fix-volatile-msg/
│   ├── cline-openai-cache-key/
│   ├── cline-pin-timestamp/
│   ├── roo-fix-volatile-msg/
│   ├── roo-bedrock-custom-arn/
│   ├── continue-fix-volatile-msg/
│   ├── continue-enable-defaults/
│   ├── continue-gemini-explicit/
│   ├── opencode-detect-openai-compat/
│   ├── opencode-bedrock-doc-blocks/
│   ├── opencode-mistral-cache-key/
│   ├── aider-1h-ttl/
│   └── aider-cache-default-on/
├── audits/                       ← evidence: completed audits + queued stubs
│   ├── cline.md
│   ├── roo-code.md
│   ├── aider.md
│   ├── opencode.md
│   ├── continue.md
│   ├── codex-cli.md              ← (reference, already correct)
│   ├── claude-code.md
│   ├── hermes-nous.md
│   ├── codex-desktop.md
│   ├── devin-cli.md
│   ├── windsurf-cascade.md
│   ├── antigravity.md
│   ├── grok-cli.md
│   └── queued stubs: crush, goose, aichat, gptme, avante-nvim, kilo-code
├── docs/                         ← the underlying API mechanics
│   ├── concepts/                 ← per-provider caching reference
│   ├── gotchas.md                ← 16 numbered footguns
│   ├── verification.md           ← how to confirm caching on wire
│   └── scorecard.md              ← completed audits graded at a glance
├── tools/                        ← scripts to verify caching + doc consistency
│   ├── check_cache.py            ← fire request twice, dump cache_* fields
│   ├── check_docs_consistency.py ← assert counts/tables/links don't drift
│   ├── audit_harness.sh
│   └── replay_harness.md
└── AGENTS.md                     ← entry point for AI agents reading this repo

Why this exists

If your agent harness sends 30,000 tokens of system prompt + tools per turn, on Claude 4.7 Opus that's $0.15 per turn uncached vs $0.015 cached — a 10x difference. A 50-turn coding session costs $7.50 vs $0.75. You're paying 10x what you should be because the harness you use either:

  • doesn't set cache_control at all,
  • sets it on volatile content that thrashes the cache,
  • doesn't set prompt_cache_key for OpenAI,
  • has caching gated behind a config flag you never set, or
  • just doesn't implement it for one of your providers.

None of these are hard to fix. They're all 5-15 line diffs. The hard part is knowing which one applies to your harness and getting it right. This repo does that work for you.


The grade card

13 completed harness audits, dated 2026-05-27. The original 7 include the default Claude Desktop Code baseline, source-recon audits for Codex CLI, Aider, OpenCode, Roo Code, Cline, and Continue, plus extended source/wire/local-install audits for Hermes/Nous, Codex Desktop, Devin CLI, Windsurf/Cascade, Antigravity, and Grok CLI. Six more files in audits/ are queued stubs, not completed audits.

HarnessAnthropicOpenAIBedrockGeminiManaged/other
Claude Desktop Codeworking (default Desktop Code verified)n/an/an/an/a
Codex CLIn/aworkingn/an/an/a
Aiderworkingautomaticn/an/an/a
OpenCodeworkingworkingpartialn/an/a
Roo Codepartialworkingpartialn/an/a
Clinepartialbrokenunverifiedn/an/a
Continuepartialpartialpartialbrokenn/a
Hermes / Nousworkingworking (Responses)n/aunverifiedxAI working
Codex Desktopn/aworking*n/an/aChatGPT Codex backend inferred
Devin CLIn/an/an/an/aunverified (opaque protobuf)
Windsurf / Cascaden/an/an/an/aunverified (desktop not captured)
Antigravityn/an/an/aunverifiedunverified (desktop not captured)
Grok CLIn/an/an/an/aworking (xAI CLI proxy cached tokens)

* RE-backed or inferred from captured/companion wire shape where public source is unavailable; see the linked audit for caveats.

Full per-provider breakdown with file:line citations in docs/scorecard.md.


Headline findings

  1. The "last 2 user messages" pattern is a copy-paste bug that propagated Cline → Roo → Continue. All three burn a breakpoint on the volatile current turn. Same one-line fix in each.
  2. Cline OpenAI native is silently broken — no prompt_cache_key, no prefix-stability work. Users on Cline+OpenAI pay full price.
  3. Gemini explicit caching is universally unimplemented. Only implicit (best-effort, free) caching engages, even on long sessions with massive stable system prompts where explicit gives a guaranteed 75% discount.
  4. Codex CLI is the reference for OpenAI-side caching — thread_id as cache key, preserved across compaction and into sub-agents.
  5. OpenCode's system-prompt split is the best Anthropic pattern.
  6. Hermes / Nous has real multi-provider cache plumbing — source covers Anthropic/OpenRouter/Nous/Qwen and xAI wire capture showed prompt_cache_key, x-grok-conv-id, and non-zero cached tokens.
  7. Closed managed surfaces need transport-aware capture. Grok CLI now verifies through its versioned chat proxy with non-zero cached tokens; Devin remains protobuf, while Windsurf and Antigravity still need desktop-driven captures.

Trust but verify

Every skill ships with a Verify section that captures the wire and confirms the fix landed. Don't take our word for it — the tools/check_cache.py script fires any request body twice (cold + warm) and prints the diff of cache_* token fields.

Run it before and after applying a skill. You should see cache_read_input_tokens (Anthropic) or cached_tokens (OpenAI) or cachedContentTokenCount (Gemini) go from 0 to most of your input.


Contributing

We accept new skills, new harness audits, and corrections. See CONTRIBUTING.md. The bar is: a captured request body + a verified hit-rate change. We don't take vibe submissions.

By participating you agree to our Code of Conduct.

Security

Found something that looks like a credential leak path, a request construction bug that leaks user secrets, or any other security issue? See SECURITY.md for the disclosure process. Don't open a public issue.

Changelog

Releases tracked in CHANGELOG.md.

License

Skills and audit prose: CC-BY-4.0. Code (tools/): MIT.


Trustgrade B

  • passBody integrity

    Whether the stored document is plausibly the kind of file the artifact declares, rather than something fetched by mistake.

  • passType matchnot applicable to this artifact type

    Whether the artifact is really the kind of thing its metadata claims it is.

  • passFreshness

    How long since the source repository was last pushed to.

  • passPrompt injection

    Scans the artifact's own text for instructions aimed at your agent rather than at you.

  • warnLicenseno SPDX license detected

    Whether the source repository declares an SPDX license permissive enough to redistribute.

How the grade is calculated

Each check contributes 0 points when it passes, 1 when it warns, and 2 when it fails. The total maps to a letter:

  • Aevery check passed
  • Bone warning
  • Ctwo warnings
  • Dprompt injection or body integrity failed, or three warnings
  • Fone of those failed, and something else is wrong

These are automated hygiene checks, not a security audit, and not a dependency or vulnerability scan. A grade of A means nothing was flagged — not that the artifact is safe.

Versions

  • git-273167ced47e2026-08-04