Find out when your Claude prompt cache stops working on Bedrock.
Bedrock can reuse the start of your prompt, like your instructions and tool list, and charge about 10% of the normal price for it. When that quietly stops, every request still works and you pay full price again. Usually the bill is the first sign. CacheCanary shows how often your requests use the cache, tells you why one didn't, and catches the problem in CI before it ships.
Free and open source (Apache 2.0). The tool runs on your machine and sends nothing anywhere.
What a broken cache costs
The example uses Anthropic's list price for Claude Sonnet 4.6 ($3 per million input tokens). Bedrock prices can vary by region.
How prompt caching works, in 30 seconds
- You add a cache point to your request, usually right after your instructions and tools. Everything before it is the part Bedrock can reuse.
- That part has to be exactly the same on every request. Change one character, like today's date, and Bedrock can't reuse it.
- It also has to be long enough: at least 1,024 tokens on Sonnet 4.6 and 4,096 on Haiku 4.5. Shorter prompts are never cached, and nothing tells you.
- A cached prompt lasts 5 minutes (or 1 hour if you ask for it), and every request that uses it starts the clock again.
What CacheCanary does
1. Shows whether you're losing money right now
Turn on Bedrock's invocation logging (to S3 or CloudWatch), then point CacheCanary at the log files. It tells you, per model, how often requests were served from the cache.
$ cachecanary logs bedrock-logs/*.json.gz --min-hit 0.8 us.anthropic.claude-sonnet-4-6: hit 50% over 8 calls (read 11496, write 11496, uncached 80, unparsed 0) <-- below threshold
Here only half the calls used the cache, below the 80% target, so the command fails. Run it as a scheduled CI job and that failure becomes your alert.
2. Tells you why a request missed the cache
Save a request your app sends to Bedrock as a JSON file. With boto3's Converse API that's one line:
# kwargs: what you pass to client.converse(**kwargs)
json.dump(kwargs, open("request.json", "w"))
Then check it, or compare it with an earlier one:
$ cachecanary lint request.json [warn] dynamic-in-prefix: Found ISO date text inside the cached prefix. If it changes per request, every call misses. (system[0]) $ cachecanary diff yesterday.json today.json system-changed: The system prompt changed inside the cached prefix (look for dates, IDs or per-user text). (system[0]) $ cachecanary lint request.json --model us.anthropic.claude-haiku-4-5-20251001-v1:0 [error] prefix-too-short: Prefix up to this checkpoint is ~1892 tokens; claude-haiku-4-5 needs at least 4096. The request succeeds but nothing is cached. (system[0])
In plain words: the first request has a date in its system prompt, so it changes every day. The second shows the system prompt changed between yesterday and today. The third prompt is too short for Haiku 4.5, so it's never cached, even though the same prompt caches fine on Sonnet 4.6. system[0] points to the first block of the system prompt.
3. Catches it before it ships
Add it to GitHub Actions. Problems show up as comments on the pull request, next to the file that caused them.
- uses: Haarris/cachecanary@v0
with:
command: lint
args: tests/fixtures/agent_request.json --model us.anthropic.claude-sonnet-4-6
If your CI has AWS access, probe goes further: it sends the request to Bedrock twice and fails if the second one wasn't served from the cache.
Why caching breaks
Often nobody changed the prompt on purpose. These are the usual causes:
- A library upgrade moved the system prompt or dropped the cache point. This happened with LiteLLM in July 2026: cache hits fell from about 90% to 25-45% and spend went up 2-3x for six days.
- You switched to a new model ID or an inference profile (an AWS name that points to a model), and your library doesn't recognise it, so it stops asking for caching.
- A small edit put today's date, a request ID or the user's name into the system prompt.
- Your tools are listed in a different order on each request.
- An agent ran a dozen tools at once. Bedrock only searches back about 20 pieces of the conversation (messages, tool calls and tool results) for the cached part, so it can lose track of it.
Anthropic added cache diagnostics to its own API, so you can see why a request missed. They don't cover Bedrock. That's the gap CacheCanary fills.
Built on how Bedrock actually behaves
The checks follow AWS's documented caching rules. The core behaviour was verified against live Amazon Bedrock, including streaming responses, both cache lifetimes (5 minutes and 1 hour), and each model's minimum prompt size. The log reader is tested against real Bedrock invocation logs.
Coming next
A hosted dashboard: cache hit rate and wasted spend per app, an alert when the rate drops, and the reason for each miss. It runs inside your own AWS account, so prompts never leave it.
If you'd use it, email hello@cachecanary.com. I read every message.