CacheCanary

Find out when your Claude prompt cache stops working on Bedrock.

Bedrock can reuse the start of your prompt, like your instructions and tool list, and charge about 10% of the normal price for it. When that quietly stops, every request still works and you pay full price again. Usually the bill is the first sign. CacheCanary shows how often your requests use the cache, tells you why one didn't, and catches the problem in CI before it ships.

pip install cachecanary

Free and open source (Apache 2.0). The tool runs on your machine and sends nothing anywhere.

What a broken cache costs

90% offis what cached prompt text costs on Bedrock. When the cache breaks, that text goes back to full price.
$2,700 a monthextra for one agent with a 10,000-token prompt and 100,000 requests a month.
2-3xdaily spend for LiteLLM users for six days in July 2026, after one library update.

The example uses Anthropic's list price for Claude Sonnet 4.6 ($3 per million input tokens). Bedrock prices can vary by region.

How prompt caching works, in 30 seconds

What CacheCanary does

1. Shows whether you're losing money right now

Turn on Bedrock's invocation logging (to S3 or CloudWatch), then point CacheCanary at the log files. It tells you, per model, how often requests were served from the cache.

$ cachecanary logs bedrock-logs/*.json.gz --min-hit 0.8
us.anthropic.claude-sonnet-4-6: hit 50% over 8 calls (read 11496, write 11496, uncached 80, unparsed 0)  <-- below threshold

Here only half the calls used the cache, below the 80% target, so the command fails. Run it as a scheduled CI job and that failure becomes your alert.

2. Tells you why a request missed the cache

Save a request your app sends to Bedrock as a JSON file. With boto3's Converse API that's one line:

# kwargs: what you pass to client.converse(**kwargs)
json.dump(kwargs, open("request.json", "w"))

Then check it, or compare it with an earlier one:

$ cachecanary lint request.json
[warn] dynamic-in-prefix: Found ISO date text inside the cached prefix. If it changes per request, every call misses. (system[0])

$ cachecanary diff yesterday.json today.json
system-changed: The system prompt changed inside the cached prefix (look for dates, IDs or per-user text). (system[0])

$ cachecanary lint request.json --model us.anthropic.claude-haiku-4-5-20251001-v1:0
[error] prefix-too-short: Prefix up to this checkpoint is ~1892 tokens; claude-haiku-4-5 needs at least 4096. The request succeeds but nothing is cached. (system[0])

In plain words: the first request has a date in its system prompt, so it changes every day. The second shows the system prompt changed between yesterday and today. The third prompt is too short for Haiku 4.5, so it's never cached, even though the same prompt caches fine on Sonnet 4.6. system[0] points to the first block of the system prompt.

3. Catches it before it ships

Add it to GitHub Actions. Problems show up as comments on the pull request, next to the file that caused them.

- uses: Haarris/cachecanary@v0
  with:
    command: lint
    args: tests/fixtures/agent_request.json --model us.anthropic.claude-sonnet-4-6

If your CI has AWS access, probe goes further: it sends the request to Bedrock twice and fails if the second one wasn't served from the cache.

Why caching breaks

Often nobody changed the prompt on purpose. These are the usual causes:

Anthropic added cache diagnostics to its own API, so you can see why a request missed. They don't cover Bedrock. That's the gap CacheCanary fills.

Built on how Bedrock actually behaves

The checks follow AWS's documented caching rules. The core behaviour was verified against live Amazon Bedrock, including streaming responses, both cache lifetimes (5 minutes and 1 hour), and each model's minimum prompt size. The log reader is tested against real Bedrock invocation logs.

17Claude models with their cache rules built in, including Sonnet 4.6, Haiku 4.5 and Opus 5.5
Both APIsConverse and InvokeModel, streaming or not
No AWS neededto check a saved request. lint and diff run offline in about 30 ms.

Coming next

A hosted dashboard: cache hit rate and wasted spend per app, an alert when the rate drops, and the reason for each miss. It runs inside your own AWS account, so prompts never leave it.

If you'd use it, email hello@cachecanary.com. I read every message.