claude-cost-audit-mcp
An MCP server for exact Claude API cost calculation — correct prompt-cache economics (write 1.25x at 5-minute
TTL / 2x at 1-hour TTL, read 0.1x), cache break-even analysis, and a scanner for the single most common
Claude-cost mistake: estimating tokens with OpenAI's tiktoken.
The mistake this catches
Per Anthropic's own documentation: "Do not use tiktoken. It's OpenAI's tokenizer. It undercounts Claude tokens by
~15-20% on typical text, and by much more on code or non-English input." Any cost estimate or context-budget
check built on tiktoken for a Claude model is silently wrong — this scans source for that exact pattern
(foreign tokenizer + Claude/Anthropic usage in the same file) and points to the fix
(messages.count_tokens).
Cache economics people get wrong
Caching doesn't pay off starting from the second request the way most people assume. The real break-even is
2 reuses at 5-minute TTL, but 3 reuses at 1-hour TTL — because the 1-hour write costs more (2x vs 1.25x).
check_cache_breakeven computes this from Anthropic's real multipliers instead of a guessed rule of thumb, and
flags prefixes below the model's own minimum-cacheable-token threshold (which is not monotonic across model
generations — 512 tokens on the newest models, 4096 on some older ones).
Tools
calculate_message_cost
Exact $ cost from token counts, using Anthropic's real current per-model pricing (including Claude Sonnet 5's time-limited introductory rate, resolved automatically from the date you pass).
check_cache_breakeven
Given expected request volume reusing the same prefix, tells you whether caching saves money and at which TTL.
audit_tokenizer_usage
Scans source for tiktoken/gpt-tokenizer/etc. alongside Claude API usage.
Use it
Hosted (recommended): MCPize — free tier, $7/mo Pro.
Self-host:
npm install
node server.js
Part of a small suite
mcp-schema-audit-mcp, cron-schedule-audit-mcp, regex-safety-audit-mcp.
License
MIT