Odel
Claude Cost Audit

Claude Cost Audit

@tylerscomic-labJavaScriptMITUpdated 1mo ago

Exact Claude API cost calc with real cache economics, plus a tiktoken-misuse scanner.

View on GitHub
Server endpointStreamable HTTPOAuthProbed

This is the third-party server itself — Odel doesn't run it. Hitting this URL directly talks straight to the upstream server with no auth or proxying. Connect through Odel to front it with managed auth.

claude-cost-audit-mcp

License: MIT Live on MCPize

An MCP server for exact Claude API cost calculation — correct prompt-cache economics (write 1.25x at 5-minute TTL / 2x at 1-hour TTL, read 0.1x), cache break-even analysis, and a scanner for the single most common Claude-cost mistake: estimating tokens with OpenAI's tiktoken.

The mistake this catches

Per Anthropic's own documentation: "Do not use tiktoken. It's OpenAI's tokenizer. It undercounts Claude tokens by ~15-20% on typical text, and by much more on code or non-English input." Any cost estimate or context-budget check built on tiktoken for a Claude model is silently wrong — this scans source for that exact pattern (foreign tokenizer + Claude/Anthropic usage in the same file) and points to the fix (messages.count_tokens).

Cache economics people get wrong

Caching doesn't pay off starting from the second request the way most people assume. The real break-even is 2 reuses at 5-minute TTL, but 3 reuses at 1-hour TTL — because the 1-hour write costs more (2x vs 1.25x). check_cache_breakeven computes this from Anthropic's real multipliers instead of a guessed rule of thumb, and flags prefixes below the model's own minimum-cacheable-token threshold (which is not monotonic across model generations — 512 tokens on the newest models, 4096 on some older ones).

Tools

calculate_message_cost

Exact $ cost from token counts, using Anthropic's real current per-model pricing (including Claude Sonnet 5's time-limited introductory rate, resolved automatically from the date you pass).

check_cache_breakeven

Given expected request volume reusing the same prefix, tells you whether caching saves money and at which TTL.

audit_tokenizer_usage

Scans source for tiktoken/gpt-tokenizer/etc. alongside Claude API usage.

Use it

Hosted (recommended): MCPize — free tier, $7/mo Pro.

Self-host:

npm install
node server.js

Part of a small suite

mcp-schema-audit-mcp, cron-schedule-audit-mcp, regex-safety-audit-mcp.

License

MIT