Client libraries
batchwatch only works if people instrument their pipelines, and nobody writes raw HTTP calls to donate data. These are the libraries that make it two lines.
| Package | Language | Tests | State |
|---|---|---|---|
python/ | Python 3.8+ | 96, all passing | works; PyPI release on the way |
typescript/ | TypeScript / JS, Node 20+ | 84, all passing (built + tested in CI) | works; built and tested in CI; npm release on the way |
dotnet/ | C#, net8.0 | built and tested in CI (xunit) | works; NuGet release on the way |
go/ | Go 1.21+ | 83, all passing (incl. -race) | works; module-proxy release on the way |
ruby/ | Ruby 3.0+ | 87, all passing (minitest) | works; RubyGems release on the way |
php/ | PHP 8.2+ | all passing | works; on Packagist, tagged release on the way |
java/ | Java 17+ | 88, all passing | works; Maven Central release on the way |
rust/ | Rust 1.63+ | 94, all passing (cargo test) | works, std-only; http-only (see note); crates.io release on the way |
cpp/ | C++17 (POSIX) | 85, all passing | works, stdlib+sockets only; source-only by design; http-only (see note) |
All nine expose the same surface, and a conformance check fails CI if any language falls behind — see The same surface, in all nine.
client/batchwatch.py in the repo root is the original single-file client
and is left untouched. clients/python/ is the packaged version of it, plus
spooling.
A note on TLS (Rust and C++). Every client except Rust and C++ gets
TLS from its standard library and talks to https://batchwatch.dev directly.
Rust's and C++'s standard libraries have no TLS, and both are written with
zero external dependencies on purpose, so their transport is http://
only. Pointed at the default https:// URL they spool rather than
deliver (the measurement is kept, not lost) until you point them at an
http:// endpoint or a local TLS-terminating proxy. Their READMEs say so.
What every client does the same way
It fails open. A batchwatch outage must never stop a user's job. Every
submission happens off the caller's thread with a short timeout, every error
is swallowed and logged at debug level, and the only call you await —
should_batch() — returns your default when it cannot answer, never a
guess. The default is "run it synchronously": being wrong that way costs
money, being wrong the other way blows a deadline. Each package has a test
that runs against a dead port and a hung socket.
Two lines to adopt. should_batch() before you submit, track() around
the call.
It never sends content. No prompts, no completions, no file names. The
body is built from one allowlist — provider, model, mode, endpoint, request
count, token counts, timestamps, status — and everything else is dropped by a
single function - _scrub in Python, Scrub in Go, clean in TypeScript,
sanitize in PHP, Ruby and Java, strip in Rust and C++ - on the way out.
Each package has
a test that asserts this on what the server actually received, with a
positive control so it cannot pass by sending nothing at all.
output_tokens defaults to null, never 0. Output costs five to six times
as much as input, so a saving computed on zero output is systematically too
low — 3.4x too low in the case that led to this rule — and nothing in the
response reveals it. Absence must stay absence all the way to the server.
Explicitly passing 0 still sends 0: zero is a measurement.
It spools to disk. An undeliverable completed measurement is appended
to a JSONL file and replayed later via POST /v1/calls/complete. Losing
measurements when the network is bad means losing them exactly when they are
most interesting.
It does the annoying parts. Beyond the two-line advisory path, every client carries the same high-level surface so you never hand-roll it:
- The high-level batch job —
batch(...)hands the client the two callables (batch-create + a synchronous fallback) and it owns the rest: a deadline guard that shifts to the fallback when the wait runs long, a poll loop with exponential backoff, jitter, a rate-limit floor and a first cadence informed by the model's measured p50, and partial-completion handling that splits a batch into landed / failed / expired mapped bycustom_id(never by index) with an idempotent retry of only the failed subset. We take the callable, never the payload — the deadline fallback is reported down the same accuracy path a completion uses, so nothing new is sent. - Read your own contributions —
my_calls()andkey_status()(GET /v1/calls/mine,/v1/keys/current): the per-key readback for verifying a measurement landed and checking your tier/quota. - Subscribe to outage alerts —
subscribe()/subscriptions()/unsubscribe()against/v1/subscriptions, the "own the outage moment" channel.
Unlike the measurement path, these last two do not fail open: they are explicit actions against a per-key route, so without a key they raise rather than silently pretend. The job path is the user's own job, so a misuse (a result before a submit, a deadline with no fallback) raises loudly too — only telemetry fails open.
No dependencies. Standard library only, in all nine.
The same surface, in all nine
Every client promises the same capabilities, and — since a feature can land in
one language, its card be closed in good faith, and the other eight silently lag
— conformance/ is the check that stops that. manifest.json
declares the promised surface; check.py greps each SDK's own source and its
own tests (per language, never a loose cross-language match) and fails CI on
any unexplained gap. Exemptions must be explicit and justified in the manifest.
Run it with python clients/conformance/check.py --list.
The spool format
One JSON object per line, in the shape /v1/calls/complete accepts:
{"provider":"openai","model":"gpt-5.6-sol","mode":"batch","requests":1,
"endpoint":null,"input_tokens":9720,"output_tokens":null,"status":"completed",
"started_at":"2026-08-25T10:00:00Z","ended_at":"2026-08-25T10:04:00Z"}
Identical across all nine clients, so a file written by one can be flushed
by another. Default location is $BATCHWATCH_SPOOL, otherwise
batchwatch-spool.jsonl in the temp directory.
Two consequences worth knowing before you rely on it:
- Spooling needs an API key.
/v1/calls/completetakes the caller's own timestamps, so it is closed to anonymous callers — see the reasoning insrc/index.js. A client without a token therefore does not spool at all: a file that can never be sent is a disk leak, not data safety. - Replay can duplicate. If the original
PATCHreached the server but the response did not, the spooled copy arrives as a second row. That is the deliberate trade: a duplicate is visible in the dataset, a lost measurement is not.
The file is capped (5 MB by default). Past the cap, measurements are dropped rather than filling the user's disk.
Publishing
Every client works today: install it from the repo (each README shows how) and it runs. Registry publishing is the next step, and it is in flight — the PHP package is already on Packagist, the .NET and TypeScript clients build and test in CI on every push, and the rest are being wired up (PyPI, npm, a Go module proxy, RubyGems, Maven Central, crates.io). Client CI itself lands in #184.
Two properties are deliberate design decisions, not gaps, and each has a workaround in the relevant README:
- Rust and C++ speak
http://only — their standard libraries carry no TLS and both are zero-dependency by design. Point them at a local TLS-terminating proxy in front ofbatchwatch.dev, and they deliver directly (see the TLS note above). - The POSIX-socket clients (C++) are first-class on Linux and other POSIX platforms; Windows needs a Winsock shim.