Odel
token optimizer

token optimizer

Local
@omkar98541PythonApache-2.0Updated 1mo ago

Reversible context compression for AI agents: cut token usage up to 94%, originals retrievable

slimctx — the token optimizer for AI agents

CI PyPI MCP Registry License: Apache 2.0 Python 3.9+ Dependencies: zero

Zero-dependency, fully-reversible context compression for AI agents.

slimctx compresses what your agent reads — tool outputs, logs, JSON, source files, prose — before it reaches the LLM. Same answers, fraction of the tokens. Pure Python stdlib: no ML models, no downloads, no network calls, ever. Auditable end to end in ~1,600 lines.

slimctx demo: 61,700 tokens compressed to 298 in 29ms, FATAL lines preserved, byte-exact retrieval
Live output of python3 benchmarks/demo.py — run it yourself, nothing is staged.

from slimctx import Pipeline, Config

pipe = Pipeline(Config(target_tokens=32_000))
result = pipe.compress(messages)        # OpenAI/Anthropic-style dicts
print(result.savings_ratio)             # e.g. 0.82

original = pipe.retrieve("a1b2c3d4...")  # byte-exact original, any time

Results (synthetic workloads modeled on real agent traffic)

WorkloadBeforeAfterSavingsKey facts kept
Code search (100 results)5,55791684%
SRE incident debugging61,699298100%
GitHub issue triage12,83697592%
Codebase exploration5,7342,76052%

Every run also verifies that each planted "needle" (the FIXME, the OOMKill, the outlier) survives compression, and that every lossy transform is byte-exact reversible. Reproduce with python3 benchmarks/bench.py.

How it works

messages ──► ContentRouter ──► one of:
                ├─ JSON  : lossless tabularization (repeated keys → header,
                │          constant columns → legend), then relevance-ranked
                │          row selection only if still over budget
                ├─ LOG   : Drain-style template mining — repeated lines
                │          collapse to `pattern [x1432]`; errors verbatim
                ├─ CODE  : AST skeleton — signatures + docstrings kept,
                │          bodies elided EXCEPT those relevant to the query
                └─ TEXT  : extractive sentence selection (BM25 + salience
                           + position), verbatim, never paraphrased

The four guarantees

  1. Universal reversibility. Before any lossy transform, the original goes into a content-addressed store (memory / SQLite / bring-your-own cipher) and the output carries a [slimctx-ref <hash> ...] marker. The model — or you — can always get the byte-exact original back.
  2. Errors are never dropped. Every compressor pins error/warning content: log errors pass verbatim, salient JSON rows are kept, salient sentences outrank filler.
  3. Deterministic output. Same input → byte-identical output, across runs and processes. Compressed prefixes stay stable, so provider prompt-caches (Anthropic/OpenAI) keep hitting.
  4. Net gain or no-op. If a transform doesn't save enough tokens to pay for its marker, the original is kept untouched. The live zone (system prompt + last N messages) is never modified at all.

Why not just use Headroom?

Headroom is the established project in this space and is more featureful today (provider proxy with SSE streaming, agent wrappers, cross-agent memory, an ML compression model). slimctx makes a different set of trade-offs, aimed at locked-down / client-site deployments:

Headroomslimctx
ReversibilityJSON only (CCR); dropped text is goneevery lossy transform
Log handlinggeneric text scoringtemplate mining ([x1432] collapse)
Code handlingAST skeletonAST skeleton + query-relevant bodies kept
DependenciesRust core, ONNX runtime, 261MB HF modelstdlib only
Network egressHuggingFace pull on first runnone, ever
Store encryptionnone (plaintext SQLite)cipher hook (bring your own)
Determinismcache-aligner componentby construction (pure functions + memo)
Audit surface~10s of KLOC across 3 languages~1,200 lines of Python

If you need the proxy/wrap ecosystem, use Headroom. If you need something you can read in an afternoon, run air-gapped, and certify for a client environment, use slimctx.

Install / test

pip install slimctx           # from PyPI — or vendor the slimctx/ directory
python -m pytest tests/ -q    # 26 tests: invariants, not examples
python3 benchmarks/bench.py   # reproduce the numbers above

Integration sketches

As a library (any framework): call pipe.compress(messages) right before your provider SDK call; expose pipe.retrieve as a tool named retrieve so the model can pull originals.

As an MCP server (GitHub Copilot, Claude Code, Cursor, ...): ships built in, stdlib-only:

python3 -m slimctx.mcp_server --db ~/.slimctx/store.db

See USAGE.md for the GitHub Copilot (.vscode/mcp.json) setup and a security deployment checklist.

Encrypted store:

from cryptography.fernet import Fernet          # optional, your choice
f = Fernet(key)
store = SqliteStore("ccr.db", cipher=(f.encrypt, f.decrypt))
pipe = Pipeline(store=store)

License

Apache-2.0. Original implementation — no code derived from Headroom.