Odel
AILANG Parse

AILANG Parse

@sunholo-dataHTMLApache-2.0Updated 5 days ago

Deterministic DOCX/PPTX/XLSX/PDF parser: track changes, comments, headers, footers, merged cells.

Server endpointStreamable HTTPNo authProbed

This is the third-party server itself — Odel doesn't run it. Hitting this URL directly talks straight to the upstream server with no auth or proxying. Connect through Odel to front it with managed auth.

AILANG Parse

AILANG Registry PyPI npm Go MCP Registry CI

Universal document parsing and generation in AILANG. Extracts structured content from DOCX, PPTX, XLSX, PDF, and image files into JSON and markdown — and writes documents back out in 9 formats. To author a document, write Markdown and convert it; see Writing documents in Markdown.

Office formats (DOCX, PPTX, XLSX) use deterministic XML parsing — no AI, no cloud, instant results. PDFs default to the deterministic pdftotext backend (poppler) — also no AI, no cloud — with docling and liteparse as local alternatives and pluggable AI (Gemini, Claude, local Ollama) for scanned/image-only pages via --pdf-backend ai. Images delegate to whatever AI model you plug in. AILANG Parse is AI-agnostic: swap --pdf-backend/--ai to change the backend, zero code changes.

Install

Requires AILANG CLI.

# Clone and symlink
git clone https://github.com/sunholo-data/ailang-parse.git
ln -s "$(pwd)/ailang-parse/bin/docparse" /usr/local/bin/docparse

SDKs

Use AILANG Parse from your language of choice:

pip install ailang-parse          # Python
npm install @ailang/parse         # JavaScript/TypeScript
go get github.com/sunholo-data/ailang-parse-go  # Go

Quick Start

# Office documents (deterministic, no AI needed)
docparse report.docx
docparse slides.pptx
docparse spreadsheet.xlsx

# PDF (deterministic pdftotext by default — no AI); images (AI auto-enabled)
docparse document.pdf
docparse photo.png

# Options
docparse report.docx describe        # AI image descriptions
docparse report.docx summarize       # AI document summary
docparse contract.pdf                # PDF: deterministic pdftotext (default)
docparse scan.pdf --pdf-backend ai --ai gemini-2.5-flash  # Scanned PDF needs AI

# Format conversion
docparse report.docx --convert output.html
docparse data.csv --convert report.docx
docparse notes.md --convert slides.pptx
docparse notes.md --convert offer.docx --reference-doc letterhead.docx

# AI document generation
ailang run --entry main --caps IO,FS,Env,AI --ai gemini-2.5-flash \
  docparse/main.ail --generate report.docx --prompt "Q1 sales report with tables"

Output

Every run produces:

  • docparse/data/output.json — Structured JSON with typed blocks
  • docparse/data/output.md — LLM-ready markdown

What AILANG Parse Extracts

FeatureDOCXPPTXXLSXBest Competitor
Tables with merged cellsYesYesYesRaw OOXML only
Track changes (redlining)YesPandoc (3/3)
Comments (interleaved)YesRaw OOXML (2/2)
Headers/footersYesKreuzberg (2/3)
Text boxes / VML shapesYesYesRaw OOXML (1/2)
Equations (§22.1)YesNone
Field codes (§17.16)YesKreuzberg, OOXML
Speaker notesYesNone
Multi-sheet extractionYesKreuzberg

OfficeDocBench (69 files, 11 formats, 7 metrics): AILANG Parse 93.9% composite with 100% coverage vs nearest competitor 68.0% coverage-adjusted. 8 parsers compared including Raw OOXML, Pandoc, Kreuzberg, MarkItDown, Unstructured, Docling. Scores include aspirational ECMA-376 spec targets that intentionally lower our score.

Supported Formats

Parsing (16 formats): DOCX, PPTX, XLSX, ODT, ODP, ODS, HTML, Markdown, CSV, EPUB, EML, MBOX, TEX, RTF, PDF, images (JPG/PNG)

Generation (9 formats): DOCX, PPTX, XLSX, ODT, ODP, ODS, HTML, Markdown, QMD (Quarto)

Writing documents in Markdown

Markdown is the input an LLM can write, so it is the practical way to generate a document: write markdown, convert to any of the nine output formats.

docparse report.md --convert report.docx

What survives the trip: YAML front matter (title/author/date → document properties), bold/italic/code/strike as real character formatting, links as real hyperlinks, images (local paths are read and embedded), fenced code blocks, blockquotes, nested lists, thematic breaks, and tables with alignment and column spans.

Headers, footers, comments and tracked changes have no Markdown syntax; those are preserved when converting from a document that already contains them.

Styling a generated DOCX from a template

--reference-doc is the Quarto/Pandoc reference-doc feature: an existing .docx supplies the look, the Markdown supplies the content.

docparse annex.md --convert annex.docx --reference-doc letterhead.docx

The template's styles.xml, numbering.xml, theme, embedded fonts, headers, footers and page setup are applied to the new content. Everything the merge does not regenerate is carried through byte-for-byte, so the letterhead, logo and licensed fonts come out exactly as they went in.

What comes from where:

Templatepage size, margins, headers, footers, page numbering, fonts, theme, colours
Your documentthe body content, and docProps/core.xml (title/author)
Mergedstyles.xml (ours fill only the styleIds the template lacks), numbering.xml (our list definitions take ids above the template's), [Content_Types].xml, both .rels

Two consequences worth knowing:

  • The template's headers and footers win. A source document's own headers are dropped rather than mixed with the letterhead. The page furniture all lives in the template's body <w:sectPr>, which is lifted whole.
  • The template's comments are dropped along with its body, and so are commentsExtended.xml and people.xml. Comments in the source document still come through.

Two flags refine a multi-section template:

  • --reference-section N picks which of the template's sections supplies the page setup, headers and footers — 1 is the first section, Word's numbering. The default is the last section (the body-level one, what the flag-less behaviour has always lifted). A multi-section template's wanted furniture is often an earlier section's — the master agreement's CONFIDENTIAL footer, not the Annex's missing one.
  • --table-style NAME binds generated tables to a table style the template defines (matched on styleId, then style name). Without it, the style named Table is used if the template has one, else the first table style that is not the implicit Normal Table. Under a bound style the generator stops emitting its own hardcoded borders — the style carries them.

An unreadable or non-DOCX reference is an error and writes nothing — a silent fallback to the built-in styling would produce a plausible file missing exactly the letterhead it was asked for. DOCX output only.

Architecture

docparse/
├── types/document.ail           # Block ADT (11 variants)
├── services/
│   ├── format_router.ail        # Format detection (36 inline tests)
│   ├── zip_extract.ail          # ZIP layer (9 inline tests)
│   ├── docx_parser.ail          # DOCX XML → Blocks (6 inline tests)
│   ├── pptx_parser.ail          # PPTX slides → Blocks
│   ├── xlsx_parser.ail          # XLSX worksheets → Blocks
│   ├── direct_ai_parser.ail     # PDF/image → Blocks (AI)
│   ├── layout_ai.ail            # AI self-healing (optional)
│   ├── output_formatter.ail     # JSON + markdown output
│   └── docparse_browser.ail     # WASM browser adapter
└── main.ail                     # CLI entry point

91 contracts, 50+ inline tests. Of the 91, Z3 proves 14 outright; the rest are checked at runtime under --prove/--verify-contracts in CI, and skip statically because parser code is recursive and higher-order, which is outside Z3's decidable fragment.

AI Configuration

AILANG Parse uses AILANG's AI effect — any model AILANG supports works:

docparse scan.pdf --ai gemini-2.5-flash          # Google (default; fast)
docparse scan.pdf --ai gemini-3-flash-preview    # Google (slower; thinking model)
docparse scan.pdf --ai granite-docling           # Local Ollama (free)
docparse scan.pdf --ai claude-haiku-4-5          # Anthropic

AI usage is bounded by capability budgets (AI @limit=200 on main), so costs are predictable.

Dev Commands

docparse --check       # Type-check all modules
docparse --test        # Run inline tests
docparse --prove       # Static Z3 contract verification

Benchmarks

uv run benchmarks/run_benchmarks.py --suite office     # Structural (no API, instant)
uv run benchmarks/run_benchmarks.py --suite pdf         # PDF extraction (needs AI)
uv run benchmarks/run_benchmarks.py --competitors       # Compare to Docling etc.

See benchmarks/ for details.

License

Apache 2.0