Odel
web2md

web2md

Local
@io-oi-aiTypeScriptMITUpdated 3 days ago

MCP Server for Web2MD — convert URLs to Markdown from Claude Desktop, Cursor, etc.

web2md-core

Turn messy HTML into clean, LLM-ready Markdown.

This is the extraction and conversion engine behind Web2MD. It is a pure library — give it an HTML string, get Markdown back. No network calls, no API key, no account. Nothing in this package talks to a server.

npm install web2md-core

Why convert at all

Feeding raw HTML to a language model wastes most of your context window on markup, navigation, and ads. Converting first cuts that down and gives the model a document it can actually follow.

The library reports both numbers so you can see the difference:

import { convertToMarkdown } from 'web2md-core'

const result = convertToMarkdown(html, { url: 'https://example.com/post' })

console.log(result.markdown)
console.log(result.metadata.originalTokenCount, '→', result.metadata.tokenCount)
// e.g. 417 → 281

convertToMarkdown returns null when it cannot find a main content block — check for that rather than assuming a result.

What it does

  • Finds the actual article. Strips navigation, sidebars, ads, cookie banners, and footers, keeping the content a reader came for.
  • Preserves structure. Headings, lists, tables, and fenced code blocks survive the round trip — that structure is what lets a model answer questions about one specific section.
  • Reports tokens. Estimated counts for both the original HTML and the cleaned Markdown, plus helpers to split or trim for a target context window.
  • Runs anywhere. Uses linkedom for parsing, so it works in Node without a browser.

API

convertToMarkdown(html, options?)

The main entry point. Note the signature takes two arguments — the URL goes inside options, not as a positional parameter:

convertToMarkdown(html, { url: 'https://example.com/post' })
OptionDefaultMeaning
urlSource URL. Used to resolve relative links and fill metadata.url.
includeLinksfalseKeep <a> as Markdown links. Off by default because link URLs are often the bulk of the tokens on navigation-heavy pages.
includeImagesfalseKeep images. Off by default for the same reason.
includeMetafalsePrepend a metadata block (title, source, timestamp).
customRuleA CustomRule for site-specific extraction.
detectCodeLanguagefalseTry to infer the language of fenced code blocks.

includeLinks and includeImages default to off. That is deliberate — the primary use case is feeding an LLM, where both are usually noise. Turn them on when you are archiving rather than summarising.

Other exports

quickConvert(html, url?)        // same result, but with links, images and
                                // metadata turned ON — the "archive it" preset
extractContent(html, url?)      // main content element, before conversion
htmlToMarkdown(html, options?)  // low-level conversion, no extraction
countTokens(text)               // token estimate
splitByTokens(md, limit)        // chunk for RAG ingestion
optimizeForContextWindow(md, model)
htmlLooksLikeLoginWall(html)    // detect login walls so you can fail loudly
MODEL_CONTEXT_LIMITS            // context sizes for common models

Markdown → sanitized HTML (via DOMPurify), for previewing output:

renderMarkdownSync(md)
renderMarkdownFull(md)          // async; includes syntax highlighting
renderMarkdownWithFormulas(md)  // KaTeX math

getPageHTML() and getSelectionHTML() read document directly and therefore only work in a browser. They throw in Node — that boundary is intentional.

Site-specific extraction

Generic extraction handles most pages. When a site needs special treatment, pass a rule:

convertToMarkdown(html, {
  url: 'https://example.com/thread',
  customRule: {
    name: 'Example forum',
    domain: 'example.com',
    contentSelector: '.thread-body',
    removeSelectors: ['.signature', '.ad-slot'],
  },
})

Scope

This package covers extraction and conversion. It does not include Web2MD's browser extension, hosted API, or account system — those stay in the product.

Contributions to extraction quality are especially welcome: if a site converts badly, an issue with the URL and what went wrong is genuinely useful.

License

MIT