Odel
rag mcp

rag mcp

Local
@jaimenbellPythonMITUpdated 2w ago

Minimal RAG-over-a-corpus MCP retrieval: search_knowledge returns cited chunks. Local embeddings.

rag-mcp

CI

A minimal, honest RAG-over-a-corpus MCP retrieval tool. One tool, search_knowledge(query, k), that embeds a query, vector-searches a local corpus, and returns passages with citations (source + heading + chunk index) so answers are traceable.

Built to slot into the mcp-factory manifest model. Fully local + $0 (no paid embedding API).

Why it's safe to put in front of a real corpus

  • Cited - every hit carries source + heading + chunk_index.
  • Auth-scoped - results are confined to the configured corpus root; sources that escape it (absolute paths, .. traversal) are refused.
  • Fail-soft - a down or empty store returns a structured error, never an exception that crashes the calling agent.
  • Bounded - k is clamped to [1, 20]; empty queries are rejected.
  • Version-pinned deps (requirements.txt).

Stack

LayerChoice
Embeddingslocal ONNX all-MiniLM-L6-v2 (384-dim, CPU, $0) -- default. bge-large-en-v1.5 (1024-dim, 512-token context) available opt-in via RAG_MCP_EMBEDDER=bge; see CUTOVER.md.
Vector storeChromaDB embedded PersistentClient (zero-infra)
Servermcp Python SDK 2.x, stdio transport, protocol revision 2026-07-28

Protocol revision

Pinned to mcp==2.0.0, the first SDK release implementing MCP protocol revision 2026-07-28. The server serves both eras on the same stdio connection -- the client's first frame picks:

Client opens withNegotiated revisionNotes
a per-request _meta envelope (or a server/discover probe)2026-07-28stateless per-request envelope; no initialize
the classic initialize handshake2025-11-25handshake era caps here -- expected, not a downgrade

2026-07-28 is not reachable via the initialize handshake; it is a "modern" revision reached through server/discover or an inline _meta version stamp. Era selection is automatic and per-connection -- there is no server-side flag.

tests/test_protocol_version.py asserts both paths end-to-end, so a dependency rollback that silently drops the server to an older revision fails CI instead of passing quietly.

Quick start

python -m venv .venv && .venv/Scripts/python -m pip install -r requirements.txt

# Ingest a corpus (markdown). Incremental by default: only files whose content
# changed since the last run are re-embedded.
python -m rag_mcp.cli ingest path/to/docs --db ./store.chroma

# Force a rebuild in place (ignore the manifest, re-embed everything)
python -m rag_mcp.cli ingest path/to/docs --db ./store.chroma --full

# One-off query (corpus root = the auth scope)
python -m rag_mcp.cli query "your question" --db ./store.chroma --corpus path/to/docs -k 5

# Run as an MCP server (stdio); configure via env first
#   RAG_MCP_CORPUS_ROOT, RAG_MCP_DB_PATH, RAG_MCP_COLLECTION, RAG_MCP_EMBEDDER
python run_server.py        # operational entrypoint (referenced by mcp.yaml)
python -m rag_mcp           # same server, via the packaged console entry point
rag-mcp                     # after `pip install jaimenbell-rag-mcp` -- console script

Keeping the index fresh (incremental ingest)

Ingest is incremental by default. A manifest inside the store dir records a SHA-256 of each file's decoded text; a run re-embeds only what actually changed, and prunes what upsert alone never could (chunks of deleted/renamed notes, and trailing chunks of notes that got shorter).

Measured on a live 2808-file / 26.6 MiB corpus (bge, CPU):

RunCost
tick with no changes~0.7s (walk + read + hash everything)
full re-embed~2h33m (50,109 chunks at ~5.5 chunks/sec)

That is what makes a frequent schedule affordable: reingest.bat is meant to run every 15 minutes instead of once daily at 03:00, which had left a note written at 03:05 invisible to search_knowledge for nearly 24 hours.

The manifest is only trusted when the run identity matches -- embedder, embedding dimension, collection and chunking parameters. Change any of them and every file is re-embedded, so an embedder swap can never be silently half-applied. A missing, corrupt, or mismatched manifest, or a manifest against an empty store, all degrade to a full rebuild; nothing degrades to a wrong skip.

Snapshot de-duplication

The manifest's skip is a whole-file hash, so it cannot see the duplication that actually hurts retrieval: a daily snapshot series (fleet-health-2026-07-23.md and friends) repeats yesterday's paragraphs verbatim inside a file whose hash still changed. Measured on the live vault, one ## RED Bots status line took five distinct values across fifteen consecutive files and crowded a top-10 with byte-identical copies of itself, burying the document that explained it at rank 16.

Ingest therefore also de-duplicates at chunk level, but only within a dated series and only against the immediately preceding snapshot. The first occurrence is always embedded and keeps its own date as its source; later verbatim repeats are not embedded, and instead extend the survivor's repeat_dates metadata, which search_knowledge returns as snapshot_date / also_unchanged_on / snapshots_covered. So "what did this say on date X" is still answerable -- that is why the series is de-duplicated rather than excluded. A value that changes and later returns is kept, because it is a new fact rather than a repeat.

Scope is narrow and stated with the rule in rag_mcp/snapshots.py: filename ending in -YYYY-MM-DD, at least 3 such files sharing a directory and stem, byte-identical under an identical heading. On the live corpus that is 316 of 2,814 files and collapses 842 of 50,428 chunks (17.5% of series chunks, 1.67% corpus-wide) while touching zero ordinary notes. Disable with --no-snapshot-dedupe.

--full rebuilds in place (ignores the manifest, keeps the store); --clean deletes the store first. Both still WRITE a manifest, so the next run is cheap. reingest-clean.bat (weekly) remains a belt-and-braces reset.

As an MCP server

Register via mcp.yaml (validated against mcp-factory's Manifest loader). The tool is search_knowledge(query, k); it reads the store configured by the RAG_MCP_* env vars.

Tests

python -m pytest        # 149 passed

Layout

rag_mcp/
  chunking.py   heading-scoped, overlapping markdown chunks
  store.py      VectorStore (Chroma) + Embedder protocol (MiniLM default + BgeEmbedder opt-in + offline HashEmbedder)
  ingest.py     idempotent ingest pipeline with source/heading/chunk-index metadata; incremental by default
  manifest.py   per-file content hashes -> skip unchanged files, prune stale chunks
  search.py     search_knowledge: cited, auth-scoped, fail-soft, bounded
  server.py     MCP stdio server exposing search_knowledge
  config.py     env-driven Config
  cli.py        ingest + query CLI
  __main__.py   console entrypoint (`python -m rag_mcp` / `rag-mcp` script); fails loud on missing config
run_server.py   operational MCP entrypoint (referenced by mcp.yaml)
mcp.yaml        manifest (mcp-factory model)

Commercial support

Maintained by Jaimen Bell. For production MCP integrations, custom servers, or agent-reliability work, see jaimenbell.dev.

Building your own MCP server? The MCP Starter Kit has templates, a build playbook, and packaging war-stories from shipping this one.

mcp-name: io.github.jaimenbell/rag-mcp