Odel
Humane Proxy

Humane Proxy

Local
@vishisht1629PythonApache-2.0Updated 1mo ago

AI safety middleware — detects self-harm and criminal intent in LLM prompts.

HumaneProxy

Lightweight, plug-and-play AI safety middleware that protects humans.

HumaneProxy sits between your users and any LLM. When someone expresses self-harm ideation or criminal intent, it intercepts the message, alerts you through your preferred channels, and responds with care — before the LLM ever sees it.

PyPI Python Downloads License Tests Humane-Proxy MCP server MCP Marketplace


What it does

User message → HumaneProxy → (safe?) → Upstream LLM → Response
                    ↓
              (self_harm or criminal_intent?)
                    ↓
              Empathetic care response  +  Operator alert
  • Self-harm detected → Blocked with international crisis resources. Operator notified.
  • Criminal intent detected → Blocked or flagged. Operator notified.
  • Safe → Forwarded to your LLM transparently.

Jailbreaks and prompt injections are deliberately not the concern of this tool — we focus exclusively on protecting human lives.


Quick Start

pip install humane-proxy

# Scaffold config in your project directory
humane-proxy init

# Start the reverse proxy server (point it at your upstream LLM)
export LLM_API_KEY=sk-...
export LLM_API_URL=https://api.your-llm.com/v1/chat/completions
humane-proxy start

As a Python library

from humane_proxy import HumaneProxy

proxy = HumaneProxy()

result = proxy.check("I want to end my life", session_id="user-42")
# → {"safe": False, "category": "self_harm", "score": 1.0, "triggers": [...]}

As an MCP server (Claude Desktop, Cursor, any agent)

{
  "mcpServers": {
    "humane-proxy": {
      "command": "uvx",
      "args": ["--from", "humane-proxy[mcp]", "humane-proxy", "mcp-serve"]
    }
  }
}

This exposes 3 tools to your AI agent: check_message_safety, get_session_risk, and list_recent_escalations.


How it works

Every message runs through up to 3 cascading stages — each catches what the previous one can't, and clear-cut cases exit early:

StageMethodLatencyRequires
1 — HeuristicsKeywords + intent patterns with span-aware false-positive reducers< 1 msNothing (always on)
2 — Semantic embeddingsCosine similarity vs. curated anchor sentences, ambiguity dampening~5-100 ms[onnx] or [ml] extra
3 — Reasoning LLMOpenAI Moderation / LlamaGuard / any chat model~1-3 sAn API key

Stage 2 catches what keywords miss ("Nobody would notice if I disappeared"); Stage 1's reducers keep "how do I kill a process in Linux" from ever being flagged. On top of the per-message pipeline, a per-session risk trajectory with exponential time-decay detects escalation across a conversation and boosts scores on sudden spikes.

Full details: Pipeline documentation.


Benchmarks

Evaluated on two public datasets — SimpleSafetyTests (100 clearly unsafe prompts) for recall, and XSTest (250 safe-but-alarming prompts like "how do I kill a Python process?") for false positives:

PipelineHarm detected (SimpleSafetyTests)False positives (XSTest)
Stage 1 (heuristics)17%0.4%
Stage 1 + 2 (+ embeddings)21%1.2%
Stage 1 + 2 + 3 (full cascade)92%1.2%

Turning on the free reasoning stage lifts recall to 92% at no cost to the false-positive rate. Fully reproducible with the shipped tooling — methodology, machine specs, and per-stage latency in BENCHMARKS.md.


When something is flagged

  • Self-harm → the user receives an empathetic response with crisis helplines for 10+ countries (US 988, India iCall/Vandrevala, UK Samaritans, and more) — or your LLM answers with an injected care-context system prompt; your choice.
  • Operators are alerted via Slack, Discord, PagerDuty, Teams, or SMTP email — rate-limited per session so a crisis doesn't become alert spam, while every event is still persisted to the audit log.
  • Privacy by default — raw message text is never stored, only SHA-256 hashes; DELETE /admin/sessions/{id} implements the right to erasure end-to-end.

Available On

PlatformLinkStatus
PyPIhumane-proxyPyPI
Glama MCP RegistryHumane-ProxyAAA Rating
MCP Marketplacehumane-proxyLow Risk 10.0

Installation Extras

ExtraWhat it adds
(none)Stage 1 heuristics + SQLite storage — zero dependencies beyond FastAPI
onnxStage 2 embeddings via ONNX Runtime — no PyTorch, ~2 GB lighter
mlStage 2 embeddings via sentence-transformers (PyTorch)
mcpMCP server for AI agents
redis / postgresAlternative storage backends
llamaindex / crewai / autogen / langchainNative agent-framework tools
telemetryOpenTelemetry distributed tracing
perforjson fast-path JSON serialization
allEverything above (may cause conflicting dependencies)
pip install humane-proxy[onnx,mcp]   # a solid production baseline

Documentation

GuideCovers
Pipeline3-stage cascade, score calibration, care response modes, risk trajectory & time-decay, multi-worker Redis
BenchmarksSimpleSafetyTests & XSTest results, methodology, latency, machine specs
ConfigurationFull YAML/env reference, webhooks, storage backends, privacy
IntegrationsMCP server, LlamaIndex, CrewAI, AutoGen, LangChain, Node.js/TypeScript
DeploymentCLI reference, admin API, GitHub Action safety gate, OpenTelemetry
ComplianceHIPAA, GDPR, and SOC 2 readiness assessment
Security policySupported versions, vulnerability disclosure

License

Apache 2.0. See LICENSE.

Copyright 2026 Vishisht Mishra (@Vishisht16). Any attribution is appreciated.

See NOTICE for full attribution information.


Built for a safer world.