Governed GPU inference ops (vLLM + Ray Serve): latency RCA, scaling, drain, 39 tools.
Runs locally over stdio
This server isn't hosted — your MCP client launches it from a package registry. Use one of the commands below, or drop the config into your client (e.g. Claude Desktop).
Run
uvx inference-aiops
MCP client config
{
"mcpServers": {
"inference-aiops": {
"command": "uvx",
"args": [
"inference-aiops"
]
}
}
}