Benchmarks local LLM inference speed (tokens/sec) on your own hardware via MCP tools.
Runs locally over stdio
This server isn't hosted — your MCP client launches it from a package registry. Use one of the commands below, or drop the config into your client (e.g. Claude Desktop).
Run
uvx inferbench-cli
MCP client config
{
"mcpServers": {
"inferbench": {
"command": "uvx",
"args": [
"inferbench-cli"
]
}
}
}