Prompt evals over MCP: run a prompt on your dataset, score each output 1-5 with an LLM judge.
Probing this server's capabilities…