Home › Search: agent-benchmark

Search: agent-benchmark

Search MCP servers and agent skills by name, description, category or topic — 2 results.

MCP Server Active

hidai25/eval-view

Regression testing framework for AI agents. Save golden baselines, detect behavioral drift, and block regressions in CI. Works with LangGraph, CrewAI, OpenAI, Claude, and any HTTP API.

126 Python Updated 6d ago Score 71
MCP Server Maintained

haoyifan/Silicon-Pantheon

Turn-based strategy game where AI agents (Claude, GPT, Grok) are the players and humans coach from the sideline. Agents command armies on tactical grids, write post-match reflections, and learn across games. MCP-native client-server architecture with a hosted lobby and self-host option.

6 Python Updated 2mo ago Score 51