Home › Search: evaluation

Search: evaluation

Search MCP servers and agent skills by name, description, category or topic — 26 results.

Agent Skill Active

muratcankoylan/evaluation

Build evaluation frameworks for agent systems

17.5k Python Updated 18d ago Score 85
Agent Skill Active

huggingface/hugging-face-evaluation

Model evaluation with vLLM/lighteval and eval tables

10.9k Python Updated 2d ago Score 85
Agent Skill Active

NVIDIA/Model-Optimizer/evaluation

Evaluates accuracy of quantized or unquantized LLMs using NeMo Evaluator Launcher (NEL).

2.8k Python Updated today Score 84
Agent Skill Active

santifer/career-ops

14-skill collection for AI-powered job search: JD evaluation with A-F scoring, ATS-optimized PDF generation, portal scanners (Greenhouse/Ashby/Lever), interview prep with STAR+R, batch processing, and a Go dashboard TUI

62.5k JavaScript Updated today Score 90
MCP Server Official Active

DataEval/dingo

MCP server for the Dingo: a comprehensive data quality evaluation tool. Server Enables interaction with Dingo's rule-based and LLM-based evaluation capabilities and rules&prompts listing.

731 Python Updated yesterday Score 89
Agent Skill Active

NVIDIA/Model-Optimizer/accessing-mlflow

Query and browse evaluation results stored in MLflow.

2.8k Python Updated today Score 84
Agent Skill Active

NVIDIA/Model-Optimizer/launching-evals

Run, monitor, analyze, and debug LLM evaluations via nemo-evaluator-launcher.

2.8k Python Updated today Score 84
Agent Skill Active

NVIDIA/Model-Optimizer/monitor

Monitor submitted jobs (PTQ, evaluation, deployment) on SLURM clusters.

2.8k Python Updated today Score 84
Agent Skill Active

NVIDIA/NeMo-Evaluator/byob

Create custom LLM evaluation benchmarks using the BYOB decorator framework.

2.8k Python Updated today Score 84
Agent Skill Active

NVIDIA/NeMo-Evaluator-Launcher/accessing-mlflow

Query and browse evaluation results stored in MLflow.

2.8k Python Updated today Score 84
Agent Skill Active

NVIDIA/NeMo-Evaluator-Launcher/launching-evals

Run, monitor, analyze, and debug LLM evaluations via nemo-evaluator-launcher.

2.8k Python Updated today Score 84
MCP Server Active

hidai25/eval-view

Regression testing framework for AI agents. Save golden baselines, detect behavioral drift, and block regressions in CI. Works with LangGraph, CrewAI, OpenAI, Claude, and any HTTP API.

126 Python Updated 6d ago Score 71
Agent Skill Active

smixs/creative-director-skill

AI creative director with recursive self-assessment: 20+ methodologies (SIT, TRIZ, Bisociation, SCAMPER, Synectics), 3-axis evaluation calibrated against Cannes/D&AD/HumanKind, 5-phase process from brief to presentation

125 Python Updated 9d ago Score 71
MCP Server Active

sidclawhq/platform

Governance proxy for MCP servers. Wraps any upstream server with policy evaluation, human approval workflows, and hash-chain audit trails. 18+ framework integrations. Apache 2.0 SDK.

12 TypeScript Updated 3d ago Score 61
MCP Server Stale

jasonjmcghee/claude-debugs-for-you

An MCP Server and VS Code Extension which enables (language agnostic) automatic debugging via breakpoints and expression evaluation.

514 TypeScript Updated 7mo ago Score 60
MCP Server Active

iris-eval/mcp-server

MCP-native agent evaluation and observability server with trace logging, output quality evaluation, cost tracking, 12 built-in eval rules, real-time dashboard, and PII detection.

7 TypeScript Updated 4d ago Score 59
MCP Server Active

OrygnsCode/opa-mcp-server

Open Policy Agent (OPA) and Rego policy toolkit. 32 tools spanning authoring (format, lint, check, deps), evaluation (eval, test, bench, coverage), and OPA REST control (policies, data, decisions, compile). Wraps the OPA CLI and the [Regal](https://github.com/StyraInc/regal) linter, with AI-assisted helpers for explaining decisions, generating test skeletons, and suggesting fixes.

7 TypeScript Updated today Score 59
MCP Server Official Stale

heurist-network/heurist-mesh-mcp-server

Access specialized web3 AI agents for blockchain analysis, smart contract security auditing, token metrics evaluation, and on-chain interactions through the Heurist Mesh network. Provides comprehensive tools for DeFi analysis, NFT valuation, and transaction monitoring across multiple blockchains

66 Python Updated 4mo ago Score 56
MCP Server Active

verifyax/verifyax-mcp

MCP server for the VerifyAX platform. Enables agent evaluation, simulation testing, and functional/non-functional verification workflows through natural language.

1 TypeScript Updated 4d ago Score 53
MCP Server Active

shuji-bonji/xcomet-mcp-server

Translation quality evaluation using xCOMET models. Provides quality scoring (0-1), error detection with severity levels (minor/major/critical), and optimized batch processing with 25x speedup.

2 TypeScript Updated 5d ago Score 50
MCP Server Active

ankitkapur1992-hlido/hlido-mcp

Independent trust scores, claim audits, and comparisons for AI agents — queryable by your agent over MCP. Hosted Cloudflare Worker at hlido.eu/mcp (no install). Returns a 0–100 score, tier verdict, per-claim PASS/FAIL audit, and signed evidence for a reviewed agent. From [Hlido](https://hlido.eu).

0 JavaScript Updated 17d ago Score 50
MCP Server Maintained

mumez/pharo-smalltalk-interop-mcp-server

Pharo Smalltalk integration enabling code evaluation, class/method introspection, package management, test execution, and project installation for interactive development with Pharo images.

12 Python Updated 1mo ago Score 49
MCP Server Maintained

supertrained/rhumb

Agent-native tool intelligence across 1,000+ scored services. 21 MCP tools: discover services, check AN Scores, compare alternatives, resolve capabilities to ranked providers, execute through 3 credential modes (managed, BYOK, agent vault), track costs with receipts, and inspect failure modes. Zero-signup option via x402 micropayments.

3 Python Updated 1mo ago Score 49
MCP Server Stale

tan-yong-sheng/ai-vision-mcp

Multimodal AI vision MCP server for image, video, and object detection analysis. Enables UI/UX evaluation, visual regression testing, and interface understanding using Google Gemini and Vertex AI.

69 TypeScript Updated 3mo ago Score 46
MCP Server Maintained

sh6drack/zen-mcp

Zen Browser automation via WebDriver BiDi. 20 tools for navigation, form filling, screenshots, and JavaScript evaluation. No Selenium or Playwright required.

15 JavaScript Updated 2mo ago Score 45
MCP Server Stale

mattjoyce/mcp-persona-sessions

Enable AI assistants to conduct structured, persona-driven sessions including interview preparation, personal reflection, and coaching conversations. Built-in timer management and performance evaluation tools.

12 Python Updated 5mo ago Score 39