Search: evaluation
Search MCP servers and agent skills by name, description, category or topic — 26 results.
huggingface/hugging-face-evaluation
Model evaluation with vLLM/lighteval and eval tables
NVIDIA/Model-Optimizer/evaluation
Evaluates accuracy of quantized or unquantized LLMs using NeMo Evaluator Launcher (NEL).
santifer/career-ops
14-skill collection for AI-powered job search: JD evaluation with A-F scoring, ATS-optimized PDF generation, portal scanners (Greenhouse/Ashby/Lever), interview prep with STAR+R, batch processing, and a Go dashboard TUI
DataEval/dingo
MCP server for the Dingo: a comprehensive data quality evaluation tool. Server Enables interaction with Dingo's rule-based and LLM-based evaluation capabilities and rules&prompts listing.
NVIDIA/Model-Optimizer/accessing-mlflow
Query and browse evaluation results stored in MLflow.
NVIDIA/Model-Optimizer/launching-evals
Run, monitor, analyze, and debug LLM evaluations via nemo-evaluator-launcher.
NVIDIA/Model-Optimizer/monitor
Monitor submitted jobs (PTQ, evaluation, deployment) on SLURM clusters.
NVIDIA/NeMo-Evaluator/byob
Create custom LLM evaluation benchmarks using the BYOB decorator framework.
NVIDIA/NeMo-Evaluator-Launcher/accessing-mlflow
Query and browse evaluation results stored in MLflow.
NVIDIA/NeMo-Evaluator-Launcher/launching-evals
Run, monitor, analyze, and debug LLM evaluations via nemo-evaluator-launcher.
hidai25/eval-view
Regression testing framework for AI agents. Save golden baselines, detect behavioral drift, and block regressions in CI. Works with LangGraph, CrewAI, OpenAI, Claude, and any HTTP API.
smixs/creative-director-skill
AI creative director with recursive self-assessment: 20+ methodologies (SIT, TRIZ, Bisociation, SCAMPER, Synectics), 3-axis evaluation calibrated against Cannes/D&AD/HumanKind, 5-phase process from brief to presentation
sidclawhq/platform
Governance proxy for MCP servers. Wraps any upstream server with policy evaluation, human approval workflows, and hash-chain audit trails. 18+ framework integrations. Apache 2.0 SDK.
jasonjmcghee/claude-debugs-for-you
An MCP Server and VS Code Extension which enables (language agnostic) automatic debugging via breakpoints and expression evaluation.
iris-eval/mcp-server
MCP-native agent evaluation and observability server with trace logging, output quality evaluation, cost tracking, 12 built-in eval rules, real-time dashboard, and PII detection.
OrygnsCode/opa-mcp-server
Open Policy Agent (OPA) and Rego policy toolkit. 32 tools spanning authoring (format, lint, check, deps), evaluation (eval, test, bench, coverage), and OPA REST control (policies, data, decisions, compile). Wraps the OPA CLI and the [Regal](https://github.com/StyraInc/regal) linter, with AI-assisted helpers for explaining decisions, generating test skeletons, and suggesting fixes.
heurist-network/heurist-mesh-mcp-server
Access specialized web3 AI agents for blockchain analysis, smart contract security auditing, token metrics evaluation, and on-chain interactions through the Heurist Mesh network. Provides comprehensive tools for DeFi analysis, NFT valuation, and transaction monitoring across multiple blockchains
verifyax/verifyax-mcp
MCP server for the VerifyAX platform. Enables agent evaluation, simulation testing, and functional/non-functional verification workflows through natural language.
shuji-bonji/xcomet-mcp-server
Translation quality evaluation using xCOMET models. Provides quality scoring (0-1), error detection with severity levels (minor/major/critical), and optimized batch processing with 25x speedup.
ankitkapur1992-hlido/hlido-mcp
Independent trust scores, claim audits, and comparisons for AI agents — queryable by your agent over MCP. Hosted Cloudflare Worker at hlido.eu/mcp (no install). Returns a 0–100 score, tier verdict, per-claim PASS/FAIL audit, and signed evidence for a reviewed agent. From [Hlido](https://hlido.eu).
mumez/pharo-smalltalk-interop-mcp-server
Pharo Smalltalk integration enabling code evaluation, class/method introspection, package management, test execution, and project installation for interactive development with Pharo images.
supertrained/rhumb
Agent-native tool intelligence across 1,000+ scored services. 21 MCP tools: discover services, check AN Scores, compare alternatives, resolve capabilities to ranked providers, execute through 3 credential modes (managed, BYOK, agent vault), track costs with receipts, and inspect failure modes. Zero-signup option via x402 micropayments.
tan-yong-sheng/ai-vision-mcp
Multimodal AI vision MCP server for image, video, and object detection analysis. Enables UI/UX evaluation, visual regression testing, and interface understanding using Google Gemini and Vertex AI.
sh6drack/zen-mcp
Zen Browser automation via WebDriver BiDi. 20 tools for navigation, form filling, screenshots, and JavaScript evaluation. No Selenium or Playwright required.
mattjoyce/mcp-persona-sessions
Enable AI assistants to conduct structured, persona-driven sessions including interview preparation, personal reflection, and coaching conversations. Built-in timer management and performance evaluation tools.