‹ The Index

Agent Eval Harness

plugin

Generic agentic evaluation for skills and agents. Provides end-to-end skills to analyze, test, score, review, and iteratively improve agent skills with MLflow support for experiment tracking, tracing, and reporting.

Works with: Claude Code, Cursor, Claude Desktop, Codex CLI, Gemini CLI, Cline, Windsurf, VS Code

Category: Dev Tools & CI — see all ranked ›

Work: Model evaluation · Agent development

Who it is for: AI engineer

thin

“Registry README about registry.yaml and nightly SHA pinning; this plugin is never named.”

Security audit

Not scanned yet. We audit npm-published capabilities for known advisories, install-time scripts and permission surface; this one has no npm package we can resolve, or has not reached the queue.

source ↗  ·  plugin:redhat-global-engineering/ge-public-skills/agent-eval-harness

Already running this? npx tashan-cli doctor checks your whole config against the Index — how it works ›