‹ The Index
Eval
skill
Evaluate and rank agent results by metric or LLM judge for an AgentHub session. Use when the user runs /hub:eval or asks to score, compare, or pick a winner among completed AgentHub agents.
Works with: Claude Code, Cursor, Codex CLI
Category: Dev Tools & CI — see all ranked ›
Work: Model evaluation
Who it is for: AI engineer
- Adoption: 1 repos
- Health: active
- GitHub stars: 23,181
- Contributors: 30
- License: MIT
Security audit
Not scanned yet. We audit npm-published capabilities for known advisories, install-time scripts and permission surface; this one has no npm package we can resolve, or has not reached the queue.
source ↗ · skill:alirezarezvani/eval
Already running this? npx tashan-cli doctor checks your whole config against the Index — how it works ›