Features
Agent evaluation features
See how AgentClash turns agent runs into reviewable scorecards, replay evidence, and reusable challenge packs for CI-ready evaluation.
Explore AgentClash features for agent scorecards, replay evidence, and challenge packs that turn real tasks into repeatable evaluations.
Agent scorecards
Scorecards for correctness, cost, latency, and evidence quality.
Agent replay
Replay tool calls, artifacts, and evidence for agent debugging.
Challenge packs
Repeatable agent evaluation workloads with scoring and CI gates.
Synthetic dataset generation
Hosted weak-vs-strong generation for eval-ready datasets.