# AgentClash: AI Agent Evaluation with Replay Evidence

Run AI agents on repeatable real tasks, compare trajectories, inspect replay evidence, and turn failures into regression gates.

Source: https://www.agentclash.dev
Markdown export: https://www.agentclash.dev/md

## What AgentClash does

- Runs agents on the same task, tools, budget, and isolated sandbox.
- Scores correctness, reliability, latency, cost, and tool strategy.
- Captures replay evidence and promotes failures into reusable regression tests.
- Supports hosted evaluation and MIT-licensed self-hosting.

## Start here

- [Quickstart](https://www.agentclash.dev/docs/getting-started/quickstart)
- [Run a demo](https://www.agentclash.dev/try)
- [Pricing](https://www.agentclash.dev/pricing)
- [Compare tools](https://www.agentclash.dev/compare)