# Evaluate coding agents on real repositories

Coding agents fail on messy repos, flaky tests, and tool limits — not on polished benchmark prompts. AgentClash evaluates patches, test runs, artifacts, and cost on workloads your team actually ships.

Source: https://www.agentclash.dev/use-cases/coding-agent-evaluation
Markdown export: https://www.agentclash.dev/md/use-cases/coding-agent-evaluation

## What coding agent eval should check

Look for correct fixes, sane tool usage, reproducible artifacts, and stable cost/latency — not just a plausible diff in a demo.

- OpenTelemetry-compatible trace import
- Pinned datasets and golden test cases
- Baseline versus candidate regression checks
- Replay trails for tool calls, outputs, and artifacts
- Scorecards for correctness, cost, latency, and evidence
- CI gates for prompt, model, RAG, and tool changes

## Coding eval workflow

### 1. Import the evidence

Start from OpenTelemetry traces, curated datasets, support transcripts, or a real failure your team already saw.

### 2. Pin the baseline

Record the current accepted behavior so every prompt, model, RAG, or tool change has a fair comparison point.

### 3. Replay the evidence

Inspect tool calls, outputs, artifacts, latency, cost, and judge evidence when a candidate gets worse.

### 4. Gate the release

Compare candidate and baseline runs, then fail CI before a regression reaches users.

## Start with a real bug

Promote a failed coding-agent run into a challenge pack, then eval model and harness changes before they reach users.

- [Agent evals](https://www.agentclash.dev/agent-evals): Real-task agent evals with replay evidence and CI gates.
- [LLM agent evaluation](https://www.agentclash.dev/llm-agent-evaluation): Evaluate LLM agents on full trajectories, not one-shot answers.
- [Compare tools](https://www.agentclash.dev/compare): See how AgentClash differs from prompt-eval platforms.
- [Datasets overview](https://www.agentclash.dev/docs/guides/datasets-overview): Import examples, record baselines, sync regression suites, and gate CI.
- [Dataset CI gates](https://www.agentclash.dev/docs/guides/dataset-ci-gates): Fail builds when a candidate regresses against a pinned baseline.
- [CI/CD agent gates](https://www.agentclash.dev/docs/guides/ci-cd-agent-gates): Block pull requests when agent behavior gets worse.

## Coding agent evaluation FAQ

### Can AgentClash evaluate agents that edit code?

Yes. Challenge packs can run agents in sandboxes with repository fixtures, test commands, and artifact checks for patches and logs.

### How do teams compare coding agents fairly?

Run every candidate on the same repo state, tool policy, and time budget, then compare scorecards and replay evidence.

### Can coding agent evals run in CI?

Yes. Wire challenge packs into pull request gates so model, prompt, or tool changes cannot merge when correctness regresses.