# Challenge packs for repeatable agent evaluation

Challenge packs are AgentClash's unit of agent evaluation: a real task, tool policy, scoring rules, and artifacts encoded once so every model or harness change reruns the same workload.

Source: https://www.agentclash.dev/features/challenge-packs
Markdown export: https://www.agentclash.dev/md/features/challenge-packs

## What challenge packs encode

Inputs, sandbox resources, allowed tools, validators, judges, and pass conditions — everything needed for a fair, repeatable agent eval.

- OpenTelemetry-compatible trace import
- Pinned datasets and golden test cases
- Baseline versus candidate regression checks
- Replay trails for tool calls, outputs, and artifacts
- Scorecards for correctness, cost, latency, and evidence
- CI gates for prompt, model, RAG, and tool changes

## Pack lifecycle

### 1. Import the evidence

Start from OpenTelemetry traces, curated datasets, support transcripts, or a real failure your team already saw.

### 2. Pin the baseline

Record the current accepted behavior so every prompt, model, RAG, or tool change has a fair comparison point.

### 3. Replay the evidence

Inspect tool calls, outputs, artifacts, latency, cost, and judge evidence when a candidate gets worse.

### 4. Gate the release

Compare candidate and baseline runs, then fail CI before a regression reaches users.

## Author packs with docs

Read the challenge pack docs and authoring guide, then promote escaped failures into packs your whole team can run.

- [Agent evals](https://www.agentclash.dev/agent-evals): Real-task agent evals with replay evidence and CI gates.
- [LLM agent evaluation](https://www.agentclash.dev/llm-agent-evaluation): Evaluate LLM agents on full trajectories, not one-shot answers.
- [Compare tools](https://www.agentclash.dev/compare): See how AgentClash differs from prompt-eval platforms.
- [Challenge packs docs](https://www.agentclash.dev/docs/challenge-packs): Overview of pack structure, scoring, and execution modes.
- [Write a challenge pack](https://www.agentclash.dev/docs/guides/write-a-challenge-pack): Step-by-step authoring guide for your first pack.
- [Datasets overview](https://www.agentclash.dev/docs/guides/datasets-overview): Import examples, record baselines, sync regression suites, and gate CI.
- [Dataset CI gates](https://www.agentclash.dev/docs/guides/dataset-ci-gates): Fail builds when a candidate regresses against a pinned baseline.

## Challenge packs FAQ

### What is a challenge pack?

A challenge pack is a versioned agent evaluation workload with inputs, tool policy, scoring rules, and expected artifacts.

### Can challenge packs run locally and in CI?

Yes. The same pack can power exploratory evals, hosted runs, and pull request gates.

### How do teams create challenge packs?

Start from a real failure or release risk, encode it as YAML, and iterate with replay evidence until the scoring rules match what reviewers expect.