agent-eval
affaan-m/ECC
재현 가능한 작업에서 코딩 에이전트들을 직접 비교하여 합격률, 비용, 소요 시간 및 일관성 지표를 분석합니다.
...모든 것을 확장하십시오에이전트 평가 스킬
재현 가능한 작업에서 코딩 에이전트들을 직접 비교할 수 있는 경량 CLI 도구입니다. "어떤 코딩 에이전트가 가장 좋을까?"라는 모든 비교는 vibes에서 실행되며, 이 도구는 이를 체계화합니다.
사용 시점
- 자신의 코드베이스에서 코딩 에이전트(Claude Code, Aider, Codex 등)를 비교할 때
- 새로운 도구나 모델을 도입하기 전에 에이전트의 성능을 측정할 때
- 에이전트가 모델이나 툴을 업데이트했을 때 회귀 테스트를 수행할 때
- 팀을 위해 데이터에 기반한 에이전트 선정 결정을 내릴 때
설치
참고: 소스 코드를 검토한 후, agent-eval 저장소에서 해당 도구를 설치하십시오.
핵심 개념
YAML 작업 정의
태스크를 선언적으로 정의합니다. 각 태스크는 수행할 작업, 처리할 파일, 성공 여부를 판단하는 방법을 명시합니다:
name: add-retry-logic
description: Add exponential backoff retry to the HTTP client
repo: ./my-project
files:
- src/http_client.py
prompt: |
Add retry logic with exponential backoff to all HTTP requests.
Max 3 retries. Initial delay 1s, max delay 30s.
judge:
- type: pytest
command: pytest tests/test_http_client.py -v
- type: grep
pattern: "exponential_backoff|retry"
files: src/http_client.py
commit: "abc1234" # pin to specific commit for reproducibility
Git 워크트리 격리
각 에이전트 실행에는 고유한 Git 워크트리가 할당되며, Docker가 필요하지 않습니다. 이를 통해 재현성 격리가 보장되므로, 에이전트 간에 상호 간섭이 발생하거나 기본 저장소가 손상되는 것을 방지할 수 있습니다.
수집된 메트릭
| 메트릭 | 측정 항목 |
|---|---|
| 통과율 | 에이전트가 심사 기준을 통과하는 코드를 생성했습니까? |
| 비용 | 작업당 API 사용량 (사용 가능한 경우) |
| 시간 | 완료까지 소요된 실제 시간(초) |
| 일관성 | 반복 실행 시 통과율 (예: 3/3 = 100%) |
워크플로우
1. 작업 정의
YAML 파일이 포함된 tasks/ YAML 파일이 들어 있는 디렉터리를 생성합니다(작업당 하나씩):
mkdir tasks
# Write task definitions (see template above)
2. 에이전트 실행
작업에 대해 에이전트를 실행합니다:
agent-eval run --task tasks/add-retry-logic.yaml --agent claude-code --agent aider --runs 3
각 실행 시:
- 지정된 커밋을 기반으로 새로운 git 작업 디렉터리를 생성합니다
- 에이전트에 프롬프트를 전달합니다
- 심사 기준을 실행합니다
- 합격/불합격, 비용 및 시간을 기록합니다
3. 결과 비교
비교 보고서 생성:
agent-eval report --format table
Task: add-retry-logic (3 runs each)
┌──────────────┬───────────┬────────┬────────┬─────────────┐
│ Agent │ Pass Rate │ Cost │ Time │ Consistency │
├──────────────┼───────────┼────────┼────────┼─────────────┤
│ claude-code │ 3/3 │ $0.12 │ 45s │ 100% │
│ aider │ 2/3 │ $0.08 │ 38s │ 67% │
└──────────────┴───────────┴────────┴────────┴─────────────┘
평가 유형
코드 기반(결정론적)
judge:
- type: pytest
command: pytest tests/ -v
- type: command
command: npm run build
패턴 기반
judge:
- type: grep
pattern: "class.*Retry"
files: src/**/*.py
모델 기반(LLM을 판정자로 활용)
judge:
- type: llm
prompt: |
Does this implementation correctly handle exponential backoff?
Check for: max retries, increasing delays, jitter.
모범 사례
- 단순한 예제가 아닌 실제 워크로드를 반영하는 3~5개의 작업으로 시작하십시오
- 에이전트당 최소 3회의 시도를 실행하여 변동성을 포착하십시오 — 에이전트는 비결정적입니다
- 결과가 며칠 또는 몇 주에 걸쳐 재현될 수 있도록 작업 YAML에 커밋을 고정하십시오
- 작업당 최소 한 개의 결정론적 심사자(테스트, 빌드)를 포함하십시오 — LLM 심사자는 노이즈를 유발합니다
- 통과율과 함께 비용도 추적하십시오 — 비용이 10배나 들지만 95%의 통과율을 보이는 에이전트는 올바른 선택이 아닐 수 있습니다
- 작업 정의를 버전 관리하세요 — 이는 테스트 고정 장치이므로 코드처럼 취급하십시오
링크
- 저장소: github.com/joaquinhuigomez/agent-eval
---
name: agent-eval
description: Compare coding agents head-to-head on reproducible tasks with pass rate, cost, time, and consistency metrics.
---
# Agent Eval Skill
A lightweight CLI tool for comparing coding agents head-to-head on reproducible tasks. Every "which coding agent is best?" comparison runs on vibes — this tool systematizes it.
## When to Activate
- Comparing coding agents (Claude Code, Aider, Codex, etc.) on your own codebase
- Measuring agent performance before adopting a new tool or model
- Running regression checks when an agent updates its model or tooling
- Producing data-backed agent selection decisions for a team
## Installation
> **Note:** Install agent-eval from its repository after reviewing the source.
## Core Concepts
### YAML Task Definitions
Define tasks declaratively. Each task specifies what to do, which files to touch, and how to judge success:
```yaml
name: add-retry-logic
description: Add exponential backoff retry to the HTTP client
repo: ./my-project
files:
- src/http_client.py
prompt: |
Add retry logic with exponential backoff to all HTTP requests.
Max 3 retries. Initial delay 1s, max delay 30s.
judge:
- type: pytest
command: pytest tests/test_http_client.py -v
- type: grep
pattern: "exponential_backoff|retry"
files: src/http_client.py
commit: "abc1234" # pin to specific commit for reproducibility
```
### Git Worktree Isolation
Each agent run gets its own git worktree — no Docker required. This provides reproducibility isolation so agents cannot interfere with each other or corrupt the base repo.
### Metrics Collected
| Metric | What It Measures |
|--------|-----------------|
| Pass rate | Did the agent produce code that passes the judge? |
| Cost | API spend per task (when available) |
| Time | Wall-clock seconds to completion |
| Consistency | Pass rate across repeated runs (e.g., 3/3 = 100%) |
## Workflow
### 1. Define Tasks
Create a `tasks/` directory with YAML files, one per task:
```bash
mkdir tasks
# Write task definitions (see template above)
```
### 2. Run Agents
Execute agents against your tasks:
```bash
agent-eval run --task tasks/add-retry-logic.yaml --agent claude-code --agent aider --runs 3
```
Each run:
1. Creates a fresh git worktree from the specified commit
2. Hands the prompt to the agent
3. Runs the judge criteria
4. Records pass/fail, cost, and time
### 3. Compare Results
Generate a comparison report:
```bash
agent-eval report --format table
```
```
Task: add-retry-logic (3 runs each)
┌──────────────┬───────────┬────────┬────────┬─────────────┐
│ Agent │ Pass Rate │ Cost │ Time │ Consistency │
├──────────────┼───────────┼────────┼────────┼─────────────┤
│ claude-code │ 3/3 │ $0.12 │ 45s │ 100% │
│ aider │ 2/3 │ $0.08 │ 38s │ 67% │
└──────────────┴───────────┴────────┴────────┴─────────────┘
```
## Judge Types
### Code-Based (deterministic)
```yaml
judge:
- type: pytest
command: pytest tests/ -v
- type: command
command: npm run build
```
### Pattern-Based
```yaml
judge:
- type: grep
pattern: "class.*Retry"
files: src/**/*.py
```
### Model-Based (LLM-as-judge)
```yaml
judge:
- type: llm
prompt: |
Does this implementation correctly handle exponential backoff?
Check for: max retries, increasing delays, jitter.
```
## Best Practices
- **Start with 3-5 tasks** that represent your real workload, not toy examples
- **Run at least 3 trials** per agent to capture variance — agents are non-deterministic
- **Pin the commit** in your task YAML so results are reproducible across days/weeks
- **Include at least one deterministic judge** (tests, build) per task — LLM judges add noise
- **Track cost alongside pass rate** — a 95% agent at 10x the cost may not be the right choice
- **Version your task definitions** — they are test fixtures, treat them as code
## Links
- Repository: [github.com/joaquinhuigomez/agent-eval](https://github.com/joaquinhuigomez/agent-eval)
모든 파일
1개 파일agent-eval 설치
스킬 파일을 다운로드하여 .claude/skills/ 디렉터리에 압축을 풀어주세요.
ZIP 다운로드저장소를 클론하고 스킬 파일을 프로젝트에 복사하세요.
git clone https://github.com/affaan-m/ECC/tree/main/skills/agent-eval # Copy SKILL.md to your .claude/skills/ directory
복사





집
