옵션

agent-eval

affaan-m/ECC affaan-m/ECC

재현 가능한 작업에서 코딩 에이전트들을 직접 비교하여 합격률, 비용, 소요 시간 및 일관성 지표를 분석합니다.

...모든 것을 확장하십시오
0
업데이트 된 시간 2026년 10월 1일

에이전트 평가 스킬

재현 가능한 작업에서 코딩 에이전트들을 직접 비교할 수 있는 경량 CLI 도구입니다. "어떤 코딩 에이전트가 가장 좋을까?"라는 모든 비교는 vibes에서 실행되며, 이 도구는 이를 체계화합니다.

사용 시점

  • 자신의 코드베이스에서 코딩 에이전트(Claude Code, Aider, Codex 등)를 비교할 때
  • 새로운 도구나 모델을 도입하기 전에 에이전트의 성능을 측정할 때
  • 에이전트가 모델이나 툴을 업데이트했을 때 회귀 테스트를 수행할 때
  • 팀을 위해 데이터에 기반한 에이전트 선정 결정을 내릴 때

설치

참고: 소스 코드를 검토한 후, agent-eval 저장소에서 해당 도구를 설치하십시오.

핵심 개념

YAML 작업 정의

태스크를 선언적으로 정의합니다. 각 태스크는 수행할 작업, 처리할 파일, 성공 여부를 판단하는 방법을 명시합니다:

name: add-retry-logic
description: Add exponential backoff retry to the HTTP client
repo: ./my-project
files:
  - src/http_client.py
prompt: |
  Add retry logic with exponential backoff to all HTTP requests.
  Max 3 retries. Initial delay 1s, max delay 30s.
judge:
  - type: pytest
    command: pytest tests/test_http_client.py -v
  - type: grep
    pattern: "exponential_backoff|retry"
    files: src/http_client.py
commit: "abc1234"  # pin to specific commit for reproducibility

Git 워크트리 격리

각 에이전트 실행에는 고유한 Git 워크트리가 할당되며, Docker가 필요하지 않습니다. 이를 통해 재현성 격리가 보장되므로, 에이전트 간에 상호 간섭이 발생하거나 기본 저장소가 손상되는 것을 방지할 수 있습니다.

수집된 메트릭

메트릭 측정 항목
통과율 에이전트가 심사 기준을 통과하는 코드를 생성했습니까?
비용 작업당 API 사용량 (사용 가능한 경우)
시간 완료까지 소요된 실제 시간(초)
일관성 반복 실행 시 통과율 (예: 3/3 = 100%)

워크플로우

1. 작업 정의

YAML 파일이 포함된 tasks/ YAML 파일이 들어 있는 디렉터리를 생성합니다(작업당 하나씩):

mkdir tasks
# Write task definitions (see template above)

2. 에이전트 실행

작업에 대해 에이전트를 실행합니다:

agent-eval run --task tasks/add-retry-logic.yaml --agent claude-code --agent aider --runs 3

각 실행 시:

  1. 지정된 커밋을 기반으로 새로운 git 작업 디렉터리를 생성합니다
  2. 에이전트에 프롬프트를 전달합니다
  3. 심사 기준을 실행합니다
  4. 합격/불합격, 비용 및 시간을 기록합니다

3. 결과 비교

비교 보고서 생성:

agent-eval report --format table
Task: add-retry-logic (3 runs each)
┌──────────────┬───────────┬────────┬────────┬─────────────┐
│ Agent        │ Pass Rate │ Cost   │ Time   │ Consistency │
├──────────────┼───────────┼────────┼────────┼─────────────┤
│ claude-code  │ 3/3       │ $0.12  │ 45s    │ 100%        │
│ aider        │ 2/3       │ $0.08  │ 38s    │  67%        │
└──────────────┴───────────┴────────┴────────┴─────────────┘

평가 유형

코드 기반(결정론적)

judge:
  - type: pytest
    command: pytest tests/ -v
  - type: command
    command: npm run build

패턴 기반

judge:
  - type: grep
    pattern: "class.*Retry"
    files: src/**/*.py

모델 기반(LLM을 판정자로 활용)

judge:
  - type: llm
    prompt: |
      Does this implementation correctly handle exponential backoff?
      Check for: max retries, increasing delays, jitter.

모범 사례

  • 단순한 예제가 아닌 실제 워크로드를 반영하는 3~5개의 작업으로 시작하십시오
  • 에이전트당 최소 3회의 시도를 실행하여 변동성을 포착하십시오 — 에이전트는 비결정적입니다
  • 결과가 며칠 또는 몇 주에 걸쳐 재현될 수 있도록 작업 YAML에 커밋을 고정하십시오
  • 작업당 최소 한 개의 결정론적 심사자(테스트, 빌드)를 포함하십시오 — LLM 심사자는 노이즈를 유발합니다
  • 통과율과 함께 비용도 추적하십시오 — 비용이 10배나 들지만 95%의 통과율을 보이는 에이전트는 올바른 선택이 아닐 수 있습니다
  • 작업 정의를 버전 관리하세요 — 이는 테스트 고정 장치이므로 코드처럼 취급하십시오

링크

  • 저장소: github.com/joaquinhuigomez/agent-eval
GitHub에서 보기
---
name: agent-eval
description: Compare coding agents head-to-head on reproducible tasks with pass rate, cost, time, and consistency metrics.
---

# Agent Eval Skill

A lightweight CLI tool for comparing coding agents head-to-head on reproducible tasks. Every "which coding agent is best?" comparison runs on vibes — this tool systematizes it.

## When to Activate

- Comparing coding agents (Claude Code, Aider, Codex, etc.) on your own codebase
- Measuring agent performance before adopting a new tool or model
- Running regression checks when an agent updates its model or tooling
- Producing data-backed agent selection decisions for a team

## Installation

> **Note:** Install agent-eval from its repository after reviewing the source.

## Core Concepts

### YAML Task Definitions

Define tasks declaratively. Each task specifies what to do, which files to touch, and how to judge success:

```yaml
name: add-retry-logic
description: Add exponential backoff retry to the HTTP client
repo: ./my-project
files:
  - src/http_client.py
prompt: |
  Add retry logic with exponential backoff to all HTTP requests.
  Max 3 retries. Initial delay 1s, max delay 30s.
judge:
  - type: pytest
    command: pytest tests/test_http_client.py -v
  - type: grep
    pattern: "exponential_backoff|retry"
    files: src/http_client.py
commit: "abc1234"  # pin to specific commit for reproducibility
```

### Git Worktree Isolation

Each agent run gets its own git worktree — no Docker required. This provides reproducibility isolation so agents cannot interfere with each other or corrupt the base repo.

### Metrics Collected

| Metric | What It Measures |
|--------|-----------------|
| Pass rate | Did the agent produce code that passes the judge? |
| Cost | API spend per task (when available) |
| Time | Wall-clock seconds to completion |
| Consistency | Pass rate across repeated runs (e.g., 3/3 = 100%) |

## Workflow

### 1. Define Tasks

Create a `tasks/` directory with YAML files, one per task:

```bash
mkdir tasks
# Write task definitions (see template above)
```

### 2. Run Agents

Execute agents against your tasks:

```bash
agent-eval run --task tasks/add-retry-logic.yaml --agent claude-code --agent aider --runs 3
```

Each run:
1. Creates a fresh git worktree from the specified commit
2. Hands the prompt to the agent
3. Runs the judge criteria
4. Records pass/fail, cost, and time

### 3. Compare Results

Generate a comparison report:

```bash
agent-eval report --format table
```

```
Task: add-retry-logic (3 runs each)
┌──────────────┬───────────┬────────┬────────┬─────────────┐
│ Agent        │ Pass Rate │ Cost   │ Time   │ Consistency │
├──────────────┼───────────┼────────┼────────┼─────────────┤
│ claude-code  │ 3/3       │ $0.12  │ 45s    │ 100%        │
│ aider        │ 2/3       │ $0.08  │ 38s    │  67%        │
└──────────────┴───────────┴────────┴────────┴─────────────┘
```

## Judge Types

### Code-Based (deterministic)

```yaml
judge:
  - type: pytest
    command: pytest tests/ -v
  - type: command
    command: npm run build
```

### Pattern-Based

```yaml
judge:
  - type: grep
    pattern: "class.*Retry"
    files: src/**/*.py
```

### Model-Based (LLM-as-judge)

```yaml
judge:
  - type: llm
    prompt: |
      Does this implementation correctly handle exponential backoff?
      Check for: max retries, increasing delays, jitter.
```

## Best Practices

- **Start with 3-5 tasks** that represent your real workload, not toy examples
- **Run at least 3 trials** per agent to capture variance — agents are non-deterministic
- **Pin the commit** in your task YAML so results are reproducible across days/weeks
- **Include at least one deterministic judge** (tests, build) per task — LLM judges add noise
- **Track cost alongside pass rate** — a 95% agent at 10x the cost may not be the right choice
- **Version your task definitions** — they are test fixtures, treat them as code

## Links

- Repository: [github.com/joaquinhuigomez/agent-eval](https://github.com/joaquinhuigomez/agent-eval)

모든 파일

1개 파일

agent-eval 설치

스킬 파일을 다운로드하여 .claude/skills/ 디렉터리에 압축을 풀어주세요.

ZIP 다운로드

저장소를 클론하고 스킬 파일을 프로젝트에 복사하세요.

git clone https://github.com/affaan-m/ECC/tree/main/skills/agent-eval # Copy SKILL.md to your .claude/skills/ directory

복사 복사
빠른 설정: 스킬 폴더를 .claude/skills/로 복사하세요. Claude가 해당 스킬을 자동으로 감지하여 사용할 것입니다.
저장소 affaan-m/ECC

관련 스킬

web-search
업데이트 된 시간 2026년 6월 29일
webapp-testing
업데이트 된 시간 2026년 6월 29일
lark-base
업데이트 된 시간 2026년 7월 5일
agentmail
업데이트 된 시간 2026년 6월 29일
OR