옵션

반복적인 실험을 통해 신뢰할 수 있는 브라우저 자동화 기술을 습득하며, 내부 에이전트를 실행해 사이트를 탐색하고, 작업이 일관되게 성공할 때까지 탐색 지침을 개선해 나갑니다.

...모든 것을 확장하십시오
0
업데이트 된 시간 2026년 9월 30일

AutoBrowse — 스스로 발전하는 브라우저 자동화 기술

반복적인 실험을 통해 신뢰할 수 있는 브라우저 자동화 기술을 쌓아보세요. 내부 에이전트가 사이트를 탐색하고 (evaluate.ts). 외부 에이전트인 여러분은 발생한 상황을 분석하여 명령어를 개선합니다(strategy.md). 테스트가 일관되게 통과될 때까지 이 과정을 반복합니다.

시작점

호출 방식은 유연합니다. 명시적인 플래그와 자유 형식의 자연어 모두 사용할 수 있습니다:

/autobrowse --task google-flights
/autobrowse --task google-flights --iterations 10 --env remote
/autobrowse --task google-flights --browser-trace
/autobrowse --tasks google-flights,amazon-add-to-cart
/autobrowse --all

# Also fine — parse freely:
/autobrowse https://flights.google.com/
/autobrowse book a flight on delta.com
/autobrowse fix the existing google-flights skill

--browser-trace (기본값: 비활성화, 원격 전용): 각 반복을 형제 browser-trace 스킬과 페어링하여, 페이지별 네트워크/콘솔/페이지 수명 주기 증거를 수집하기 위해 내부 에이전트를 CDP 캡처로 감쌉니다. 이는 --env remote; 다음과 함께 사용 시 오류가 발생합니다 --env local와 결합 시 오류가 발생합니다. 형제스킬이 존재해야 합니다. browser-trace 스킬이 ${CLAUDE_SKILL_DIR}/../browser-trace/, 그리고 BROWSERBASE_API_KEY env 변수가 필요합니다.

사용자가 --task :

  • 기존 작업 중 ${WORKSPACE}/tasks/ 사이트/의도(site/intent)와 명확하게 일치하는 경우, 이를 사용합니다.
  • 그렇지 않은 경우, 짧은 케밥 케이스(kebab-case) 이름을 선택하고 ${WORKSPACE}/tasks//task.md 다음에서 ${CLAUDE_SKILL_DIR}/references/example-task.md를 생성하고, 사용자의 발언을 바탕으로 URL/목표를 입력한 뒤 진행하십시오. 선택한 이름을 한 줄로 사용자에게 알려주십시오.

실행 방법

1단계 — 인수를 분석하고 맥락을 파악하세요

전달된 내용을 확인합니다:

  • --task → 단일 작업 모드
  • --tasks a,b,c 또는 --all → 다중 작업 모드 (하위 에이전트 생성)
  • --iterations N → 평가 횟수 → 개선 사이클 수 (기본값: 5)
  • --env local|remote → 브라우저 환경 (기본값: 로컬; 봇 차단 사이트의 경우 원격 사용)
  • --browser-trace → 브라우저 추적 통합 활성화 (기본값: 비활성화). 이는 --env remote. 만약 --env local --browser-trace 두 가지가 모두 명시적으로 전달된 경우, 다음 오류가 발생합니다: browser-trace requires Browserbase; drop --env local or drop --browser-trace.

사용자가 대신 자유 형식 텍스트를 전달한 경우, 계속 진행하기 전에 이를 위의 항목 중 하나에 매핑하십시오.

2단계 — 작업 공간 설정

모든 훈련 아티팩트(작업 정의, 전략 반복, 트레이스, 보고서)는 현재 작업 디렉터리의 워크스페이스 디렉터리에 저장되며, ~/.claude/skills/. 이를 통해 내부 에이전트의 파일 쓰기 작업이 Claude의 홈 디렉터리에 기록되지 않도록 하여 권한 관련 문제를 방지합니다.

기본 작업 공간: ${CWD}/autobrowse/

mkdir -p ./autobrowse/tasks ./autobrowse/traces ./autobrowse/reports

작업 디렉터리(./autobrowse/tasks//task.md)가 아직 존재하지 않는다면, 다음과 같이 생성하십시오:

mkdir -p ./autobrowse/tasks/
cp ${CLAUDE_SKILL_DIR}/references/example-task.md ./autobrowse/tasks//task.md
# Then edit task.md to describe the URL, inputs, steps, and expected JSON output

다음 위치의 스킬 소스 파일은 ${CLAUDE_SKILL_DIR} 에 있는 스킬 소스 코드는 읽기 전용으로 유지되며, ./autobrowse/ CWD 내의 파일에만 기록됩니다. 졸업(최종 단계)에서는 단일 파일을 ~/.claude/skills//SKILL.md.

사용 가능한 작업 목록:

ls ./autobrowse/tasks/

3단계 — 멀티태스크: 병렬 서브 에이전트 생성

여러 태스크를 실행하는 경우, 에이전트 도구를 사용하여 태스크당 하나의 하위 에이전트를 동시에 생성하십시오. 각 하위 에이전트는 해당 태스크에 대한 전체 autobrowse 루프를 실행하기 위한 독립적인 프롬프트를 수신합니다:

"당신은 작업 의 '' 스킬을 실행 중입니다. 작업 공간: (예: /path/to/project/autobrowse). 다음의 반복: 평가 → 추적 읽기 → strategy.md 개선 → 반복. 다음을 사용하십시오 --env . --workspace 를 전달합니다. 상위 호출에서 --browser-trace를 사용한 경우, 모든 반복에서 SKILL.md 루프의 traced-path 블록을 반드시 사용해야 합니다(세션 사전 생성, bb-capture 연결, --connect-url evaluate.mjs에 전달, 중지+이분법, 릴리스) — 기본 단일 명령 경로로 되돌아가지 마십시오. autobrowse 루프 지침을 정확히 따르십시오.

졸업 시, 적절한 agentskills 프론트매터(이름 + 설명)와 함께 ~/.claude/skills//SKILL.md 적절한 agentskills 프론트매터(이름 + 설명)와 함께 스킬을 배포하십시오. strategy.md를 단순히 복사하지 말고, 독립적으로 작동하는 스킬을 작성하십시오.

마지막에는 다음 내용을 포함한 구조화된 요약 보고서를 출력하십시오: 작업 이름, 최종 실행 결과(합격/불합격), 누적 총 비용, 완료된 반복 횟수, 반복별 표(반복 번호, 턴 수, 비용, 상태, 검증된 가설), 그리고 2~3가지 핵심 학습 사항."

모든 하위 에이전트를 병렬로 생성하고, 모두 완료될 때까지 기다린 후, 각 에이전트의 요약 정보를 수집하여 세션 보고서를 작성하십시오.

단일 작업의 경우, 이 단계를 건너뛰고 바로 아래의 루프를 실행하십시오.

루프 (각 작업에 대해 실행)

반복 시작

다음이 존재하는지 확인합니다 ./autobrowse/tasks//task.md 가 존재하는지 확인합니다(존재하지 않으면 템플릿을 기반으로 생성하십시오 — 2단계 참조). strategy.md 첫 실행 시 하네스에 의해 자동으로 비어 있는 상태로 생성됩니다.

필수 조건

  • ANTHROPIC_API_KEY 파일이 환경에(또는 CWD의 .env CWD 내의 파일에 있어야 합니다 — evaluate.mjs 파일을 자동으로 불러옵니다). 파일이 없으면 허스가 명확한 오류 메시지를 출력하고 종료됩니다. 다른 경로에서 키를 찾으려고 하지 마십시오.

내부 에이전트 실행

기본 경로(--browser-trace 없음) — 단일 명령어, 오케스트레이션 없음:

node ${CLAUDE_SKILL_DIR}/scripts/evaluate.mjs --task  --workspace ./autobrowse
# or for bot-protected sites:
node ${CLAUDE_SKILL_DIR}/scripts/evaluate.mjs --task  --workspace ./autobrowse --env remote

이 명령은 브라우저 세션을 실행하고 전체 추적 정보를 ./autobrowse/traces//latest/.

추적 경로(--browser-trace, 원격 전용) — 외부 하네스가 Browserbase 세션을 미리 생성하고, bb-capture 수동 관찰자로 연결한 후, 세션의 connectUrl 를 evaluate.mjs 전달하므로 모든 내부 browse 호출이 --cdp $connectUrl --session autobrowse-main (관찰자에게 전체 Network/Console 이벤트를 제공하는 표준 브라우저 추적 패턴). 이 블록을 반복마다 한 번씩 실행하며 $N 1을 인덱스로 하는 반복 횟수로 설정하여 이 블록을 반복할 때마다 한 번씩 실행하십시오:

# Preflight — fail fast if browser-trace isn't installed alongside autobrowse.
BT_DIR="${CLAUDE_SKILL_DIR}/../browser-trace"
if [ ! -f "$BT_DIR/scripts/bb-capture.mjs" ]; then
  echo "ERROR: --browser-trace requires the browser-trace skill at $BT_DIR." >&2
  echo "Install it by cloning github.com/browserbase/skills and copying skills/browser-trace/" >&2
  echo "into the same parent directory as autobrowse (e.g. ~/.claude/skills/browser-trace/)." >&2
  exit 1
fi

# a. SESSION SETUP — pre-create the keep-alive session and derive its connectUrl
sid=$(browse cloud sessions create --keep-alive --verified --proxies \
  | node -e "let s='';process.stdin.on('data',c=>s+=c).on('end',()=>process.stdout.write(JSON.parse(s).id))")
connect_url=$(browse cloud sessions get "$sid" \
  | node -e "let s='';process.stdin.on('data',c=>s+=c).on('end',()=>process.stdout.write(JSON.parse(s).connectUrl))")

RUN_ID="run-$(printf '%03d' "$N")"
TRACE_ROOT="./autobrowse/traces//$RUN_ID"
mkdir -p "$TRACE_ROOT"
export O11Y_ROOT="$TRACE_ROOT/.o11y"   # park browser-trace output inside the autobrowse run dir
export O11Y_RUN_ID="$RUN_ID"           # tells the browse CLI which run dir to write descriptors.ndjson into

# b. ATTACH BROWSER-TRACE — passive observer; runs in background
node ${CLAUDE_SKILL_DIR}/../browser-trace/scripts/bb-capture.mjs "$sid" "$RUN_ID" &
sleep 2

# c. RUN AUTOBROWSE — connectUrl flag tells evaluate.mjs to inject --cdp/--session
#    into every inner browse call. The inner agent never sees --remote.
node ${CLAUDE_SKILL_DIR}/scripts/evaluate.mjs \
  --task  --workspace ./autobrowse --env remote \
  --connect-url "$connect_url" --run-number "$N"

# d. STOP + BISECT + UNIFY — order matters; bisect needs the session to still
#    exist, and unify-trace joins the bisect output with autobrowse's trace.json
#    into a single time-ordered NDJSON the outer agent reads first each iter.
node ${CLAUDE_SKILL_DIR}/../browser-trace/scripts/stop-capture.mjs "$RUN_ID"
node ${CLAUDE_SKILL_DIR}/../browser-trace/scripts/bisect-cdp.mjs "$RUN_ID"
node ${CLAUDE_SKILL_DIR}/scripts/unify-trace.mjs \
  --trace-dir "$TRACE_ROOT" \
  --o11y-dir "$O11Y_ROOT/$RUN_ID"

# e. RELEASE
browse cloud sessions update "$sid" --status REQUEST_RELEASE

이렇게 하면 내부 에이전트 트레이스가 ./autobrowse/traces//latest/ 에 기록하고, CDP 이분 탐색 결과를 ./autobrowse/traces//latest/.o11y//에 기록합니다. 추적된 browse CLI는 또한 명령어별로 상세한 노드 설명자를 .o11y//cdp/descriptors.ndjson (페이지 구동 호출당 하나의 JSON 객체: target, tag, id, role, accessibleName, attributes, xpath, bounding-rect). 이 설명자 파일은 하류 코드 생성 단계에 입력으로 사용되며, 가설 수립에는 필수적이지 않습니다 — 추적을 읽을 때는 이 파일을 건너뛰어도 됩니다.

추적 기록 읽기

cat ./autobrowse/traces//latest/summary.md

요약 정보에는 소요 시간, 비용, 턴 수, 결정 로그 및 최종 JSON 출력이 포함됩니다.

에이전트가 실패했거나 멈춘 경우, 더 자세히 살펴보세요:

  • 읽기 ./autobrowse/traces//latest/trace.json — 실패한 턴을 검색하세요
  • 'Read' 도구를 사용하여 오류 지점 주변의 스크린샷을 확인하세요

--browser-trace를 사용한 경우 — unified-events.jsonl부터 시작하십시오. 이 하네스는 에이전트의 턴 로그와 브라우저의 CDP 파이어호스를 실행 루트에서 시간 순서대로 정렬된 하나의 NDJSON 스트림으로 결합합니다. 소스 태그가 지정된(source: "agent" | "browser")가 지정된 단일 파일로, 실제 시간 타임스탬프 순서대로 교차 배열되어 있습니다. 파일을 위에서 아래로 훑어보세요. 오류 원인은 대개 인접한 한두 줄에 있습니다(에이전트가 명령 X를 내렸고, 브라우저가 Y로 응답한 경우).

cat ./autobrowse/traces//latest/unified-events.jsonl

구조화된 파일(trace.json, .o11y//cdp/*) 역시 통합 스트림에서 더 자세한 정보가 필요한 부분을 가리킬 때 에이전트가 드릴다운 형태로 활용할 수 있습니다:

필요 사항 드릴다운 파일 또는 명령어
페이지별 합계 + 타이밍(이벤트, 네트워크 횟수, 페이지별 오류) .o11y//cdp/summary.json
모든 실패한 네트워크 요청을 한곳에 모아 보기 .o11y//cdp/network/failed.jsonl
전체 콘솔 예외 페이로드(스택 트레이스 등) .o11y//cdp/console/exceptions.jsonl
페이지별 슬라이스 (N 페이지의 이벤트만) .o11y//cdp/pages//
특정 턴에 대한 전체 추론 텍스트 / 잘림 없는 도구 출력 trace.json (다음 항목으로 필터링 turn === N)
임시 그룹화 쿼리(예: 상위 호스트, 페이지별 오류) O11Y_ROOT=./autobrowse/traces//latest/.o11y node ${CLAUDE_SKILL_DIR}/../browser-trace/scripts/query.mjs

통합 스트림이 기본값이며, 그룹화된 쿼리, 전체 텍스트 페이로드 또는 스트림 필터링으로는 얻을 수 없는 정보를 필요로 할 때만 구조화된 파일을 자세히 살펴보세요.

가설을 하나 세우세요

문제가 발생한 정확한 턴을 찾으십시오. 어떤 단일 휴리스틱이 이를 방지할 수 있었을까요?

‘ --browser-trace가설은 unified-events.jsonl의 특정 이벤트(줄 번호 또는 타임스탬프)를 인용해야 하며, 드릴다운 파일에 접근해야 했다면 해당 파일 이름을 명시해야 합니다. 이를 통해 업데이트가 직감에 의존하기보다는 증거에 기반을 두게 됩니다. 에이전트의 명령어에만 기반한 가설은 “클릭이 작동하지 않았다”라고 말할 수 있지만, 통합 스트림을 근거로 삼으면 “unified-events.jsonl의 47행: browse open 그 뒤에 Network.responseReceived 상태 403이 발생했으며 /api/checkout — 다음으로 전환 --verified --proxies."

예시:

  • "드롭다운을 클릭한 후 1초 대기 — 옵션이 클릭 가능해지기 전에 애니메이션이 표시됩니다"
  • "다음으로 바로 이동 /pay-invoice/ — 랜딩 페이지를 완전히 건너뛰기"
  • " browse fill #field_3 value ‘not’을 사용하세요 browse type — 이 필드는 포커스가 이동하면 내용이 지워집니다"
  • "페이지의 8번 단계에서 로딩 아이콘이 표시됩니다 — browse wait timeout 2000 스냅샷 전에"
  • (다음과 같이 --browser-trace) "unified-events.jsonl의 47행에서, 3개의 연속된 Network.responseReceived 이벤트가 /api/availability 403 오류가 반환되었습니다. browse open — 사이트가 지문 인식(fingerprinting)을 수행 중입니다. 다음 반복에서는 --verified --proxies."

strategy.md 업데이트

편집 ./autobrowse/tasks//strategy.md. 잘 작동했던 부분은 모두 유지하세요. 특정 오류만 수정하세요. 구체적인 휴리스틱을 추가하세요.

좋은 전략은 다음을 갖춥니다:

  • 빠른 경로: 탐색을 건너뛸 수 있는 직접 URL 또는 지름길
  • 단계별 워크플로: 시간 안내가 포함된 정확한 순서
  • 사이트별 지식: 선택기 ID, 양식 필드 이름, 성공 표시기
  • 오류 복구: X가 잘못되었을 때 취해야 할 조치

결과 평가

새로운 요약문을 읽어보세요. 통과했나요? 뚜렷한 진전이 있었나요?

  • 통과 또는 진전 → 유지, 다음 반복
  • 진전이 없거나 퇴보 → strategy.md를 이전 버전으로 되돌리고 다른 가설을 시도해 보세요

실행 가능한 스크립트 생성 (선택 사항)

작업이 수렴되면, 다음을 통해 하나 이상의 프레임워크에서 결정론적이고 실행 가능한 스크립트를 생성할 수 있습니다 scripts/codegen.mjs. 이는 프레임워크당 LLM 호출을 한 번만 수행하며, 콘텐츠 해시별로 캐시되고, 선택적으로 최신 세션과의 검증 및 실패 시 재작성 기능을 지원합니다.

node ${CLAUDE_SKILL_DIR}/scripts/codegen.mjs \
  --task  \
  --workspace ./autobrowse \
  --frameworks playwright,stagehand \
  --verify

각 프레임워크는 tasks/// 각 프레임워크는 생성된 스크립트와 독립적인 스캐폴드(package.json, tsconfig.json)가 포함됩니다. 이 디렉터리는 cd tasks//playwright && npm install && npx tsx .ts — 유일한 실행 환경 요구 사항은 BROWSERBASE_API_KEY (여기에 ANTHROPIC_API_KEY Stagehand 타깃용)입니다.

내장 프레임워크: playwright, stagehand. 사용자 정의 프레임워크를 추가하려면 --prompt-template --frameworks custom (사용자 정의 러너를 제공하거나 다음 플래그를 전달하십시오 --no-verify).

일반적인 플래그:

플래그 용도
--frameworks a,b,... 쉼표로 구분됨; 기본값 playwright
--verify / --no-verify 생성된 스크립트를 새로운 BB 세션에서 실행; 기본값 --verify
--max-retries N 검증 실패 시 재작성 횟수 제한; 기본값 2
--cache-only 캐시 미스 시 오류 발생 (CI 친화적)
--force 캐시 무효화
--dry-run 프롬프트 크기 및 비용 추정; LLM 호출하지 않음
--run 특정 값을 강제 적용 run-NNN (기본값: 최신 통과 결과)

출력은 각 프레임워크당 한 줄의 JSON으로 stdout에 출력됩니다. 선택된 프레임워크 중 하나라도 최종 상태가 passed: false.

참조 references/playwright-cdp-bridge.md 생성된 스크립트가 따르는 표준 connectOverCDP 표준 패턴에 대해서는 여기를 참조하십시오.

모든 반복이 끝난 후 — 준비가 되면 게시

작업이 최근 3회 반복 중 2회 이상 통과했거나 최대 반복 횟수 제한에 도달한 경우, 이를 Claude Code 스킬로 설치하십시오. 단순히 strategy.md를 복사해서는 안 됩니다. 스킬은 독립적으로 작동해야 하며, 이 코드베이스를 한 번도 본 적이 없는 사람에게도 유용해야 합니다. 최대 반복 횟수까지 진행했음에도 완벽히 통과하지 못한 경우, 알려진 실패 지점을 기록하되 배운 모든 내용을 문서화하십시오.

다음 위치에 파일을 작성하여 설치하십시오 ~/.claude/skills//SKILL.md:

mkdir -p ~/.claude/skills/

SKILL.md에는 다음 구조를 사용하십시오:

---
name: 
description: <1-2 sentences describing what this skill does and when to use it. Include trigger keywords.>
---

#  — Browser Skill

## Purpose
<1-2 sentences: what this automates and why it exists.>

## When to Use


## Browse CLI Reference
The inner agent uses the `browse` CLI. Key commands for this task:
- `browse stop` — kill existing session (always run before switching to remote)
- `browse open  --remote` — start a fresh Browserbase cloud session and navigate
- `browse open  --local` — start a clean local browser and navigate
- `browse tab new ` — open URL in a new tab
- `browse wait load` — wait for page to finish loading
- `browse wait timeout ` — wait a fixed amount of time for spinners or animations
- `browse wait selector ""` — wait for an element to become visible
- `browse get title` — verify you're on the right page
- `browse get text body` — extract all visible text (preferred for content extraction)
- `browse snapshot` — get accessibility tree; each node has a ref in `[X-Y]` format (e.g. `[0-5]`, `[2-147]`)
- `browse click [X-Y]` — click element by ref from the latest snapshot (include the brackets)

**Never use `--session ` flags in SKILL.md.** Named sessions are a parallel-run workaround — they contaminate skills with infrastructure concerns. Skills must work in isolation with the default session.

## Workflow

### Step 1 — Start session


### Step 2 — Navigate


### Step 3 — Extract


### Step 4 — Output


## Site-Specific Gotchas


## Failure Recovery


## Expected Output
```json


After writing the SKILL.md, confirm it's installed:
```bash
ls ~/.claude/skills//SKILL.md

이제 이 스킬은 / Claude Code에서 사용할 수 있습니다.

최종 보고서 (멀티태스크 모드)

모든 하위 에이전트가 작업을 완료한 후, 마크다운 테이블을 출력합니다:

작업 반복 횟수 최종 상태 졸업 비용
google-flights 5 ✅ 합격 예 0.42달러
amazon-add-to-cart 5 ❌ 실패 아니요 $1.20

그런 다음 영구 세션 보고서를 ./autobrowse/reports/ 작업 공간 내에 실행 내역이 영구적으로 기록되도록 하려면:

mkdir -p ./autobrowse/reports

다음과 같이 파일을 작성합니다. ./autobrowse/reports/YYYY-MM-DD-HH-MM-.md 다음과 같이 작성하십시오:

# AutoBrowse Session Report
**Date:** 
**Tasks:** 
**Environment:** remote|local
**Total cost:** $X.XX

## Results

| Task | Iterations | Pass Rate | Final Status | Graduated | Cost |
|------|-----------|-----------|--------------|-----------|------|
| ... | ... | X/5 | ✅/❌ | yes/no | $X.XX |

## Per-Task Learnings

### 
- **Key insight 1:** 
- **Key insight 2:** 
- **Failure mode fixed:** 

## Iteration Log

### 
| Iter | Turns | Cost | Status | Hypothesis tested |
|------|-------|------|--------|-------------------|
| 1 | 79 | $18.75 | ❌ fail | baseline |
| 2 | 9 | $0.26 | ✅ pass | session contamination fix |
| ... | ... | ... | ... | ... |

규칙

  • strategy.md 파일만 편집하십시오 — task.md (템플릿에서 생성하는 경우를 제외하고) 또는 evaluate.mjs
  • 작업 공간에 머무르십시오 — 모든 훈련 데이터는 ./autobrowse/로만 저장되며, ~/.claude/skills/autobrowse/로 절대 보내지 마십시오. 스킬 소스는 읽기 전용입니다.
  • 반복마다 하나의 가설만 세우세요 — 한 번에 하나의 변경 사항만 테스트하세요
  • 성공을 기반으로 발전시키세요 — 효과가 있었던 것은 유지하고, 거기에 추가하세요
  • 추적 기록을 신뢰하세요 — 내부 에이전트는 자신이 보고 수행한 내용을 정확히 보여줍니다
  • ~/.claude/skills/로 전환 — 그곳에 작성하는 유일한 파일은 최종 졸업된 파일입니다 SKILL.md
  • 이분법적 분석(bisecting)을 마치기 전에는 릴리스하지 마십시오 — --browser-trace, 각 반복 주기의 마지막 단계 순서는 절대 변경할 수 없습니다: stop-capture → bisect-cdp → browse cloud sessions update REQUEST_RELEASE. 이진 탐색은 추적이 중단될 때 세션이 여전히 존재해야만 가능합니다.
GitHub에서 보기
---
name: autobrowse
description: Builds reliable browser automation skills through iterative experimentation, running an inner agent to browse sites and improving navigation instructions until tasks pass consistently.
license: MIT
---

# AutoBrowse — Self-Improving Browser Skill

Build reliable browser automation skills through iterative experimentation. An inner agent browses the site (`evaluate.ts`). You — the outer agent — read what happened and improve the instructions (`strategy.md`). Repeat until it passes consistently.

## Entry Points

Invocation is flexible — both explicit flags and free-form natural language work:

```
/autobrowse --task google-flights
/autobrowse --task google-flights --iterations 10 --env remote
/autobrowse --task google-flights --browser-trace
/autobrowse --tasks google-flights,amazon-add-to-cart
/autobrowse --all

# Also fine — parse freely:
/autobrowse https://flights.google.com/
/autobrowse book a flight on delta.com
/autobrowse fix the existing google-flights skill
```

`--browser-trace` (default off, remote-only): pairs each iteration with the sibling `browser-trace` skill — wraps the inner agent in a CDP capture for per-page network/console/page-lifecycle evidence. Implies `--env remote`; errors if combined with `--env local`. Requires the sibling `browser-trace` skill present at `${CLAUDE_SKILL_DIR}/../browser-trace/`, and the `BROWSERBASE_API_KEY` env var.

When the user drops a URL or free-form instruction instead of `--task <name>`:
- If an existing task in `${WORKSPACE}/tasks/` clearly matches the site/intent, use it.
- Otherwise, pick a short kebab-case name, create `${WORKSPACE}/tasks/<name>/task.md` from `${CLAUDE_SKILL_DIR}/references/example-task.md`, fill in the URL/goal based on what the user said, and proceed. Tell the user the chosen name in one line.

---

## How to run

### Step 1 — Parse arguments and orient

Check what was passed:
- `--task <name>` → single task mode
- `--tasks a,b,c` or `--all` → multi-task mode (spawn sub-agents)
- `--iterations N` → how many evaluate → improve cycles (default: 5)
- `--env local|remote` → browser environment (default: local; use remote for bot-protected sites)
- `--browser-trace` → opt in to the browser-trace integration (default off). Implies `--env remote`. If `--env local --browser-trace` are both passed explicitly, error with: `browser-trace requires Browserbase; drop --env local or drop --browser-trace.`

If the user passed free-form text instead, map it to one of the above before continuing.

### Step 2 — Set up the workspace

All training artifacts (task definitions, strategy iterations, traces, reports) live in a workspace directory in the **current working directory** — NOT inside `~/.claude/skills/`. This keeps the inner agent's file writes out of Claude's home dir and away from permission friction.

Default workspace: `${CWD}/autobrowse/`

```bash
mkdir -p ./autobrowse/tasks ./autobrowse/traces ./autobrowse/reports
```

If the task directory (`./autobrowse/tasks/<task>/task.md`) doesn't exist yet, scaffold it:

```bash
mkdir -p ./autobrowse/tasks/<task>
cp ${CLAUDE_SKILL_DIR}/references/example-task.md ./autobrowse/tasks/<task>/task.md
# Then edit task.md to describe the URL, inputs, steps, and expected JSON output
```

The skill source at `${CLAUDE_SKILL_DIR}` stays read-only — only `./autobrowse/` in CWD gets written to during training. Graduation (final step) writes a single file to `~/.claude/skills/<task>/SKILL.md`.

List available tasks:
```bash
ls ./autobrowse/tasks/
```

### Step 3 — Multi-task: spawn parallel sub-agents

If running multiple tasks, use the Agent tool to spawn one sub-agent per task simultaneously. Each sub-agent receives a self-contained prompt to run the full autobrowse loop for its task:

> "You are running the autobrowse skill for task `<name>`. Workspace: `<absolute-path-to-workspace>` (e.g. `/path/to/project/autobrowse`). Run `<N>` iterations of: evaluate → read trace → improve strategy.md → repeat. Use `--env <env>`. Pass `--workspace <workspace>` to every evaluate.mjs invocation. If the parent invocation used `--browser-trace`, you MUST use the traced-path block of the SKILL.md loop for every iteration (pre-create session, attach bb-capture, pass `--connect-url` to evaluate.mjs, stop+bisect, release) — do not fall back to the default single-command path. Follow the autobrowse loop instructions exactly.
>
> When graduating, install the skill to `~/.claude/skills/<task-name>/SKILL.md` with proper agentskills frontmatter (name + description). Do not just copy strategy.md — write a self-contained skill.
>
> At the end, output a structured summary with: task name, pass/fail on final run, total cumulative cost, iterations completed, per-iteration table (iter number, turns, cost, status, hypothesis tested), and 2-3 bullet key learnings."

Spawn all sub-agents in parallel, wait for all to complete, then collect their summaries and write the session report.

**For single task**, skip this step and run the loop directly below.

---

## The Loop (run this for each task)

### Iteration start

Check that `./autobrowse/tasks/<task>/task.md` exists (scaffold it from the template if not — see Step 2). `strategy.md` is auto-created empty by the harness on first run.

### Requirements

- `ANTHROPIC_API_KEY` must be in the environment (or in a `.env` file in CWD — `evaluate.mjs` auto-loads it). If missing, the harness prints a clear error and exits; don't hunt for keys in other paths.

### Run the inner agent

**Default path (no `--browser-trace`)** — single command, no orchestration:

```bash
node ${CLAUDE_SKILL_DIR}/scripts/evaluate.mjs --task <task-name> --workspace ./autobrowse
# or for bot-protected sites:
node ${CLAUDE_SKILL_DIR}/scripts/evaluate.mjs --task <task-name> --workspace ./autobrowse --env remote
```

This runs the browser session and writes a full trace to `./autobrowse/traces/<task>/latest/`.

**Traced path (`--browser-trace`, remote only)** — the outer harness pre-creates a Browserbase session, attaches `bb-capture` as a passive observer, and passes the session's `connectUrl` to `evaluate.mjs` so every inner `browse` call uses `--cdp $connectUrl --session autobrowse-main` (the canonical browser-trace pattern that gives observers full Network/Console events). Run this block once per iteration with `$N` set to the 1-indexed iteration number:

```bash
# Preflight — fail fast if browser-trace isn't installed alongside autobrowse.
BT_DIR="${CLAUDE_SKILL_DIR}/../browser-trace"
if [ ! -f "$BT_DIR/scripts/bb-capture.mjs" ]; then
  echo "ERROR: --browser-trace requires the browser-trace skill at $BT_DIR." >&2
  echo "Install it by cloning github.com/browserbase/skills and copying skills/browser-trace/" >&2
  echo "into the same parent directory as autobrowse (e.g. ~/.claude/skills/browser-trace/)." >&2
  exit 1
fi

# a. SESSION SETUP — pre-create the keep-alive session and derive its connectUrl
sid=$(browse cloud sessions create --keep-alive --verified --proxies \
  | node -e "let s='';process.stdin.on('data',c=>s+=c).on('end',()=>process.stdout.write(JSON.parse(s).id))")
connect_url=$(browse cloud sessions get "$sid" \
  | node -e "let s='';process.stdin.on('data',c=>s+=c).on('end',()=>process.stdout.write(JSON.parse(s).connectUrl))")

RUN_ID="run-$(printf '%03d' "$N")"
TRACE_ROOT="./autobrowse/traces/<task-name>/$RUN_ID"
mkdir -p "$TRACE_ROOT"
export O11Y_ROOT="$TRACE_ROOT/.o11y"   # park browser-trace output inside the autobrowse run dir
export O11Y_RUN_ID="$RUN_ID"           # tells the browse CLI which run dir to write descriptors.ndjson into

# b. ATTACH BROWSER-TRACE — passive observer; runs in background
node ${CLAUDE_SKILL_DIR}/../browser-trace/scripts/bb-capture.mjs "$sid" "$RUN_ID" &
sleep 2

# c. RUN AUTOBROWSE — connectUrl flag tells evaluate.mjs to inject --cdp/--session
#    into every inner browse call. The inner agent never sees --remote.
node ${CLAUDE_SKILL_DIR}/scripts/evaluate.mjs \
  --task <task-name> --workspace ./autobrowse --env remote \
  --connect-url "$connect_url" --run-number "$N"

# d. STOP + BISECT + UNIFY — order matters; bisect needs the session to still
#    exist, and unify-trace joins the bisect output with autobrowse's trace.json
#    into a single time-ordered NDJSON the outer agent reads first each iter.
node ${CLAUDE_SKILL_DIR}/../browser-trace/scripts/stop-capture.mjs "$RUN_ID"
node ${CLAUDE_SKILL_DIR}/../browser-trace/scripts/bisect-cdp.mjs "$RUN_ID"
node ${CLAUDE_SKILL_DIR}/scripts/unify-trace.mjs \
  --trace-dir "$TRACE_ROOT" \
  --o11y-dir "$O11Y_ROOT/$RUN_ID"

# e. RELEASE
browse cloud sessions update "$sid" --status REQUEST_RELEASE
```

This writes the inner-agent trace to `./autobrowse/traces/<task-name>/latest/` and the CDP bisect to `./autobrowse/traces/<task-name>/latest/.o11y/<run-id>/`. The traced `browse` CLI also emits per-command rich node descriptors to `.o11y/<run-id>/cdp/descriptors.ndjson` (one JSON object per page-driving call: target tag/id/role/accessibleName/attributes/xpath/bounding-rect). The descriptors file feeds downstream codegen; it is **not** required for hypothesis formation — skip it when reading the trace.

### Read the trace

```bash
cat ./autobrowse/traces/<task-name>/latest/summary.md
```

The summary has duration, cost, turns, the decision log, and the final JSON output.

If the agent failed or got stuck, look deeper:
- Read `./autobrowse/traces/<task-name>/latest/trace.json` — search for the failure turn
- Read screenshots around the failure point with the Read tool

**When `--browser-trace` was used — start with `unified-events.jsonl`.** The harness joins the agent's turn log and the browser's CDP firehose into one time-ordered NDJSON stream at the run root. One file, source-tagged (`source: "agent" | "browser"`), interleaved by wall-clock timestamp. Skim it top-to-bottom; the failure cause is usually one or two adjacent lines (the agent issued command X, the browser responded with Y).

```bash
cat ./autobrowse/traces/<task-name>/latest/unified-events.jsonl
```

The structured files (`trace.json`, `.o11y/<run-id>/cdp/*`) are **also agent-consumable as drill-downs** when the unified stream points at something you need more of:

| Need | Drill-down file or command |
|---|---|
| Per-page totals + timing (events, network counts, errors by page) | `.o11y/<run-id>/cdp/summary.json` |
| All failed network requests in one place | `.o11y/<run-id>/cdp/network/failed.jsonl` |
| Full console exception payloads (stacktraces, etc.) | `.o11y/<run-id>/cdp/console/exceptions.jsonl` |
| Per-page slice (only events on page N) | `.o11y/<run-id>/cdp/pages/<pid>/` |
| Full reasoning text / untruncated tool outputs for a specific turn | `trace.json` (filter by `turn === N`) |
| Ad-hoc grouped query (e.g. top hosts, errors-by-page) | `O11Y_ROOT=./autobrowse/traces/<task-name>/latest/.o11y node ${CLAUDE_SKILL_DIR}/../browser-trace/scripts/query.mjs <run-id> <cmd>` |

The unified stream is the default; drill into structured files only when you need a grouped query, a full-text payload, or filtering the stream can't give you.

### Form one hypothesis

Find the exact turn where things went wrong. What single heuristic would have prevented it?

Under `--browser-trace`, the hypothesis must cite a **specific event from `unified-events.jsonl`** (line number or timestamp) — or name the drill-down file if you had to descend into one. This keeps updates evidence-grounded rather than vibes-driven. A hypothesis based only on the agent's commands might say "the click didn't work"; grounded in the unified stream, it can say "line 47 of unified-events.jsonl: `browse open` was followed by `Network.responseReceived` status 403 on `/api/checkout` — switch to `--verified --proxies`."

Examples:
- "After clicking the dropdown, wait 1s — options animate in before they're clickable"
- "Navigate directly to `/pay-invoice/` — skip the landing page entirely"
- "Use `browse fill #field_3 value` not `browse type` — this field clears on focus"
- "The page shows a spinner at turn 8 — add `browse wait timeout 2000` before snapshot"
- (with `--browser-trace`) "At line 47 of unified-events.jsonl, 3 consecutive `Network.responseReceived` events on `/api/availability` returned 403 right after `browse open` — the site is fingerprinting; the next iter needs `--verified --proxies`."

### Update strategy.md

Edit `./autobrowse/tasks/<task-name>/strategy.md`. Keep everything that worked. Fix the specific failure. Add a concrete heuristic.

Good strategies have:
- **Fast path**: direct URL or shortcuts to skip exploration
- **Step-by-step workflow**: exact sequence with timing notes
- **Site-specific knowledge**: selector IDs, form field names, success indicators
- **Failure recovery**: what to do when X goes wrong

### Judge the result

Read the new summary. Did it pass? Make clear progress?
- **Pass or progress** → keep, next iteration
- **No progress or regression** → revert strategy.md to the previous version and try a different hypothesis

### Generate a runnable script (optional)

Once the task has converged, you can produce a deterministic, runnable script
in one or more frameworks via `scripts/codegen.mjs`. This is one shot of an
LLM call per framework, cached by content hash, with optional verify-against-
fresh-session and rewrite-on-failure.

```bash
node ${CLAUDE_SKILL_DIR}/scripts/codegen.mjs \
  --task <name> \
  --workspace ./autobrowse \
  --frameworks playwright,stagehand \
  --verify
```

Each framework gets its own subdirectory under `tasks/<name>/<framework>/`
with the emitted script and a self-contained scaffold (`package.json`,
`tsconfig.json`). The directory is runnable standalone with
`cd tasks/<name>/playwright && npm install && npx tsx <name>.ts` — the only
runtime requirement is `BROWSERBASE_API_KEY` (plus `ANTHROPIC_API_KEY` for
the Stagehand target).

Builtin frameworks: `playwright`, `stagehand`. Add a custom framework with
`--prompt-template <path> --frameworks custom` (and provide your own runner
or pass `--no-verify`).

Common flags:

| Flag | Purpose |
|---|---|
| `--frameworks a,b,...` | Comma-separated; default `playwright` |
| `--verify` / `--no-verify` | Run the produced script against a fresh BB session; default `--verify` |
| `--max-retries N` | Rewrite-on-verify-failure cap; default 2 |
| `--cache-only` | Error if cache miss (CI-friendly) |
| `--force` | Bust the cache |
| `--dry-run` | Estimate prompt size + cost; don't call the LLM |
| `--run <id>` | Force a specific `run-NNN` (default: latest passing) |

Output is one JSON line per framework on stdout. Non-zero exit if any
selected framework's final state is `passed: false`.

See `references/playwright-cdp-bridge.md` for the canonical
`connectOverCDP` patterns the emitted scripts follow.

### After all iterations — publish if ready

If the task passed on 2+ of the last 3 iterations **or has reached the max iteration limit**, install it as a Claude Code skill. **Do not just copy strategy.md** — the skill must be self-contained and useful to someone who has never seen this codebase. If graduating at max iterations without a clean pass, note the known failure point but still document everything learned.

Install by writing to `~/.claude/skills/<task-name>/SKILL.md`:

```bash
mkdir -p ~/.claude/skills/<task-name>
```

Use this structure for the SKILL.md:

```markdown
---
name: <task-name>
description: <1-2 sentences describing what this skill does and when to use it. Include trigger keywords.>
---

# <Task Title> — Browser Skill

## Purpose
<1-2 sentences: what this automates and why it exists.>

## When to Use
<When should someone reach for this skill.>

## Browse CLI Reference
The inner agent uses the `browse` CLI. Key commands for this task:
- `browse stop` — kill existing session (always run before switching to remote)
- `browse open <url> --remote` — start a fresh Browserbase cloud session and navigate
- `browse open <url> --local` — start a clean local browser and navigate
- `browse tab new <url>` — open URL in a new tab
- `browse wait load` — wait for page to finish loading
- `browse wait timeout <ms>` — wait a fixed amount of time for spinners or animations
- `browse wait selector "<selector>"` — wait for an element to become visible
- `browse get title` — verify you're on the right page
- `browse get text body` — extract all visible text (preferred for content extraction)
- `browse snapshot` — get accessibility tree; each node has a ref in `[X-Y]` format (e.g. `[0-5]`, `[2-147]`)
- `browse click [X-Y]` — click element by ref from the latest snapshot (include the brackets)

**Never use `--session <name>` flags in SKILL.md.** Named sessions are a parallel-run workaround — they contaminate skills with infrastructure concerns. Skills must work in isolation with the default session.

## Workflow

### Step 1 — Start session
<exact browse commands in order>

### Step 2 — Navigate
<exact URL and verification steps>

### Step 3 — Extract
<exact extraction commands>

### Step 4 — Output
<what JSON to emit, referencing the schema below>

## Site-Specific Gotchas
<Bullet list of every hard-won heuristic from the iterations. This is the core value of the skill.>

## Failure Recovery
<What to do when navigation fails, session is contaminated, or extraction returns garbage>

## Expected Output
```json
<paste the exact expected output schema from task.md>
```
```

After writing the SKILL.md, confirm it's installed:
```bash
ls ~/.claude/skills/<task-name>/SKILL.md
```

The skill is now available as `/<task-name>` in Claude Code.

---

## Final report (multi-task mode)

After all sub-agents complete, print a markdown table:

| Task | Iterations | Final Status | Graduated | Cost |
|------|-----------|--------------|-----------|------|
| google-flights | 5 | ✅ pass | yes | $0.42 |
| amazon-add-to-cart | 5 | ❌ fail | no | $1.20 |

Then write a persistent session report to `./autobrowse/reports/` so there's a durable record of the run inside the workspace:

```bash
mkdir -p ./autobrowse/reports
```

Write the file `./autobrowse/reports/YYYY-MM-DD-HH-MM-<tasks>.md` with:

```markdown
# AutoBrowse Session Report
**Date:** <ISO date>
**Tasks:** <comma-separated list>
**Environment:** remote|local
**Total cost:** $X.XX

## Results

| Task | Iterations | Pass Rate | Final Status | Graduated | Cost |
|------|-----------|-----------|--------------|-----------|------|
| ... | ... | X/5 | ✅/❌ | yes/no | $X.XX |

## Per-Task Learnings

### <task-name>
- **Key insight 1:** <what the agent learned>
- **Key insight 2:** <another heuristic>
- **Failure mode fixed:** <what was failing and how it was resolved>

## Iteration Log

### <task-name>
| Iter | Turns | Cost | Status | Hypothesis tested |
|------|-------|------|--------|-------------------|
| 1 | 79 | $18.75 | ❌ fail | baseline |
| 2 | 9 | $0.26 | ✅ pass | session contamination fix |
| ... | ... | ... | ... | ... |
```

---

## Rules

- **Only edit `strategy.md`** — never touch `task.md` (unless creating it from the template) or `evaluate.mjs`
- **Stay in the workspace** — all training writes go to `./autobrowse/`, never to `~/.claude/skills/autobrowse/`. The skill source is read-only.
- **One hypothesis per iteration** — test one change at a time
- **Build on wins** — keep what worked, add to it
- **Trust the trace** — the inner agent shows exactly what it saw and did
- **Graduate to `~/.claude/skills/`** — the only file you write there is the final graduated `SKILL.md`
- **Don't release before bisecting** — under `--browser-trace`, the order at the end of each iteration is non-negotiable: `stop-capture` → `bisect-cdp` → `browse cloud sessions update REQUEST_RELEASE`. Bisect depends on the session still existing when the trace stops.

autobrowse 설치

스킬 파일을 다운로드하여 .claude/skills/ 디렉터리에 압축을 풀어주세요.

ZIP 다운로드

저장소를 클론하고 스킬 파일을 프로젝트에 복사하세요.

git clone https://github.com/browserbase/skills/tree/main/skills/autobrowse # Copy SKILL.md to your .claude/skills/ directory

복사 복사
빠른 설정: 스킬 폴더를 .claude/skills/로 복사하세요. Claude가 해당 스킬을 자동으로 감지하여 사용할 것입니다.
저장소 browserbase/skills

관련 스킬

klingai-upgrade-migration
업데이트 된 시간 2026년 7월 3일
Verification &amp; Quality Assurance
업데이트 된 시간 2026년 6월 29일
base44-cli
업데이트 된 시간 2026년 6월 29일
Railway CLI Management
업데이트 된 시간 2026년 7월 2일
OR