选项

通过迭代实验,运行内部代理来浏览网站,并不断优化导航指令,直至任务能够稳定通过,从而掌握可靠的浏览器自动化技能。

...展开全部
0
更新时间 2026-09-30

AutoBrowse — 自我提升的浏览器自动化技能

通过迭代实验,培养可靠的浏览器自动化技能。一个内部代理负责浏览网站(evaluate.ts)。而你——作为外部代理——则分析发生的情况,并优化操作指令(strategy.md)。重复这一过程,直到系统能够稳定通过测试。

入口点

调用方式灵活——既支持显式参数,也支持自由形式的自然语言:

/autobrowse --task google-flights
/autobrowse --task google-flights --iterations 10 --env remote
/autobrowse --task google-flights --browser-trace
/autobrowse --tasks google-flights,amazon-add-to-cart
/autobrowse --all

# 以下方式同样有效——可自由解析:
/autobrowse https://flights.google.com/
/autobrowse 在 delta.com 上预订机票
/autobrowse 修复现有的 google-flights 技能

--browser-trace(默认关闭,仅限远程模式):将每次迭代与同级的browser-trace技能配对——将内部代理封装在 CDP 捕获中,以获取每页的网络/控制台/页面生命周期证据。 隐含--env remote;若与--env local 结合使用将引发错误。要求同级browser-trace技能位于${CLAUDE_SKILL_DIR}/../browser-trace/ 目录下,且需设置BROWSERBASE_API_KEY环境变量。

当用户输入 URL 或自由格式指令而非--task 时:

  • 如果${WORKSPACE}/tasks/中已有任务与该网站/意图明显匹配,则使用该任务。
  • 否则,选择一个简短的鞑靼式命名,根据${CLAUDE_SKILL_DIR}/references/example-task.md 生成${WORKSPACE}/tasks//task.md,根据用户所述内容填写 URL/目标,然后继续执行。 用一句话告知用户所选的名称。

运行方法

步骤 1 — 解析参数并确定方向

检查传入的参数:

  • --task→ 单任务模式
  • --tasks a,b,c或--all→ 多任务模式(启动子代理)
  • --iterations N→ 评估 → 优化循环的次数(默认:5)
  • --env local|remote→ 浏览器环境(默认:local;访问受机器人防护的网站时使用 remote)
  • --browser-trace→ 启用浏览器跟踪集成(默认关闭)。 该选项默认包含--env remote。若同时显式指定了--env local 和 --browser-trace,将报错提示:browser-trace 需要 Browserbase;请移除 --env local 或移除 --browser-trace。

如果用户传入的是自由格式文本,请将其映射到上述选项之一后再继续。

步骤 2 — 设置工作区

所有训练成果(任务定义、策略迭代、跟踪记录、报告)均存储在当前工作目录中的工作区目录内——而非~/.claude/skills/ 目录内。此举可确保内部代理的文件写入操作不进入 Claude 的主目录,从而避免权限冲突。

默认工作区:${CWD}/autobrowse/

mkdir -p ./autobrowse/tasks ./autobrowse/traces ./autobrowse/reports

如果任务目录(./autobrowse/tasks//task.md)尚未存在,请创建该目录:

mkdir -p ./autobrowse/tasks/
cp ${CLAUDE_SKILL_DIR}/references/example-task.md ./autobrowse/tasks//task.md
# 然后编辑 task.md 文件,描述 URL、输入、步骤以及预期的 JSON 输出

位于${CLAUDE_SKILL_DIR}的技能源文件保持只读状态——训练过程中仅向当前工作目录(CWD)下的./autobrowse/目录写入数据。毕业(最后一步)会向~/.claude/skills//SKILL.md 写入一个文件。

列出可用任务:

ls ./autobrowse/tasks/

步骤 3 — 多任务:并行启动子代理

若运行多个任务,请使用 Agent 工具为每个任务同时启动一个子代理。每个子代理会收到一个自包含的提示,用于执行其任务的完整autobrowse 循环:

autobrowse “您正在为任务 。工作区: (例如/path/to/project/autobrowse)。执行 以下循环:评估 → 读取跟踪信息 → 改进 strategy.md → 重复。 请使用--env 。向每次 evaluate.mjs 调用传递--workspace。如果父调用使用了--browser-trace,您必须在每次迭代中使用 SKILL.md 循环中的 traced-path 代码块(预创建会话、附加 bb-capture、将--connect-url参数传递给 evaluate.mjs、停止+二分法排查、发布)——切勿回退到默认的单命令路径。请严格遵循autobrowse 中的循环操作指南。

毕业时,请将技能安装到~/.claude/skills//SKILL.md,并确保 frontmatter 中包含正确的 agentskills 信息(名称 + 描述)。请勿仅复制 strategy.md —— 应编写一个自包含的技能。

最后,输出一份结构化总结,包含:任务名称、最终运行结果(通过/失败)、累计总成本、已完成的迭代次数、每轮迭代表格(迭代编号、回合数、成本、状态、测试的假设),以及2-3条关键收获的要点。"

并行启动所有子代理,等待全部完成后,收集其摘要并撰写会话报告。

对于单个任务,请跳过此步骤,直接运行下方的循环。

循环(针对每个任务运行)

迭代开始

检查./autobrowse/tasks//task.md是否存在(若不存在,请根据模板生成——参见步骤 2)。strategy.md文件会在首次运行时由测试框架自动生成(初始为空)。

要求

  • 环境中必须存在ANTHROPIC_API_KEY(或位于当前工作目录下的.env 文件中——evaluate.mjs会自动加载该文件)。若缺失,测试框架将输出明确的错误信息并退出;请勿在其他路径中查找该密钥。

运行内部代理

默认路径(不带--browser-trace参数)——单条命令,无需编排:

node ${CLAUDE_SKILL_DIR}/scripts/evaluate.mjs --task --workspace ./autobrowse
# 或针对受机器人保护的网站:
node ${CLAUDE_SKILL_DIR}/scripts/evaluate.mjs --task --workspace ./autobrowse --env remote

这将启动浏览器会话,并将完整的跟踪记录写入./autobrowse/traces//latest/。

跟踪路径(--browser-trace,仅限远程)——外部测试框架会预先创建一个 Browserbase 会话,将bb-capture作为被动观察者附加,并将该会话的connectUrl传递给evaluate.mjs,以便每个内部browse调用都使用--cdp $connectUrl --sessionautobrowse-main(标准 browser-trace 模式,可为观察者提供完整的 Network/Console 事件)。每轮迭代执行一次此代码块,其中$N设置为从 1 开始计数的迭代编号:

# 预检 — 若 browser-trace 未与autobrowse 一同安装,则快速报错。
BT_DIR="${CLAUDE_SKILL_DIR}/../browser-trace"
if [ ! -f "$BT_DIR/scripts/bb-capture.mjs" ]; then
  echo "错误:--browser-trace 需要在 $BT_DIR 目录下安装 browser-trace 技能。" >&2
  echo "请克隆 github.com/browserbase/skills,并将 skills/browser-trace/" >&2
  echo "复制到与autobrowse 相同的父目录下(例如 ~/.claude/skills/browser-trace/)。" >&2
  exit 1
fi

# a. 会话设置 — 预先创建保持活动状态的会话并推导其 connectUrl
sid=$(browse cloud sessions create --keep-alive --verified --proxies \
  | node -e "let s='';process.stdin.on('data',c=>s+=c).on('end',()=>process.stdout.write(JSON.parse(s).id))")
connect_url=$(browse cloud sessions get "$sid" \
  | node -e "let s='';process.stdin.on('data',c=>s+=c).on('end',()=>process.stdout.write(JSON.parse(s).connectUrl))")

RUN_ID="run-$(printf '%03d' "$N")"
TRACE_ROOT="./autobrowse/traces//$RUN_ID"
mkdir -p "$TRACE_ROOT"
export O11Y_ROOT="$TRACE_ROOT/.o11y"   # 将浏览器跟踪输出保存在autobrowse 运行目录中
export O11Y_RUN_ID="$RUN_ID"           # 告知 browse CLI 应将 descriptors.ndjson 写入哪个运行目录

# b. 附加浏览器跟踪 — 被动观察者;在后台运行
node ${CLAUDE_SKILL_DIR}/../browser-trace/scripts/bb-capture.mjs "$sid" "$RUN_ID" &
sleep 2

# c. 运行AUTOBROWSE — connectUrl 参数指示 evaluate.mjs 将 --cdp/--session
#    注入到每个内部 browse 调用中。内部代理永远不会看到 --remote。
node ${CLAUDE_SKILL_DIR}/scripts/evaluate.mjs \
  --task --workspace ./autobrowse --env remote \
  --connect-url "$connect_url" --run-number "$N"

# d. STOP + BISECT + UNIFY — 执行顺序至关重要;bisect 需要会话仍
#    存在,而 unify-trace 会将 bisect 的输出与autobrowse 的 trace.json 合并
#    合并为一个按时间排序的 NDJSON,外部代理在每次迭代中首先读取该文件。
node ${CLAUDE_SKILL_DIR}/../browser-trace/scripts/stop-capture.mjs "$RUN_ID"
node ${CLAUDE_SKILL_DIR}/../browser-trace/scripts/bisect-cdp.mjs "$RUN_ID"
node ${CLAUDE_SKILL_DIR}/scripts/unify-trace.mjs \
  --trace-dir "$TRACE_ROOT" \
  --o11y-dir "$O11Y_ROOT/$RUN_ID"

# e. 发布
browse cloud sessions update "$sid" --status REQUEST_RELEASE

这会将代理内部跟踪写入./autobrowse/traces//latest/,并将 CDP 二分法写入./autobrowse/traces//latest/.o11y//。 被追踪的浏览CLI 还会向.o11y//cdp/descriptors.ndjson输出按命令划分的丰富节点描述符(每个页面驱动调用对应一个 JSON 对象:target tag/id/role/accessibleName/attributes/xpath/bounding-rect)。 该描述符文件用于为下游代码生成提供数据;它并非假设形成所必需的——在阅读跟踪信息时可跳过该部分。

读取跟踪记录

cat ./autobrowse/traces//latest/summary.md

摘要包含持续时间、成本、迭代次数、决策日志以及最终的 JSON 输出。

如果智能体失败或卡住,请进一步检查:

  • 阅读./autobrowse/traces//latest/trace.json— 查找失败的回合
  • 使用 Read 工具查看故障点附近的截图

当使用了--browser-trace选项时——请从unified-events.jsonl 文件开始分析。测试框架会将代理的回合日志和浏览器的 CDP 数据流合并为一个按时间排序的 NDJSON 流,并存储在运行根目录下。 这是一个带有源标签(source: "agent" | "browser")的文件,各数据按墙钟时间戳交错排列。请从上到下快速浏览;故障原因通常出现在一两行相邻的记录中(代理发出命令 X,浏览器响应为 Y)。

cat ./autobrowse/traces//latest/unified-events.jsonl

当统一数据流指向您需要进一步深入分析的内容时,结构化文件(trace.json、.o11y//cdp/*)也可作为代理的钻取数据源:

需求 钻取文件或命令
每页总计 + 时序数据(事件、网络计数、按页面划分的错误) .o11y//cdp/summary.json
所有失败的网络请求汇总 .o11y//cdp/network/failed.jsonl
完整的控制台异常有效载荷(堆栈跟踪等) .o11y//cdp/console/exceptions.jsonl
按页面切片(仅包含第 N 页上的事件) .o11y//cdp/pages//
特定回合的完整推理文本/未截断的工具输出 trace.json(按轮次 === N 过滤)
临时分组查询(例如:顶级主机、按页面分类的错误) O11Y_ROOT=./autobrowse/traces//latest/.o11y 节点 ${CLAUDE_SKILL_DIR}/../browser-trace/scripts/query.mjs

统一数据流是默认选项;仅当您需要分组查询、全文有效载荷,或数据流无法满足的过滤需求时,才需深入到结构化文件中。

提出一个假设

找出问题确切发生的转折点。如果采用哪一项单一启发式方法,本可以避免这一问题?

在--browser-trace 模式下,假设必须引用 unified-events.jsonl中的具体事件(行号或时间戳)——若需深入分析,则需指名相应的钻取文件。这样能确保更新基于确凿证据,而非凭主观感觉。 仅基于用户操作的假设可能表述为“点击未生效”; 若基于统一事件流,则可表述为“unified-events.jsonl 的第 47 行:browse open之后,/api/checkout出现Network.responseReceived状态码 403 —— 请切换至--verified --proxies 模式”。

示例:

  • “点击下拉菜单后等待 1 秒——选项在可点击前会先以动画形式显示”
  • “直接导航至/pay-invoice/— 完全跳过登录页面”
  • “使用‘填入’操作设置 #field_3 的值,而非‘浏览’类型——该字段在获得焦点时会清空”
  • “页面在第 8 步显示加载图标 — 在快照前为浏览操作添加2000 毫秒的等待超时”
  • (使用--browser-trace 选项) “在 unified-events.jsonl 的第 47 行,/api/availability上的 3 个连续Network.responseReceived事件在打开浏览器窗口后立即返回 403 状态码——该网站正在进行指纹识别;下一次迭代需使用--verified --proxies 参数。”

更新 strategy.md

编辑./autobrowse/tasks//strategy.md。保留所有已验证有效的内容。修复具体故障。添加具体的启发式规则。

优秀的策略应具备:

  • 快速路径:直接 URL 或可跳过探索的快捷方式
  • 分步工作流:带有时间注释的精确操作序列
  • 特定网站知识:选择器 ID、表单字段名称、成功指示符
  • 故障恢复:当 X 出错时该如何处理

评估结果

阅读新摘要。是否通过?是否取得明显进展?

  • 通过或有进展→ 保留,进入下一次迭代
  • 无进展或出现退步→ 将 strategy.md 还原至上一版本,并尝试其他假设

生成可运行脚本(可选)

一旦任务收敛,您可以通过scripts/codegen.mjs 在一种或多种框架中生成一个确定性的、可运行的脚本。这是针对每个框架的 单次 LLM 调用,按内容哈希进行缓存,并可选地支持 “与最新会话进行验证”和“失败时重写”功能。

node ${CLAUDE_SKILL_DIR}/scripts/codegen.mjs \
  --task \
  --workspace ./autobrowse \
  --frameworks playwright,stagehand \
  --verify

每个框架都在tasks/、/ 和/下拥有自己的子目录, 其中包含生成的脚本和一个自包含的框架(package.json、 tsconfig.json)。 该目录可通过以下命令独立运行: cd tasks//playwright && npm install && npx tsx.ts——唯一的 运行时要求是BROWSERBASE_API_KEY(针对 Stagehand 目标还需 ANTHROPIC_API_KEY)。

内置框架:playwright、stagehand。使用 --prompt-template --frameworks custom添加自定义框架(并提供您自己的运行器 或传入--no-verify)。

常用标志:

参数 用途
--frameworks a,b,... 以逗号分隔;默认值为playwright
--verify/--no-verify 在全新的 BB 会话中运行生成的脚本;默认值为--verify
--max-retries N 验证失败时的重写次数上限;默认值为 2
--仅缓存 若缓存未命中则报错(适合持续集成)
--force 清空缓存
--dry-run 估算提示词长度及计算成本;不调用大型语言模型
--run 强制使用特定的运行-NNN(默认:最新通过的)

输出为每个框架一行 JSON 数据,输出至标准输出。若 任何选定框架的最终状态为“false”,则返回非零退出状态。

请参阅references/playwright-cdp-bridge.md,了解生成的脚本遵循的规范 connectOverCDP模式。

所有迭代结束后——若已准备就绪则发布

如果任务在最近 3 次迭代中有 2 次及以上通过,或已达到最大迭代次数限制,则将其安装为 Claude Code 技能。请勿简单复制 strategy.md—— 该技能必须自成体系,并对从未接触过此代码库的人具有实用价值。若在达到最大迭代次数时仍未完全通过测试,请记录已知的失败点,但仍需完整记录所学内容。

通过在~/.claude/skills/ 目录下编写/SKILL.md 文件进行安装:

mkdir -p ~/.claude/skills/

SKILL.md 文件请采用以下结构:

---
name:
description:<1-2 sentences describing what this skill does and when to use it. Include trigger keywords.>
---

# — 浏览器技能

## 目的
<1-2 sentences: what this automates and why it exists.>

## 适用场景


## 浏览 CLI 参考文档
内部代理使用 `browse` CLI。此任务的关键命令:
- `browse stop` — 终止现有会话(切换到远程模式前务必执行)
- `browse open --remote` — 启动一个全新的 Browserbase 云会话并进行浏览
- `browse open --local` — 启动一个干净的本地浏览器并进行浏览
- `browse tab new` — 在新标签页中打开 URL
- `browse wait load` — 等待页面加载完成
- `browse wait timeout` — 等待固定时间,以处理加载转圈图标或动画效果
- `browse wait selector ""` — 等待某个元素变得可见
- `browse get title` — 验证当前是否处于正确页面
- `browse get text body` — 提取所有可见文本(推荐用于内容提取)
- `browse snapshot` — 获取辅助功能树;每个节点都有一个格式为 `[X-Y]` 的引用(例如 `[0-5]`、`[2-147]`)
- `browse click [X-Y]` — 根据最新快照中的引用点击元素 (包括方括号)

**切勿在 SKILL.md 中使用 `--session` 参数。** 命名会话是一种并行运行的变通方案——它们会将基础设施问题引入技能中。技能必须在默认会话下独立运行。

## 工作流程

### 步骤 1 — 启动会话


### 步骤 2 — 导航


### 步骤 3 — 提取


### 步骤 4 — 输出


## 特定于网站的注意事项


## 故障恢复


## 预期输出
```json


编写完 SKILL.md 后,请确认其已安装:
```bash
ls ~/.claude/skills//SKILL.md

该技能现已在 Claude Code 中以/的形式提供。

最终报告(多任务模式)

当所有子代理任务完成后,输出一个 Markdown 表格:

任务 迭代次数 最终状态 已毕业 成本
google-flights 5 ✅ 通过 是 0.42美元
亚马逊-加入购物车 5 ❌ 失败 否 $1.20

然后将持久化会话报告写入./autobrowse/reports/,以便在工作区中保留本次运行的持久记录:

mkdir -p ./autobrowse/reports

创建文件./autobrowse/reports/YYYY-MM-DD-HH-MM-.md,内容如下:

#AutoBrowse 会话报告
**日期:**
**任务:**
**环境:** 远程|本地
**总成本:** $X.XX

## 结果

| 任务 | 迭代次数 | 通过率 | 最终状态 | 是否通过 | 成本 |
|------|-----------|-----------|--------------|-----------|------|
| ... | ... | X/5 | ✅/❌ | 是/否 | $X.XX |

## 各任务经验总结

###
- **关键见解 1:**
- **关键见解 2:**
- **已修复的故障模式:**

## 迭代日志

###
| 迭代 | 轮次 | 成本 | 状态 | 验证的假设 |
|------|-------|------|--------|-------------------|
| 1 | 79 | $18.75 | ❌ 失败 | 基线 |
| 2 | 9 | $0.26 | ✅ 通过 | 会话污染修复 |
| ... | ... | ... | ... | ... |

规则

  • 仅编辑strategy.md—— 切勿修改task.md(除非是从模板创建)或evaluate.mjs
  • 请留在工作区内— 所有训练写入操作均应指向./autobrowse/,切勿写入~/.claude/skills/autobrowse/。技能源文件为只读。
  • 每次迭代仅测试一个假设— 每次只测试一项更改
  • 基于成功经验进行扩展——保留行之有效的方法,并在其基础上进行补充
  • 信任追踪记录——内部代理会精确展示其所见所为
  • 迁移至~/.claude/skills/—— 该目录下唯一可写入的文件是最终通过验证的SKILL.md
  • 在进行二分法排查前不要发布—— 在--browser-trace 模式下,每次迭代结束时的操作顺序是不可更改的:stop-capture→bisect-cdp→浏览云会话并更新 REQUEST_RELEASE。二分法排查依赖于在追踪停止时会话仍然存在。
在 GitHub 上查看
---
name: autobrowse
description: Builds reliable browser automation skills through iterative experimentation, running an inner agent to browse sites and improving navigation instructions until tasks pass consistently.
license: MIT
---

# AutoBrowse — Self-Improving Browser Skill

Build reliable browser automation skills through iterative experimentation. An inner agent browses the site (`evaluate.ts`). You — the outer agent — read what happened and improve the instructions (`strategy.md`). Repeat until it passes consistently.

## Entry Points

Invocation is flexible — both explicit flags and free-form natural language work:

```
/autobrowse --task google-flights
/autobrowse --task google-flights --iterations 10 --env remote
/autobrowse --task google-flights --browser-trace
/autobrowse --tasks google-flights,amazon-add-to-cart
/autobrowse --all

# Also fine — parse freely:
/autobrowse https://flights.google.com/
/autobrowse book a flight on delta.com
/autobrowse fix the existing google-flights skill
```

`--browser-trace` (default off, remote-only): pairs each iteration with the sibling `browser-trace` skill — wraps the inner agent in a CDP capture for per-page network/console/page-lifecycle evidence. Implies `--env remote`; errors if combined with `--env local`. Requires the sibling `browser-trace` skill present at `${CLAUDE_SKILL_DIR}/../browser-trace/`, and the `BROWSERBASE_API_KEY` env var.

When the user drops a URL or free-form instruction instead of `--task <name>`:
- If an existing task in `${WORKSPACE}/tasks/` clearly matches the site/intent, use it.
- Otherwise, pick a short kebab-case name, create `${WORKSPACE}/tasks/<name>/task.md` from `${CLAUDE_SKILL_DIR}/references/example-task.md`, fill in the URL/goal based on what the user said, and proceed. Tell the user the chosen name in one line.

---

## How to run

### Step 1 — Parse arguments and orient

Check what was passed:
- `--task <name>` → single task mode
- `--tasks a,b,c` or `--all` → multi-task mode (spawn sub-agents)
- `--iterations N` → how many evaluate → improve cycles (default: 5)
- `--env local|remote` → browser environment (default: local; use remote for bot-protected sites)
- `--browser-trace` → opt in to the browser-trace integration (default off). Implies `--env remote`. If `--env local --browser-trace` are both passed explicitly, error with: `browser-trace requires Browserbase; drop --env local or drop --browser-trace.`

If the user passed free-form text instead, map it to one of the above before continuing.

### Step 2 — Set up the workspace

All training artifacts (task definitions, strategy iterations, traces, reports) live in a workspace directory in the **current working directory** — NOT inside `~/.claude/skills/`. This keeps the inner agent's file writes out of Claude's home dir and away from permission friction.

Default workspace: `${CWD}/autobrowse/`

```bash
mkdir -p ./autobrowse/tasks ./autobrowse/traces ./autobrowse/reports
```

If the task directory (`./autobrowse/tasks/<task>/task.md`) doesn't exist yet, scaffold it:

```bash
mkdir -p ./autobrowse/tasks/<task>
cp ${CLAUDE_SKILL_DIR}/references/example-task.md ./autobrowse/tasks/<task>/task.md
# Then edit task.md to describe the URL, inputs, steps, and expected JSON output
```

The skill source at `${CLAUDE_SKILL_DIR}` stays read-only — only `./autobrowse/` in CWD gets written to during training. Graduation (final step) writes a single file to `~/.claude/skills/<task>/SKILL.md`.

List available tasks:
```bash
ls ./autobrowse/tasks/
```

### Step 3 — Multi-task: spawn parallel sub-agents

If running multiple tasks, use the Agent tool to spawn one sub-agent per task simultaneously. Each sub-agent receives a self-contained prompt to run the full autobrowse loop for its task:

> "You are running the autobrowse skill for task `<name>`. Workspace: `<absolute-path-to-workspace>` (e.g. `/path/to/project/autobrowse`). Run `<N>` iterations of: evaluate → read trace → improve strategy.md → repeat. Use `--env <env>`. Pass `--workspace <workspace>` to every evaluate.mjs invocation. If the parent invocation used `--browser-trace`, you MUST use the traced-path block of the SKILL.md loop for every iteration (pre-create session, attach bb-capture, pass `--connect-url` to evaluate.mjs, stop+bisect, release) — do not fall back to the default single-command path. Follow the autobrowse loop instructions exactly.
>
> When graduating, install the skill to `~/.claude/skills/<task-name>/SKILL.md` with proper agentskills frontmatter (name + description). Do not just copy strategy.md — write a self-contained skill.
>
> At the end, output a structured summary with: task name, pass/fail on final run, total cumulative cost, iterations completed, per-iteration table (iter number, turns, cost, status, hypothesis tested), and 2-3 bullet key learnings."

Spawn all sub-agents in parallel, wait for all to complete, then collect their summaries and write the session report.

**For single task**, skip this step and run the loop directly below.

---

## The Loop (run this for each task)

### Iteration start

Check that `./autobrowse/tasks/<task>/task.md` exists (scaffold it from the template if not — see Step 2). `strategy.md` is auto-created empty by the harness on first run.

### Requirements

- `ANTHROPIC_API_KEY` must be in the environment (or in a `.env` file in CWD — `evaluate.mjs` auto-loads it). If missing, the harness prints a clear error and exits; don't hunt for keys in other paths.

### Run the inner agent

**Default path (no `--browser-trace`)** — single command, no orchestration:

```bash
node ${CLAUDE_SKILL_DIR}/scripts/evaluate.mjs --task <task-name> --workspace ./autobrowse
# or for bot-protected sites:
node ${CLAUDE_SKILL_DIR}/scripts/evaluate.mjs --task <task-name> --workspace ./autobrowse --env remote
```

This runs the browser session and writes a full trace to `./autobrowse/traces/<task>/latest/`.

**Traced path (`--browser-trace`, remote only)** — the outer harness pre-creates a Browserbase session, attaches `bb-capture` as a passive observer, and passes the session's `connectUrl` to `evaluate.mjs` so every inner `browse` call uses `--cdp $connectUrl --session autobrowse-main` (the canonical browser-trace pattern that gives observers full Network/Console events). Run this block once per iteration with `$N` set to the 1-indexed iteration number:

```bash
# Preflight — fail fast if browser-trace isn't installed alongside autobrowse.
BT_DIR="${CLAUDE_SKILL_DIR}/../browser-trace"
if [ ! -f "$BT_DIR/scripts/bb-capture.mjs" ]; then
  echo "ERROR: --browser-trace requires the browser-trace skill at $BT_DIR." >&2
  echo "Install it by cloning github.com/browserbase/skills and copying skills/browser-trace/" >&2
  echo "into the same parent directory as autobrowse (e.g. ~/.claude/skills/browser-trace/)." >&2
  exit 1
fi

# a. SESSION SETUP — pre-create the keep-alive session and derive its connectUrl
sid=$(browse cloud sessions create --keep-alive --verified --proxies \
  | node -e "let s='';process.stdin.on('data',c=>s+=c).on('end',()=>process.stdout.write(JSON.parse(s).id))")
connect_url=$(browse cloud sessions get "$sid" \
  | node -e "let s='';process.stdin.on('data',c=>s+=c).on('end',()=>process.stdout.write(JSON.parse(s).connectUrl))")

RUN_ID="run-$(printf '%03d' "$N")"
TRACE_ROOT="./autobrowse/traces/<task-name>/$RUN_ID"
mkdir -p "$TRACE_ROOT"
export O11Y_ROOT="$TRACE_ROOT/.o11y"   # park browser-trace output inside the autobrowse run dir
export O11Y_RUN_ID="$RUN_ID"           # tells the browse CLI which run dir to write descriptors.ndjson into

# b. ATTACH BROWSER-TRACE — passive observer; runs in background
node ${CLAUDE_SKILL_DIR}/../browser-trace/scripts/bb-capture.mjs "$sid" "$RUN_ID" &
sleep 2

# c. RUN AUTOBROWSE — connectUrl flag tells evaluate.mjs to inject --cdp/--session
#    into every inner browse call. The inner agent never sees --remote.
node ${CLAUDE_SKILL_DIR}/scripts/evaluate.mjs \
  --task <task-name> --workspace ./autobrowse --env remote \
  --connect-url "$connect_url" --run-number "$N"

# d. STOP + BISECT + UNIFY — order matters; bisect needs the session to still
#    exist, and unify-trace joins the bisect output with autobrowse's trace.json
#    into a single time-ordered NDJSON the outer agent reads first each iter.
node ${CLAUDE_SKILL_DIR}/../browser-trace/scripts/stop-capture.mjs "$RUN_ID"
node ${CLAUDE_SKILL_DIR}/../browser-trace/scripts/bisect-cdp.mjs "$RUN_ID"
node ${CLAUDE_SKILL_DIR}/scripts/unify-trace.mjs \
  --trace-dir "$TRACE_ROOT" \
  --o11y-dir "$O11Y_ROOT/$RUN_ID"

# e. RELEASE
browse cloud sessions update "$sid" --status REQUEST_RELEASE
```

This writes the inner-agent trace to `./autobrowse/traces/<task-name>/latest/` and the CDP bisect to `./autobrowse/traces/<task-name>/latest/.o11y/<run-id>/`. The traced `browse` CLI also emits per-command rich node descriptors to `.o11y/<run-id>/cdp/descriptors.ndjson` (one JSON object per page-driving call: target tag/id/role/accessibleName/attributes/xpath/bounding-rect). The descriptors file feeds downstream codegen; it is **not** required for hypothesis formation — skip it when reading the trace.

### Read the trace

```bash
cat ./autobrowse/traces/<task-name>/latest/summary.md
```

The summary has duration, cost, turns, the decision log, and the final JSON output.

If the agent failed or got stuck, look deeper:
- Read `./autobrowse/traces/<task-name>/latest/trace.json` — search for the failure turn
- Read screenshots around the failure point with the Read tool

**When `--browser-trace` was used — start with `unified-events.jsonl`.** The harness joins the agent's turn log and the browser's CDP firehose into one time-ordered NDJSON stream at the run root. One file, source-tagged (`source: "agent" | "browser"`), interleaved by wall-clock timestamp. Skim it top-to-bottom; the failure cause is usually one or two adjacent lines (the agent issued command X, the browser responded with Y).

```bash
cat ./autobrowse/traces/<task-name>/latest/unified-events.jsonl
```

The structured files (`trace.json`, `.o11y/<run-id>/cdp/*`) are **also agent-consumable as drill-downs** when the unified stream points at something you need more of:

| Need | Drill-down file or command |
|---|---|
| Per-page totals + timing (events, network counts, errors by page) | `.o11y/<run-id>/cdp/summary.json` |
| All failed network requests in one place | `.o11y/<run-id>/cdp/network/failed.jsonl` |
| Full console exception payloads (stacktraces, etc.) | `.o11y/<run-id>/cdp/console/exceptions.jsonl` |
| Per-page slice (only events on page N) | `.o11y/<run-id>/cdp/pages/<pid>/` |
| Full reasoning text / untruncated tool outputs for a specific turn | `trace.json` (filter by `turn === N`) |
| Ad-hoc grouped query (e.g. top hosts, errors-by-page) | `O11Y_ROOT=./autobrowse/traces/<task-name>/latest/.o11y node ${CLAUDE_SKILL_DIR}/../browser-trace/scripts/query.mjs <run-id> <cmd>` |

The unified stream is the default; drill into structured files only when you need a grouped query, a full-text payload, or filtering the stream can't give you.

### Form one hypothesis

Find the exact turn where things went wrong. What single heuristic would have prevented it?

Under `--browser-trace`, the hypothesis must cite a **specific event from `unified-events.jsonl`** (line number or timestamp) — or name the drill-down file if you had to descend into one. This keeps updates evidence-grounded rather than vibes-driven. A hypothesis based only on the agent's commands might say "the click didn't work"; grounded in the unified stream, it can say "line 47 of unified-events.jsonl: `browse open` was followed by `Network.responseReceived` status 403 on `/api/checkout` — switch to `--verified --proxies`."

Examples:
- "After clicking the dropdown, wait 1s — options animate in before they're clickable"
- "Navigate directly to `/pay-invoice/` — skip the landing page entirely"
- "Use `browse fill #field_3 value` not `browse type` — this field clears on focus"
- "The page shows a spinner at turn 8 — add `browse wait timeout 2000` before snapshot"
- (with `--browser-trace`) "At line 47 of unified-events.jsonl, 3 consecutive `Network.responseReceived` events on `/api/availability` returned 403 right after `browse open` — the site is fingerprinting; the next iter needs `--verified --proxies`."

### Update strategy.md

Edit `./autobrowse/tasks/<task-name>/strategy.md`. Keep everything that worked. Fix the specific failure. Add a concrete heuristic.

Good strategies have:
- **Fast path**: direct URL or shortcuts to skip exploration
- **Step-by-step workflow**: exact sequence with timing notes
- **Site-specific knowledge**: selector IDs, form field names, success indicators
- **Failure recovery**: what to do when X goes wrong

### Judge the result

Read the new summary. Did it pass? Make clear progress?
- **Pass or progress** → keep, next iteration
- **No progress or regression** → revert strategy.md to the previous version and try a different hypothesis

### Generate a runnable script (optional)

Once the task has converged, you can produce a deterministic, runnable script
in one or more frameworks via `scripts/codegen.mjs`. This is one shot of an
LLM call per framework, cached by content hash, with optional verify-against-
fresh-session and rewrite-on-failure.

```bash
node ${CLAUDE_SKILL_DIR}/scripts/codegen.mjs \
  --task <name> \
  --workspace ./autobrowse \
  --frameworks playwright,stagehand \
  --verify
```

Each framework gets its own subdirectory under `tasks/<name>/<framework>/`
with the emitted script and a self-contained scaffold (`package.json`,
`tsconfig.json`). The directory is runnable standalone with
`cd tasks/<name>/playwright && npm install && npx tsx <name>.ts` — the only
runtime requirement is `BROWSERBASE_API_KEY` (plus `ANTHROPIC_API_KEY` for
the Stagehand target).

Builtin frameworks: `playwright`, `stagehand`. Add a custom framework with
`--prompt-template <path> --frameworks custom` (and provide your own runner
or pass `--no-verify`).

Common flags:

| Flag | Purpose |
|---|---|
| `--frameworks a,b,...` | Comma-separated; default `playwright` |
| `--verify` / `--no-verify` | Run the produced script against a fresh BB session; default `--verify` |
| `--max-retries N` | Rewrite-on-verify-failure cap; default 2 |
| `--cache-only` | Error if cache miss (CI-friendly) |
| `--force` | Bust the cache |
| `--dry-run` | Estimate prompt size + cost; don't call the LLM |
| `--run <id>` | Force a specific `run-NNN` (default: latest passing) |

Output is one JSON line per framework on stdout. Non-zero exit if any
selected framework's final state is `passed: false`.

See `references/playwright-cdp-bridge.md` for the canonical
`connectOverCDP` patterns the emitted scripts follow.

### After all iterations — publish if ready

If the task passed on 2+ of the last 3 iterations **or has reached the max iteration limit**, install it as a Claude Code skill. **Do not just copy strategy.md** — the skill must be self-contained and useful to someone who has never seen this codebase. If graduating at max iterations without a clean pass, note the known failure point but still document everything learned.

Install by writing to `~/.claude/skills/<task-name>/SKILL.md`:

```bash
mkdir -p ~/.claude/skills/<task-name>
```

Use this structure for the SKILL.md:

```markdown
---
name: <task-name>
description: <1-2 sentences describing what this skill does and when to use it. Include trigger keywords.>
---

# <Task Title> — Browser Skill

## Purpose
<1-2 sentences: what this automates and why it exists.>

## When to Use
<When should someone reach for this skill.>

## Browse CLI Reference
The inner agent uses the `browse` CLI. Key commands for this task:
- `browse stop` — kill existing session (always run before switching to remote)
- `browse open <url> --remote` — start a fresh Browserbase cloud session and navigate
- `browse open <url> --local` — start a clean local browser and navigate
- `browse tab new <url>` — open URL in a new tab
- `browse wait load` — wait for page to finish loading
- `browse wait timeout <ms>` — wait a fixed amount of time for spinners or animations
- `browse wait selector "<selector>"` — wait for an element to become visible
- `browse get title` — verify you're on the right page
- `browse get text body` — extract all visible text (preferred for content extraction)
- `browse snapshot` — get accessibility tree; each node has a ref in `[X-Y]` format (e.g. `[0-5]`, `[2-147]`)
- `browse click [X-Y]` — click element by ref from the latest snapshot (include the brackets)

**Never use `--session <name>` flags in SKILL.md.** Named sessions are a parallel-run workaround — they contaminate skills with infrastructure concerns. Skills must work in isolation with the default session.

## Workflow

### Step 1 — Start session
<exact browse commands in order>

### Step 2 — Navigate
<exact URL and verification steps>

### Step 3 — Extract
<exact extraction commands>

### Step 4 — Output
<what JSON to emit, referencing the schema below>

## Site-Specific Gotchas
<Bullet list of every hard-won heuristic from the iterations. This is the core value of the skill.>

## Failure Recovery
<What to do when navigation fails, session is contaminated, or extraction returns garbage>

## Expected Output
```json
<paste the exact expected output schema from task.md>
```
```

After writing the SKILL.md, confirm it's installed:
```bash
ls ~/.claude/skills/<task-name>/SKILL.md
```

The skill is now available as `/<task-name>` in Claude Code.

---

## Final report (multi-task mode)

After all sub-agents complete, print a markdown table:

| Task | Iterations | Final Status | Graduated | Cost |
|------|-----------|--------------|-----------|------|
| google-flights | 5 | ✅ pass | yes | $0.42 |
| amazon-add-to-cart | 5 | ❌ fail | no | $1.20 |

Then write a persistent session report to `./autobrowse/reports/` so there's a durable record of the run inside the workspace:

```bash
mkdir -p ./autobrowse/reports
```

Write the file `./autobrowse/reports/YYYY-MM-DD-HH-MM-<tasks>.md` with:

```markdown
# AutoBrowse Session Report
**Date:** <ISO date>
**Tasks:** <comma-separated list>
**Environment:** remote|local
**Total cost:** $X.XX

## Results

| Task | Iterations | Pass Rate | Final Status | Graduated | Cost |
|------|-----------|-----------|--------------|-----------|------|
| ... | ... | X/5 | ✅/❌ | yes/no | $X.XX |

## Per-Task Learnings

### <task-name>
- **Key insight 1:** <what the agent learned>
- **Key insight 2:** <another heuristic>
- **Failure mode fixed:** <what was failing and how it was resolved>

## Iteration Log

### <task-name>
| Iter | Turns | Cost | Status | Hypothesis tested |
|------|-------|------|--------|-------------------|
| 1 | 79 | $18.75 | ❌ fail | baseline |
| 2 | 9 | $0.26 | ✅ pass | session contamination fix |
| ... | ... | ... | ... | ... |
```

---

## Rules

- **Only edit `strategy.md`** — never touch `task.md` (unless creating it from the template) or `evaluate.mjs`
- **Stay in the workspace** — all training writes go to `./autobrowse/`, never to `~/.claude/skills/autobrowse/`. The skill source is read-only.
- **One hypothesis per iteration** — test one change at a time
- **Build on wins** — keep what worked, add to it
- **Trust the trace** — the inner agent shows exactly what it saw and did
- **Graduate to `~/.claude/skills/`** — the only file you write there is the final graduated `SKILL.md`
- **Don't release before bisecting** — under `--browser-trace`, the order at the end of each iteration is non-negotiable: `stop-capture` → `bisect-cdp` → `browse cloud sessions update REQUEST_RELEASE`. Bisect depends on the session still existing when the trace stops.

安装 autobrowse

下载技能文件并将其解压到 .claude/skills/ 目录中。

下载ZIP

克隆仓库并复制技能文件到您的项目中。

git clone https://github.com/browserbase/skills/tree/main/skills/autobrowse # Copy SKILL.md to your .claude/skills/ directory

复制 复制
快速设置: 将技能文件夹复制到 .claude/skills/ Claude 将自动检测并使用该技能

相关技能

klingai-upgrade-migration
更新时间 2026-07-03
Verification &amp; Quality Assurance
更新时间 2026-06-29
base44-cli
更新时间 2026-06-29
Railway CLI Management
更新时间 2026-07-02
OR