autobrowse
browserbase/skills
通过迭代实验,运行内部代理来浏览网站,并不断优化导航指令,直至任务能够稳定通过,从而掌握可靠的浏览器自动化技能。
...展开全部AutoBrowse — 自我提升的浏览器自动化技能
通过迭代实验,培养可靠的浏览器自动化技能。一个内部代理负责浏览网站(evaluate.ts)。而你——作为外部代理——则分析发生的情况,并优化操作指令(strategy.md)。重复这一过程,直到系统能够稳定通过测试。
入口点
调用方式灵活——既支持显式参数,也支持自由形式的自然语言:
/autobrowse --task google-flights
/autobrowse --task google-flights --iterations 10 --env remote
/autobrowse --task google-flights --browser-trace
/autobrowse --tasks google-flights,amazon-add-to-cart
/autobrowse --all
# 以下方式同样有效——可自由解析:
/autobrowse https://flights.google.com/
/autobrowse 在 delta.com 上预订机票
/autobrowse 修复现有的 google-flights 技能
--browser-trace(默认关闭,仅限远程模式):将每次迭代与同级的browser-trace技能配对——将内部代理封装在 CDP 捕获中,以获取每页的网络/控制台/页面生命周期证据。 隐含--env remote;若与--env local 结合使用将引发错误。要求同级browser-trace技能位于${CLAUDE_SKILL_DIR}/../browser-trace/ 目录下,且需设置BROWSERBASE_API_KEY环境变量。
当用户输入 URL 或自由格式指令而非--task 时:
- 如果
${WORKSPACE}/tasks/中已有任务与该网站/意图明显匹配,则使用该任务。 - 否则,选择一个简短的鞑靼式命名,根据
${CLAUDE_SKILL_DIR}/references/example-task.md生成${WORKSPACE}/tasks/,根据用户所述内容填写 URL/目标,然后继续执行。 用一句话告知用户所选的名称。/task.md
运行方法
步骤 1 — 解析参数并确定方向
检查传入的参数:
--task→ 单任务模式--tasks a,b,c或--all→ 多任务模式(启动子代理)--iterations N→ 评估 → 优化循环的次数(默认:5)--env local|remote→ 浏览器环境(默认:local;访问受机器人防护的网站时使用 remote)--browser-trace→ 启用浏览器跟踪集成(默认关闭)。 该选项默认包含--env remote。若同时显式指定了--env local 和 --browser-trace,将报错提示:browser-trace 需要 Browserbase;请移除 --env local 或移除 --browser-trace。
如果用户传入的是自由格式文本,请将其映射到上述选项之一后再继续。
步骤 2 — 设置工作区
所有训练成果(任务定义、策略迭代、跟踪记录、报告)均存储在当前工作目录中的工作区目录内——而非~/.claude/skills/ 目录内。此举可确保内部代理的文件写入操作不进入 Claude 的主目录,从而避免权限冲突。
默认工作区:${CWD}/autobrowse/
mkdir -p ./autobrowse/tasks ./autobrowse/traces ./autobrowse/reports
如果任务目录(./autobrowse/tasks/)尚未存在,请创建该目录:
mkdir -p ./autobrowse/tasks/
cp ${CLAUDE_SKILL_DIR}/references/example-task.md ./autobrowse/tasks//task.md
# 然后编辑 task.md 文件,描述 URL、输入、步骤以及预期的 JSON 输出
位于${CLAUDE_SKILL_DIR}的技能源文件保持只读状态——训练过程中仅向当前工作目录(CWD)下的./autobrowse/目录写入数据。毕业(最后一步)会向~/.claude/skills/ 写入一个文件。
列出可用任务:
ls ./autobrowse/tasks/
步骤 3 — 多任务:并行启动子代理
若运行多个任务,请使用 Agent 工具为每个任务同时启动一个子代理。每个子代理会收到一个自包含的提示,用于执行其任务的完整autobrowse 循环:
autobrowse “您正在为任务
。工作区:(例如/path/to/project/autobrowse)。执行以下循环:评估 → 读取跟踪信息 → 改进 strategy.md → 重复。 请使用--env。向每次 evaluate.mjs 调用传递--workspace。如果父调用使用了--browser-trace,您必须在每次迭代中使用 SKILL.md 循环中的 traced-path 代码块(预创建会话、附加 bb-capture、将--connect-url参数传递给 evaluate.mjs、停止+二分法排查、发布)——切勿回退到默认的单命令路径。请严格遵循autobrowse 中的循环操作指南。毕业时,请将技能安装到
~/.claude/skills/并确保 frontmatter 中包含正确的 agentskills 信息(名称 + 描述)。请勿仅复制 strategy.md —— 应编写一个自包含的技能。/SKILL.md, 最后,输出一份结构化总结,包含:任务名称、最终运行结果(通过/失败)、累计总成本、已完成的迭代次数、每轮迭代表格(迭代编号、回合数、成本、状态、测试的假设),以及2-3条关键收获的要点。"
并行启动所有子代理,等待全部完成后,收集其摘要并撰写会话报告。
对于单个任务,请跳过此步骤,直接运行下方的循环。
循环(针对每个任务运行)
迭代开始
检查./autobrowse/tasks/是否存在(若不存在,请根据模板生成——参见步骤 2)。strategy.md文件会在首次运行时由测试框架自动生成(初始为空)。
要求
- 环境中必须存在
ANTHROPIC_API_KEY(或位于当前工作目录下的.env文件中——evaluate.mjs会自动加载该文件)。若缺失,测试框架将输出明确的错误信息并退出;请勿在其他路径中查找该密钥。
运行内部代理
默认路径(不带--browser-trace参数)——单条命令,无需编排:
node ${CLAUDE_SKILL_DIR}/scripts/evaluate.mjs --task --workspace ./autobrowse
# 或针对受机器人保护的网站:
node ${CLAUDE_SKILL_DIR}/scripts/evaluate.mjs --task --workspace ./autobrowse --env remote
这将启动浏览器会话,并将完整的跟踪记录写入./autobrowse/traces/
跟踪路径(--browser-trace,仅限远程)——外部测试框架会预先创建一个 Browserbase 会话,将bb-capture作为被动观察者附加,并将该会话的connectUrl传递给evaluate.mjs,以便每个内部browse调用都使用--cdp $connectUrl --sessionautobrowse-main(标准 browser-trace 模式,可为观察者提供完整的 Network/Console 事件)。每轮迭代执行一次此代码块,其中$N设置为从 1 开始计数的迭代编号:
# 预检 — 若 browser-trace 未与autobrowse 一同安装,则快速报错。
BT_DIR="${CLAUDE_SKILL_DIR}/../browser-trace"
if [ ! -f "$BT_DIR/scripts/bb-capture.mjs" ]; then
echo "错误:--browser-trace 需要在 $BT_DIR 目录下安装 browser-trace 技能。" >&2
echo "请克隆 github.com/browserbase/skills,并将 skills/browser-trace/" >&2
echo "复制到与autobrowse 相同的父目录下(例如 ~/.claude/skills/browser-trace/)。" >&2
exit 1
fi
# a. 会话设置 — 预先创建保持活动状态的会话并推导其 connectUrl
sid=$(browse cloud sessions create --keep-alive --verified --proxies \
| node -e "let s='';process.stdin.on('data',c=>s+=c).on('end',()=>process.stdout.write(JSON.parse(s).id))")
connect_url=$(browse cloud sessions get "$sid" \
| node -e "let s='';process.stdin.on('data',c=>s+=c).on('end',()=>process.stdout.write(JSON.parse(s).connectUrl))")
RUN_ID="run-$(printf '%03d' "$N")"
TRACE_ROOT="./autobrowse/traces//$RUN_ID"
mkdir -p "$TRACE_ROOT"
export O11Y_ROOT="$TRACE_ROOT/.o11y" # 将浏览器跟踪输出保存在autobrowse 运行目录中
export O11Y_RUN_ID="$RUN_ID" # 告知 browse CLI 应将 descriptors.ndjson 写入哪个运行目录
# b. 附加浏览器跟踪 — 被动观察者;在后台运行
node ${CLAUDE_SKILL_DIR}/../browser-trace/scripts/bb-capture.mjs "$sid" "$RUN_ID" &
sleep 2
# c. 运行AUTOBROWSE — connectUrl 参数指示 evaluate.mjs 将 --cdp/--session
# 注入到每个内部 browse 调用中。内部代理永远不会看到 --remote。
node ${CLAUDE_SKILL_DIR}/scripts/evaluate.mjs \
--task --workspace ./autobrowse --env remote \
--connect-url "$connect_url" --run-number "$N"
# d. STOP + BISECT + UNIFY — 执行顺序至关重要;bisect 需要会话仍
# 存在,而 unify-trace 会将 bisect 的输出与autobrowse 的 trace.json 合并
# 合并为一个按时间排序的 NDJSON,外部代理在每次迭代中首先读取该文件。
node ${CLAUDE_SKILL_DIR}/../browser-trace/scripts/stop-capture.mjs "$RUN_ID"
node ${CLAUDE_SKILL_DIR}/../browser-trace/scripts/bisect-cdp.mjs "$RUN_ID"
node ${CLAUDE_SKILL_DIR}/scripts/unify-trace.mjs \
--trace-dir "$TRACE_ROOT" \
--o11y-dir "$O11Y_ROOT/$RUN_ID"
# e. 发布
browse cloud sessions update "$sid" --status REQUEST_RELEASE
这会将代理内部跟踪写入./autobrowse/traces/,并将 CDP 二分法写入./autobrowse/traces/。 被追踪的浏览CLI 还会向.o11y/输出按命令划分的丰富节点描述符(每个页面驱动调用对应一个 JSON 对象:target tag/id/role/accessibleName/attributes/xpath/bounding-rect)。 该描述符文件用于为下游代码生成提供数据;它并非假设形成所必需的——在阅读跟踪信息时可跳过该部分。
读取跟踪记录
cat ./autobrowse/traces//latest/summary.md
摘要包含持续时间、成本、迭代次数、决策日志以及最终的 JSON 输出。
如果智能体失败或卡住,请进一步检查:
- 阅读
./autobrowse/traces/— 查找失败的回合/latest/trace.json - 使用 Read 工具查看故障点附近的截图
当使用了--browser-trace选项时——请从unified-events.jsonl 文件开始分析。测试框架会将代理的回合日志和浏览器的 CDP 数据流合并为一个按时间排序的 NDJSON 流,并存储在运行根目录下。 这是一个带有源标签(source: "agent" | "browser")的文件,各数据按墙钟时间戳交错排列。请从上到下快速浏览;故障原因通常出现在一两行相邻的记录中(代理发出命令 X,浏览器响应为 Y)。
cat ./autobrowse/traces//latest/unified-events.jsonl
当统一数据流指向您需要进一步深入分析的内容时,结构化文件(trace.json、.o11y/)也可作为代理的钻取数据源:
| 需求 | 钻取文件或命令 |
|---|---|
| 每页总计 + 时序数据(事件、网络计数、按页面划分的错误) | .o11y/ |
| 所有失败的网络请求汇总 | .o11y/ |
| 完整的控制台异常有效载荷(堆栈跟踪等) | .o11y/ |
| 按页面切片(仅包含第 N 页上的事件) | .o11y/ |
| 特定回合的完整推理文本/未截断的工具输出 | trace.json(按轮次 === N 过滤) |
| 临时分组查询(例如:顶级主机、按页面分类的错误) | O11Y_ROOT=./autobrowse/traces/ |
统一数据流是默认选项;仅当您需要分组查询、全文有效载荷,或数据流无法满足的过滤需求时,才需深入到结构化文件中。
提出一个假设
找出问题确切发生的转折点。如果采用哪一项单一启发式方法,本可以避免这一问题?
在--browser-trace 模式下,假设必须引用 unified-events.jsonl中的具体事件(行号或时间戳)——若需深入分析,则需指名相应的钻取文件。这样能确保更新基于确凿证据,而非凭主观感觉。 仅基于用户操作的假设可能表述为“点击未生效”; 若基于统一事件流,则可表述为“unified-events.jsonl 的第 47 行:browse open之后,/api/checkout出现Network.responseReceived状态码 403 —— 请切换至--verified --proxies 模式”。
示例:
- “点击下拉菜单后等待 1 秒——选项在可点击前会先以动画形式显示”
- “直接导航至
/pay-invoice/— 完全跳过登录页面” - “使用
‘填入’操作设置 #field_3 的值,而非‘浏览’类型——该字段在获得焦点时会清空” - “页面在第 8 步显示加载图标 — 在快照前
为浏览操作添加2000 毫秒的等待超时” - (使用
--browser-trace选项) “在 unified-events.jsonl 的第 47 行,/api/availability上的 3 个连续Network.responseReceived事件在打开浏览器窗口后立即返回 403 状态码——该网站正在进行指纹识别;下一次迭代需使用--verified --proxies参数。”
更新 strategy.md
编辑./autobrowse/tasks/。保留所有已验证有效的内容。修复具体故障。添加具体的启发式规则。
优秀的策略应具备:
- 快速路径:直接 URL 或可跳过探索的快捷方式
- 分步工作流:带有时间注释的精确操作序列
- 特定网站知识:选择器 ID、表单字段名称、成功指示符
- 故障恢复:当 X 出错时该如何处理
评估结果
阅读新摘要。是否通过?是否取得明显进展?
- 通过或有进展→ 保留,进入下一次迭代
- 无进展或出现退步→ 将 strategy.md 还原至上一版本,并尝试其他假设
生成可运行脚本(可选)
一旦任务收敛,您可以通过scripts/codegen.mjs
在一种或多种框架中生成一个确定性的、可运行的脚本。这是针对每个框架的
单次 LLM 调用,按内容哈希进行缓存,并可选地支持
“与最新会话进行验证”和“失败时重写”功能。
node ${CLAUDE_SKILL_DIR}/scripts/codegen.mjs \
--task \
--workspace ./autobrowse \
--frameworks playwright,stagehand \
--verify
每个框架都在tasks/、下拥有自己的子目录,
其中包含生成的脚本和一个自包含的框架(package.json、
tsconfig.json)。 该目录可通过以下命令独立运行:
cd tasks/——唯一的
运行时要求是BROWSERBASE_API_KEY(针对 Stagehand 目标还需
ANTHROPIC_API_KEY)。
内置框架:playwright、stagehand。使用
--prompt-template添加自定义框架(并提供您自己的运行器
或传入--no-verify)。
常用标志:
| 参数 | 用途 |
|---|---|
--frameworks a,b,... |
以逗号分隔;默认值为playwright |
--verify/--no-verify |
在全新的 BB 会话中运行生成的脚本;默认值为--verify |
--max-retries N |
验证失败时的重写次数上限;默认值为 2 |
--仅缓存 |
若缓存未命中则报错(适合持续集成) |
--force |
清空缓存 |
--dry-run |
估算提示词长度及计算成本;不调用大型语言模型 |
--run |
强制使用特定的运行-NNN(默认:最新通过的) |
输出为每个框架一行 JSON 数据,输出至标准输出。若
任何选定框架的最终状态为“false”,则返回非零退出状态。
请参阅references/playwright-cdp-bridge.md,了解生成的脚本遵循的规范
connectOverCDP模式。
所有迭代结束后——若已准备就绪则发布
如果任务在最近 3 次迭代中有 2 次及以上通过,或已达到最大迭代次数限制,则将其安装为 Claude Code 技能。请勿简单复制 strategy.md—— 该技能必须自成体系,并对从未接触过此代码库的人具有实用价值。若在达到最大迭代次数时仍未完全通过测试,请记录已知的失败点,但仍需完整记录所学内容。
通过在~/.claude/skills/ 目录下编写/SKILL.md 文件进行安装:
mkdir -p ~/.claude/skills/
SKILL.md 文件请采用以下结构:
---
name:
description:<1-2 sentences describing what this skill does and when to use it. Include trigger keywords.>
---
# — 浏览器技能
## 目的
<1-2 sentences: what this automates and why it exists.>
## 适用场景
## 浏览 CLI 参考文档
内部代理使用 `browse` CLI。此任务的关键命令:
- `browse stop` — 终止现有会话(切换到远程模式前务必执行)
- `browse open --remote` — 启动一个全新的 Browserbase 云会话并进行浏览
- `browse open --local` — 启动一个干净的本地浏览器并进行浏览
- `browse tab new` — 在新标签页中打开 URL
- `browse wait load` — 等待页面加载完成
- `browse wait timeout` — 等待固定时间,以处理加载转圈图标或动画效果
- `browse wait selector ""` — 等待某个元素变得可见
- `browse get title` — 验证当前是否处于正确页面
- `browse get text body` — 提取所有可见文本(推荐用于内容提取)
- `browse snapshot` — 获取辅助功能树;每个节点都有一个格式为 `[X-Y]` 的引用(例如 `[0-5]`、`[2-147]`)
- `browse click [X-Y]` — 根据最新快照中的引用点击元素 (包括方括号)
**切勿在 SKILL.md 中使用 `--session` 参数。** 命名会话是一种并行运行的变通方案——它们会将基础设施问题引入技能中。技能必须在默认会话下独立运行。
## 工作流程
### 步骤 1 — 启动会话
### 步骤 2 — 导航
### 步骤 3 — 提取
### 步骤 4 — 输出
## 特定于网站的注意事项
## 故障恢复
## 预期输出
```json
编写完 SKILL.md 后,请确认其已安装:
```bash
ls ~/.claude/skills//SKILL.md
该技能现已在 Claude Code 中以/的形式提供。
最终报告(多任务模式)
当所有子代理任务完成后,输出一个 Markdown 表格:
| 任务 | 迭代次数 | 最终状态 | 已毕业 | 成本 |
|---|---|---|---|---|
| google-flights | 5 | ✅ 通过 | 是 | 0.42美元 |
| 亚马逊-加入购物车 | 5 | ❌ 失败 | 否 | $1.20 |
然后将持久化会话报告写入./autobrowse/reports/,以便在工作区中保留本次运行的持久记录:
mkdir -p ./autobrowse/reports
创建文件./autobrowse/reports/YYYY-MM-DD-HH-MM-,内容如下:
#AutoBrowse 会话报告
**日期:**
**任务:**
**环境:** 远程|本地
**总成本:** $X.XX
## 结果
| 任务 | 迭代次数 | 通过率 | 最终状态 | 是否通过 | 成本 |
|------|-----------|-----------|--------------|-----------|------|
| ... | ... | X/5 | ✅/❌ | 是/否 | $X.XX |
## 各任务经验总结
###
- **关键见解 1:**
- **关键见解 2:**
- **已修复的故障模式:**
## 迭代日志
###
| 迭代 | 轮次 | 成本 | 状态 | 验证的假设 |
|------|-------|------|--------|-------------------|
| 1 | 79 | $18.75 | ❌ 失败 | 基线 |
| 2 | 9 | $0.26 | ✅ 通过 | 会话污染修复 |
| ... | ... | ... | ... | ... |
规则
- 仅编辑
strategy.md—— 切勿修改task.md(除非是从模板创建)或evaluate.mjs - 请留在工作区内— 所有训练写入操作均应指向
./autobrowse/,切勿写入~/.claude/skills/autobrowse/。技能源文件为只读。 - 每次迭代仅测试一个假设— 每次只测试一项更改
- 基于成功经验进行扩展——保留行之有效的方法,并在其基础上进行补充
- 信任追踪记录——内部代理会精确展示其所见所为
- 迁移至
~/.claude/skills/—— 该目录下唯一可写入的文件是最终通过验证的SKILL.md - 在进行二分法排查前不要发布—— 在
--browser-trace模式下,每次迭代结束时的操作顺序是不可更改的:stop-capture→bisect-cdp→浏览云会话并更新 REQUEST_RELEASE。二分法排查依赖于在追踪停止时会话仍然存在。
---
name: autobrowse
description: Builds reliable browser automation skills through iterative experimentation, running an inner agent to browse sites and improving navigation instructions until tasks pass consistently.
license: MIT
---
# AutoBrowse — Self-Improving Browser Skill
Build reliable browser automation skills through iterative experimentation. An inner agent browses the site (`evaluate.ts`). You — the outer agent — read what happened and improve the instructions (`strategy.md`). Repeat until it passes consistently.
## Entry Points
Invocation is flexible — both explicit flags and free-form natural language work:
```
/autobrowse --task google-flights
/autobrowse --task google-flights --iterations 10 --env remote
/autobrowse --task google-flights --browser-trace
/autobrowse --tasks google-flights,amazon-add-to-cart
/autobrowse --all
# Also fine — parse freely:
/autobrowse https://flights.google.com/
/autobrowse book a flight on delta.com
/autobrowse fix the existing google-flights skill
```
`--browser-trace` (default off, remote-only): pairs each iteration with the sibling `browser-trace` skill — wraps the inner agent in a CDP capture for per-page network/console/page-lifecycle evidence. Implies `--env remote`; errors if combined with `--env local`. Requires the sibling `browser-trace` skill present at `${CLAUDE_SKILL_DIR}/../browser-trace/`, and the `BROWSERBASE_API_KEY` env var.
When the user drops a URL or free-form instruction instead of `--task <name>`:
- If an existing task in `${WORKSPACE}/tasks/` clearly matches the site/intent, use it.
- Otherwise, pick a short kebab-case name, create `${WORKSPACE}/tasks/<name>/task.md` from `${CLAUDE_SKILL_DIR}/references/example-task.md`, fill in the URL/goal based on what the user said, and proceed. Tell the user the chosen name in one line.
---
## How to run
### Step 1 — Parse arguments and orient
Check what was passed:
- `--task <name>` → single task mode
- `--tasks a,b,c` or `--all` → multi-task mode (spawn sub-agents)
- `--iterations N` → how many evaluate → improve cycles (default: 5)
- `--env local|remote` → browser environment (default: local; use remote for bot-protected sites)
- `--browser-trace` → opt in to the browser-trace integration (default off). Implies `--env remote`. If `--env local --browser-trace` are both passed explicitly, error with: `browser-trace requires Browserbase; drop --env local or drop --browser-trace.`
If the user passed free-form text instead, map it to one of the above before continuing.
### Step 2 — Set up the workspace
All training artifacts (task definitions, strategy iterations, traces, reports) live in a workspace directory in the **current working directory** — NOT inside `~/.claude/skills/`. This keeps the inner agent's file writes out of Claude's home dir and away from permission friction.
Default workspace: `${CWD}/autobrowse/`
```bash
mkdir -p ./autobrowse/tasks ./autobrowse/traces ./autobrowse/reports
```
If the task directory (`./autobrowse/tasks/<task>/task.md`) doesn't exist yet, scaffold it:
```bash
mkdir -p ./autobrowse/tasks/<task>
cp ${CLAUDE_SKILL_DIR}/references/example-task.md ./autobrowse/tasks/<task>/task.md
# Then edit task.md to describe the URL, inputs, steps, and expected JSON output
```
The skill source at `${CLAUDE_SKILL_DIR}` stays read-only — only `./autobrowse/` in CWD gets written to during training. Graduation (final step) writes a single file to `~/.claude/skills/<task>/SKILL.md`.
List available tasks:
```bash
ls ./autobrowse/tasks/
```
### Step 3 — Multi-task: spawn parallel sub-agents
If running multiple tasks, use the Agent tool to spawn one sub-agent per task simultaneously. Each sub-agent receives a self-contained prompt to run the full autobrowse loop for its task:
> "You are running the autobrowse skill for task `<name>`. Workspace: `<absolute-path-to-workspace>` (e.g. `/path/to/project/autobrowse`). Run `<N>` iterations of: evaluate → read trace → improve strategy.md → repeat. Use `--env <env>`. Pass `--workspace <workspace>` to every evaluate.mjs invocation. If the parent invocation used `--browser-trace`, you MUST use the traced-path block of the SKILL.md loop for every iteration (pre-create session, attach bb-capture, pass `--connect-url` to evaluate.mjs, stop+bisect, release) — do not fall back to the default single-command path. Follow the autobrowse loop instructions exactly.
>
> When graduating, install the skill to `~/.claude/skills/<task-name>/SKILL.md` with proper agentskills frontmatter (name + description). Do not just copy strategy.md — write a self-contained skill.
>
> At the end, output a structured summary with: task name, pass/fail on final run, total cumulative cost, iterations completed, per-iteration table (iter number, turns, cost, status, hypothesis tested), and 2-3 bullet key learnings."
Spawn all sub-agents in parallel, wait for all to complete, then collect their summaries and write the session report.
**For single task**, skip this step and run the loop directly below.
---
## The Loop (run this for each task)
### Iteration start
Check that `./autobrowse/tasks/<task>/task.md` exists (scaffold it from the template if not — see Step 2). `strategy.md` is auto-created empty by the harness on first run.
### Requirements
- `ANTHROPIC_API_KEY` must be in the environment (or in a `.env` file in CWD — `evaluate.mjs` auto-loads it). If missing, the harness prints a clear error and exits; don't hunt for keys in other paths.
### Run the inner agent
**Default path (no `--browser-trace`)** — single command, no orchestration:
```bash
node ${CLAUDE_SKILL_DIR}/scripts/evaluate.mjs --task <task-name> --workspace ./autobrowse
# or for bot-protected sites:
node ${CLAUDE_SKILL_DIR}/scripts/evaluate.mjs --task <task-name> --workspace ./autobrowse --env remote
```
This runs the browser session and writes a full trace to `./autobrowse/traces/<task>/latest/`.
**Traced path (`--browser-trace`, remote only)** — the outer harness pre-creates a Browserbase session, attaches `bb-capture` as a passive observer, and passes the session's `connectUrl` to `evaluate.mjs` so every inner `browse` call uses `--cdp $connectUrl --session autobrowse-main` (the canonical browser-trace pattern that gives observers full Network/Console events). Run this block once per iteration with `$N` set to the 1-indexed iteration number:
```bash
# Preflight — fail fast if browser-trace isn't installed alongside autobrowse.
BT_DIR="${CLAUDE_SKILL_DIR}/../browser-trace"
if [ ! -f "$BT_DIR/scripts/bb-capture.mjs" ]; then
echo "ERROR: --browser-trace requires the browser-trace skill at $BT_DIR." >&2
echo "Install it by cloning github.com/browserbase/skills and copying skills/browser-trace/" >&2
echo "into the same parent directory as autobrowse (e.g. ~/.claude/skills/browser-trace/)." >&2
exit 1
fi
# a. SESSION SETUP — pre-create the keep-alive session and derive its connectUrl
sid=$(browse cloud sessions create --keep-alive --verified --proxies \
| node -e "let s='';process.stdin.on('data',c=>s+=c).on('end',()=>process.stdout.write(JSON.parse(s).id))")
connect_url=$(browse cloud sessions get "$sid" \
| node -e "let s='';process.stdin.on('data',c=>s+=c).on('end',()=>process.stdout.write(JSON.parse(s).connectUrl))")
RUN_ID="run-$(printf '%03d' "$N")"
TRACE_ROOT="./autobrowse/traces/<task-name>/$RUN_ID"
mkdir -p "$TRACE_ROOT"
export O11Y_ROOT="$TRACE_ROOT/.o11y" # park browser-trace output inside the autobrowse run dir
export O11Y_RUN_ID="$RUN_ID" # tells the browse CLI which run dir to write descriptors.ndjson into
# b. ATTACH BROWSER-TRACE — passive observer; runs in background
node ${CLAUDE_SKILL_DIR}/../browser-trace/scripts/bb-capture.mjs "$sid" "$RUN_ID" &
sleep 2
# c. RUN AUTOBROWSE — connectUrl flag tells evaluate.mjs to inject --cdp/--session
# into every inner browse call. The inner agent never sees --remote.
node ${CLAUDE_SKILL_DIR}/scripts/evaluate.mjs \
--task <task-name> --workspace ./autobrowse --env remote \
--connect-url "$connect_url" --run-number "$N"
# d. STOP + BISECT + UNIFY — order matters; bisect needs the session to still
# exist, and unify-trace joins the bisect output with autobrowse's trace.json
# into a single time-ordered NDJSON the outer agent reads first each iter.
node ${CLAUDE_SKILL_DIR}/../browser-trace/scripts/stop-capture.mjs "$RUN_ID"
node ${CLAUDE_SKILL_DIR}/../browser-trace/scripts/bisect-cdp.mjs "$RUN_ID"
node ${CLAUDE_SKILL_DIR}/scripts/unify-trace.mjs \
--trace-dir "$TRACE_ROOT" \
--o11y-dir "$O11Y_ROOT/$RUN_ID"
# e. RELEASE
browse cloud sessions update "$sid" --status REQUEST_RELEASE
```
This writes the inner-agent trace to `./autobrowse/traces/<task-name>/latest/` and the CDP bisect to `./autobrowse/traces/<task-name>/latest/.o11y/<run-id>/`. The traced `browse` CLI also emits per-command rich node descriptors to `.o11y/<run-id>/cdp/descriptors.ndjson` (one JSON object per page-driving call: target tag/id/role/accessibleName/attributes/xpath/bounding-rect). The descriptors file feeds downstream codegen; it is **not** required for hypothesis formation — skip it when reading the trace.
### Read the trace
```bash
cat ./autobrowse/traces/<task-name>/latest/summary.md
```
The summary has duration, cost, turns, the decision log, and the final JSON output.
If the agent failed or got stuck, look deeper:
- Read `./autobrowse/traces/<task-name>/latest/trace.json` — search for the failure turn
- Read screenshots around the failure point with the Read tool
**When `--browser-trace` was used — start with `unified-events.jsonl`.** The harness joins the agent's turn log and the browser's CDP firehose into one time-ordered NDJSON stream at the run root. One file, source-tagged (`source: "agent" | "browser"`), interleaved by wall-clock timestamp. Skim it top-to-bottom; the failure cause is usually one or two adjacent lines (the agent issued command X, the browser responded with Y).
```bash
cat ./autobrowse/traces/<task-name>/latest/unified-events.jsonl
```
The structured files (`trace.json`, `.o11y/<run-id>/cdp/*`) are **also agent-consumable as drill-downs** when the unified stream points at something you need more of:
| Need | Drill-down file or command |
|---|---|
| Per-page totals + timing (events, network counts, errors by page) | `.o11y/<run-id>/cdp/summary.json` |
| All failed network requests in one place | `.o11y/<run-id>/cdp/network/failed.jsonl` |
| Full console exception payloads (stacktraces, etc.) | `.o11y/<run-id>/cdp/console/exceptions.jsonl` |
| Per-page slice (only events on page N) | `.o11y/<run-id>/cdp/pages/<pid>/` |
| Full reasoning text / untruncated tool outputs for a specific turn | `trace.json` (filter by `turn === N`) |
| Ad-hoc grouped query (e.g. top hosts, errors-by-page) | `O11Y_ROOT=./autobrowse/traces/<task-name>/latest/.o11y node ${CLAUDE_SKILL_DIR}/../browser-trace/scripts/query.mjs <run-id> <cmd>` |
The unified stream is the default; drill into structured files only when you need a grouped query, a full-text payload, or filtering the stream can't give you.
### Form one hypothesis
Find the exact turn where things went wrong. What single heuristic would have prevented it?
Under `--browser-trace`, the hypothesis must cite a **specific event from `unified-events.jsonl`** (line number or timestamp) — or name the drill-down file if you had to descend into one. This keeps updates evidence-grounded rather than vibes-driven. A hypothesis based only on the agent's commands might say "the click didn't work"; grounded in the unified stream, it can say "line 47 of unified-events.jsonl: `browse open` was followed by `Network.responseReceived` status 403 on `/api/checkout` — switch to `--verified --proxies`."
Examples:
- "After clicking the dropdown, wait 1s — options animate in before they're clickable"
- "Navigate directly to `/pay-invoice/` — skip the landing page entirely"
- "Use `browse fill #field_3 value` not `browse type` — this field clears on focus"
- "The page shows a spinner at turn 8 — add `browse wait timeout 2000` before snapshot"
- (with `--browser-trace`) "At line 47 of unified-events.jsonl, 3 consecutive `Network.responseReceived` events on `/api/availability` returned 403 right after `browse open` — the site is fingerprinting; the next iter needs `--verified --proxies`."
### Update strategy.md
Edit `./autobrowse/tasks/<task-name>/strategy.md`. Keep everything that worked. Fix the specific failure. Add a concrete heuristic.
Good strategies have:
- **Fast path**: direct URL or shortcuts to skip exploration
- **Step-by-step workflow**: exact sequence with timing notes
- **Site-specific knowledge**: selector IDs, form field names, success indicators
- **Failure recovery**: what to do when X goes wrong
### Judge the result
Read the new summary. Did it pass? Make clear progress?
- **Pass or progress** → keep, next iteration
- **No progress or regression** → revert strategy.md to the previous version and try a different hypothesis
### Generate a runnable script (optional)
Once the task has converged, you can produce a deterministic, runnable script
in one or more frameworks via `scripts/codegen.mjs`. This is one shot of an
LLM call per framework, cached by content hash, with optional verify-against-
fresh-session and rewrite-on-failure.
```bash
node ${CLAUDE_SKILL_DIR}/scripts/codegen.mjs \
--task <name> \
--workspace ./autobrowse \
--frameworks playwright,stagehand \
--verify
```
Each framework gets its own subdirectory under `tasks/<name>/<framework>/`
with the emitted script and a self-contained scaffold (`package.json`,
`tsconfig.json`). The directory is runnable standalone with
`cd tasks/<name>/playwright && npm install && npx tsx <name>.ts` — the only
runtime requirement is `BROWSERBASE_API_KEY` (plus `ANTHROPIC_API_KEY` for
the Stagehand target).
Builtin frameworks: `playwright`, `stagehand`. Add a custom framework with
`--prompt-template <path> --frameworks custom` (and provide your own runner
or pass `--no-verify`).
Common flags:
| Flag | Purpose |
|---|---|
| `--frameworks a,b,...` | Comma-separated; default `playwright` |
| `--verify` / `--no-verify` | Run the produced script against a fresh BB session; default `--verify` |
| `--max-retries N` | Rewrite-on-verify-failure cap; default 2 |
| `--cache-only` | Error if cache miss (CI-friendly) |
| `--force` | Bust the cache |
| `--dry-run` | Estimate prompt size + cost; don't call the LLM |
| `--run <id>` | Force a specific `run-NNN` (default: latest passing) |
Output is one JSON line per framework on stdout. Non-zero exit if any
selected framework's final state is `passed: false`.
See `references/playwright-cdp-bridge.md` for the canonical
`connectOverCDP` patterns the emitted scripts follow.
### After all iterations — publish if ready
If the task passed on 2+ of the last 3 iterations **or has reached the max iteration limit**, install it as a Claude Code skill. **Do not just copy strategy.md** — the skill must be self-contained and useful to someone who has never seen this codebase. If graduating at max iterations without a clean pass, note the known failure point but still document everything learned.
Install by writing to `~/.claude/skills/<task-name>/SKILL.md`:
```bash
mkdir -p ~/.claude/skills/<task-name>
```
Use this structure for the SKILL.md:
```markdown
---
name: <task-name>
description: <1-2 sentences describing what this skill does and when to use it. Include trigger keywords.>
---
# <Task Title> — Browser Skill
## Purpose
<1-2 sentences: what this automates and why it exists.>
## When to Use
<When should someone reach for this skill.>
## Browse CLI Reference
The inner agent uses the `browse` CLI. Key commands for this task:
- `browse stop` — kill existing session (always run before switching to remote)
- `browse open <url> --remote` — start a fresh Browserbase cloud session and navigate
- `browse open <url> --local` — start a clean local browser and navigate
- `browse tab new <url>` — open URL in a new tab
- `browse wait load` — wait for page to finish loading
- `browse wait timeout <ms>` — wait a fixed amount of time for spinners or animations
- `browse wait selector "<selector>"` — wait for an element to become visible
- `browse get title` — verify you're on the right page
- `browse get text body` — extract all visible text (preferred for content extraction)
- `browse snapshot` — get accessibility tree; each node has a ref in `[X-Y]` format (e.g. `[0-5]`, `[2-147]`)
- `browse click [X-Y]` — click element by ref from the latest snapshot (include the brackets)
**Never use `--session <name>` flags in SKILL.md.** Named sessions are a parallel-run workaround — they contaminate skills with infrastructure concerns. Skills must work in isolation with the default session.
## Workflow
### Step 1 — Start session
<exact browse commands in order>
### Step 2 — Navigate
<exact URL and verification steps>
### Step 3 — Extract
<exact extraction commands>
### Step 4 — Output
<what JSON to emit, referencing the schema below>
## Site-Specific Gotchas
<Bullet list of every hard-won heuristic from the iterations. This is the core value of the skill.>
## Failure Recovery
<What to do when navigation fails, session is contaminated, or extraction returns garbage>
## Expected Output
```json
<paste the exact expected output schema from task.md>
```
```
After writing the SKILL.md, confirm it's installed:
```bash
ls ~/.claude/skills/<task-name>/SKILL.md
```
The skill is now available as `/<task-name>` in Claude Code.
---
## Final report (multi-task mode)
After all sub-agents complete, print a markdown table:
| Task | Iterations | Final Status | Graduated | Cost |
|------|-----------|--------------|-----------|------|
| google-flights | 5 | ✅ pass | yes | $0.42 |
| amazon-add-to-cart | 5 | ❌ fail | no | $1.20 |
Then write a persistent session report to `./autobrowse/reports/` so there's a durable record of the run inside the workspace:
```bash
mkdir -p ./autobrowse/reports
```
Write the file `./autobrowse/reports/YYYY-MM-DD-HH-MM-<tasks>.md` with:
```markdown
# AutoBrowse Session Report
**Date:** <ISO date>
**Tasks:** <comma-separated list>
**Environment:** remote|local
**Total cost:** $X.XX
## Results
| Task | Iterations | Pass Rate | Final Status | Graduated | Cost |
|------|-----------|-----------|--------------|-----------|------|
| ... | ... | X/5 | ✅/❌ | yes/no | $X.XX |
## Per-Task Learnings
### <task-name>
- **Key insight 1:** <what the agent learned>
- **Key insight 2:** <another heuristic>
- **Failure mode fixed:** <what was failing and how it was resolved>
## Iteration Log
### <task-name>
| Iter | Turns | Cost | Status | Hypothesis tested |
|------|-------|------|--------|-------------------|
| 1 | 79 | $18.75 | ❌ fail | baseline |
| 2 | 9 | $0.26 | ✅ pass | session contamination fix |
| ... | ... | ... | ... | ... |
```
---
## Rules
- **Only edit `strategy.md`** — never touch `task.md` (unless creating it from the template) or `evaluate.mjs`
- **Stay in the workspace** — all training writes go to `./autobrowse/`, never to `~/.claude/skills/autobrowse/`. The skill source is read-only.
- **One hypothesis per iteration** — test one change at a time
- **Build on wins** — keep what worked, add to it
- **Trust the trace** — the inner agent shows exactly what it saw and did
- **Graduate to `~/.claude/skills/`** — the only file you write there is the final graduated `SKILL.md`
- **Don't release before bisecting** — under `--browser-trace`, the order at the end of each iteration is non-negotiable: `stop-capture` → `bisect-cdp` → `browse cloud sessions update REQUEST_RELEASE`. Bisect depends on the session still existing when the trace stops.





首页
