autobrowse
browserbase/skills
透過反覆實驗來建立可靠的瀏覽器自動化技能,運作內部代理程式來瀏覽網站,並不斷改進導航指令,直到任務能穩定通過為止。
...展開全部AutoBrowse — 自我提升的瀏覽器自動化技能
透過反覆實驗,建立可靠的瀏覽器自動化技能。一個「內部代理」負責瀏覽網站(evaluate.ts)。而您——作為「外部代理」——則會分析發生了什麼情況,並據此優化操作指令(strategy.md)。重複此過程,直到系統能穩定通過測試為止。
執行入口
呼叫方式相當靈活——無論是顯式參數或自由形式的自然語言皆可:
/autobrowse --task google-flights
/autobrowse --task google-flights --iterations 10 --env remote
/autobrowse --task google-flights --browser-trace
/autobrowse --tasks google-flights,amazon-add-to-cart
/autobrowse --all
# 以下方式亦可 — 可自由解析:
/autobrowse https://flights.google.com/
/autobrowse 在 delta.com 預訂航班
/autobrowse 修正現有的 google-flights 技能
--browser-trace(預設為關閉,僅限遠端):將每次迭代與同級的browser-trace技能配對 — 將內層代理封裝在 CDP 擷取中,以取得每頁的網路/主控台/頁面生命週期證據。 隱含--env remote;若與--env local 併用則會發生錯誤。需在${CLAUDE_SKILL_DIR}/../browser-trace/ 目錄中存在同級的browser-trace技能,並需設定BROWSERBASE_API_KEY環境變數。
當使用者輸入 URL 或自由格式指令,而非使用--task 時:
- 若
${WORKSPACE}/tasks/中已有任務明確符合該網站/意圖,則使用該任務。 - 否則,選擇一個簡短的 kebab-case 命名格式,從
${CLAUDE_SKILL_DIR}/references/example-task.md建立${WORKSPACE}/tasks/,根據用戶的發言填入網址/目標,並繼續執行。 請以一行文字告知使用者所選的名稱。/task.md 檔案
執行方式
步驟 1 — 解析參數並進行定位
檢查傳入的參數:
--task→ 單一任務模式--tasks a,b,c或--all→ 多任務模式(啟動子代理)--iterations N→ 評估 → 優化循環的次數(預設:5)--env local|remote→ 瀏覽器環境(預設:local;針對受機器人防護的網站請使用 remote)--browser-trace→ 啟用瀏覽器追蹤整合功能(預設為關閉)。 此選項隱含--env remote。若同時明確傳入--env local 與 --browser-trace,將顯示錯誤訊息:browser-trace 需要 Browserbase;請移除 --env local 或移除 --browser-trace。
若使用者傳入的是自由格式文字,請在繼續執行前將其映射至上述選項之一。
步驟 2 — 設定工作區
所有訓練產出(任務定義、策略迭代、追蹤記錄、報告)皆存放於當前工作目錄中的工作區目錄內 — 並非位於~/.claude/skills/ 內。此舉可避免內部代理的檔案寫入操作進入 Claude 的家目錄,並避免權限衝突。
預設工作區:${CWD}/autobrowse/
mkdir -p ./autobrowse/tasks ./autobrowse/traces ./autobrowse/reports
若任務目錄(./autobrowse/tasks/)尚未存在,請建立其骨架:
mkdir -p ./autobrowse/tasks/
cp ${CLAUDE_SKILL_DIR}/references/example-task.md ./autobrowse/tasks//task.md
# 接著編輯 task.md 以描述 URL、輸入、步驟及預期的 JSON 輸出
位於${CLAUDE_SKILL_DIR}的技能原始檔保持唯讀狀態——僅在當前工作目錄 (CWD) 中的./autobrowse/會於訓練期間被寫入。結業(最後一步)會將單一檔案寫入~/.claude/skills/。
列出可用任務:
ls ./autobrowse/tasks/
步驟 3 — 多任務:啟動並行子代理
若執行多個任務,請使用 Agent 工具,針對每個任務同時啟動一個子代理。每個子代理會收到一個獨立的提示,用以執行其任務的完整autobrowse 迴圈:
autobrowse 「您正在執行任務
。工作區:(例如/path/to/project/autobrowse)。執行以下循環:評估 → 讀取追蹤紀錄 → 改善策略.md → 重複。 請使用--env。對每次 evaluate.mjs 呼叫,皆須傳入--workspace。若父級呼叫使用了--browser-trace,您必須在每次迭代中使用 SKILL.md 迴圈中的 traced-path 區塊(預先建立會話、附加 bb-capture、傳遞--connect-url傳遞給 evaluate.mjs、停止並二分法、發布)——切勿回退至預設的單一指令路徑。請嚴格遵循autobrowse 中的迴圈指示。完成訓練後,請將技能安裝至
~/.claude/skills/前置資訊 (agentskills) 包含正確的名稱與描述。請勿僅複製 strategy.md —— 應撰寫一個自成一體的技能。/SKILL.md,並確保 最後,請輸出結構化摘要,內容應包含:任務名稱、最終執行結果(通過/失敗)、累計總成本、已完成的迭代次數、每迭代資料表(迭代編號、回合數、成本、狀態、測試假設),以及 2 至 3 項關鍵學習要點。
並行啟動所有子代理,待全部完成後,收集其摘要並撰寫會話報告。
若為單一任務,請跳過此步驟並直接執行下方的迴圈。
迴圈(針對每個任務執行此迴圈)
迭代開始
檢查./autobrowse/tasks/是否存在(若不存在,請根據範本建立骨架 — 參見步驟 2)。strategy.md會在首次執行時由測試框架自動建立為空檔。
需求
- 環境變數中必須有
ANTHROPIC_API_KEY(或位於當前工作目錄(CWD)的.env檔案中 —evaluate.mjs會自動載入該檔案)。若缺失,測試框架會輸出明確的錯誤訊息並退出;請勿在其他路徑中搜尋金鑰。
執行內部代理
預設路徑(未指定--browser-trace)— 單一指令,無協調機制:
node ${CLAUDE_SKILL_DIR}/scripts/evaluate.mjs --task --workspace ./autobrowse
# 或針對受機器人保護的網站:
node ${CLAUDE_SKILL_DIR}/scripts/evaluate.mjs --task --workspace ./autobrowse --env remote
這會執行瀏覽器會話,並將完整的追蹤紀錄寫入./autobrowse/traces/
追蹤路徑(--browser-trace,僅限遠端)—— 外層測試框架會預先建立一個 Browserbase 會話,將bb-capture附加為被動觀察者,並將該會話的connectUrl傳遞給evaluate.mjs,以便每個內層的browse呼叫皆使用--cdp $connectUrl --sessionautobrowse-main(標準的 browser-trace 模式,可讓觀察者取得完整的 Network/Console 事件)。每輪迭代執行此區塊一次,並將$N設為從 1 開始的迭代序號:
# 預檢 — 若未與autobrowse 一同安裝 browser-trace,則立即終止。
BT_DIR="${CLAUDE_SKILL_DIR}/../browser-trace"
if [ ! -f "$BT_DIR/scripts/bb-capture.mjs" ]; then
echo "錯誤:--browser-trace 需要位於 $BT_DIR 的 browser-trace 技能。" >&2
echo "請透過克隆 github.com/browserbase/skills 並將 skills/browser-trace/" >&2
echo "複製到與autobrowse 相同的父目錄中(例如 ~/.claude/skills/browser-trace/)。" >&2
exit 1
fi
# a. 會話設定 — 預先建立保持連線的會話並推導其 connectUrl
sid=$(browse cloud sessions create --keep-alive --verified --proxies \
| node -e "let s='';process.stdin.on('data',c=>s+=c).on('end',()=>process.stdout.write(JSON.parse(s).id))")
connect_url=$(browse cloud sessions get "$sid" \
| node -e "let s='';process.stdin.on('data',c=>s+=c).on('end',()=>process.stdout.write(JSON.parse(s).connectUrl))")
RUN_ID="run-$(printf '%03d' "$N")"
TRACE_ROOT="./autobrowse/traces//$RUN_ID"
mkdir -p "$TRACE_ROOT"
export O11Y_ROOT="$TRACE_ROOT/.o11y" # 將瀏覽器追蹤輸出儲存至autobrowse 執行目錄中
export O11Y_RUN_ID="$RUN_ID" # 告知 browse CLI 應將 descriptors.ndjson 寫入哪個執行目錄
# b. 附加瀏覽器追蹤 — 被動觀察者;於背景執行
node ${CLAUDE_SKILL_DIR}/../browser-trace/scripts/bb-capture.mjs "$sid" "$RUN_ID" &
sleep 2
# c. 執行 `AUTOBROWSE ` — `connectUrl` 旗標指示 `evaluate.mjs` 將 `--cdp/--session`
# 注入每個內層的 `browse` 呼叫中。內層代理程式永遠不會看到 `--remote`。
node ${CLAUDE_SKILL_DIR}/scripts/evaluate.mjs \
--task --workspace ./autobrowse --env remote \
--connect-url "$connect_url" --run-number "$N"
# d. STOP + BISECT + UNIFY —— 執行順序至關重要;bisext 需要該會話仍
# 存在,而 unify-trace 會將 bisect 的輸出與autobrowse 的 trace.json 合併
# 並將其整合為單一按時間排序的 NDJSON,外層代理程式會在每次迭代時優先讀取此檔案。
node ${CLAUDE_SKILL_DIR}/../browser-trace/scripts/stop-capture.mjs "$RUN_ID"
node ${CLAUDE_SKILL_DIR}/../browser-trace/scripts/bisect-cdp.mjs "$RUN_ID"
node ${CLAUDE_SKILL_DIR}/scripts/unify-trace.mjs \
--trace-dir "$TRACE_ROOT" \
--o11y-dir "$O11Y_ROOT/$RUN_ID"
# e. 發佈
browse cloud sessions update "$sid" --status REQUEST_RELEASE
這會將代理內部追蹤記錄寫入./autobrowse/traces/,並將 CDP 雙分法追蹤寫入./autobrowse/traces/。 被追蹤的瀏覽CLI 還會將每項指令的豐富節點描述符輸出至.o11y/(每個驅動頁面的呼叫對應一個 JSON 物件:target tag/id/role/accessibleName/attributes/xpath/bounding-rect)。 該描述子檔案用於饋送下游的程式碼生成模組;對於假設建立並非必要 — 閱讀追蹤記錄時可跳過此部分。
讀取追蹤記錄
cat ./autobrowse/traces//latest/summary.md
摘要包含執行時間、成本、回合數、決策日誌以及最終的 JSON 輸出。
若代理程式失敗或陷入僵局,請進一步查閱:
- 讀取
./autobrowse/traces/— 搜尋發生失敗的回合/latest/trace.json - 使用 Read 工具檢視失敗點附近的螢幕截圖
若使用了--browser-trace 選項— 請從unified-events.jsonl 開始分析。測試框架會將代理的回合日誌與瀏覽器的 CDP 數據流,整合為一個按時間排序的 NDJSON 資料流,並存放於執行根目錄中。 這是一個帶有來源標籤(source: "agent" | "browser")的檔案,並以實際時間戳記交錯排列。請從上至下快速瀏覽;失敗原因通常出現在一兩行相鄰的記錄中(代理發送指令 X,瀏覽器回應 Y)。
cat ./autobrowse/traces//latest/unified-events.jsonl
當統一資料流指向您需要進一步分析的內容時,結構化檔案(trace.json、.o11y/)也可作為代理程式可用的深入分析資料:
| 需求 | 所需的深入分析檔案或指令 |
|---|---|
| 每頁總計 + 時間統計(事件、網路計數、各頁錯誤) | .o11y/ |
| 所有失敗的網路請求集中於一處 | .o11y/ |
| 完整的控制台例外資料(堆疊追蹤等) | .o11y/ |
| 按頁面切片(僅顯示第 N 頁的事件) | .o11y/ |
| 特定回合的完整推理文字/未截斷的工具輸出 | trace.json(篩選條件:回合 === N) |
| 臨時分組查詢(例如:前幾名主機、按頁面分類的錯誤) | O11Y_ROOT=./autobrowse/traces/ |
統一資料流為預設設定;僅當您需要分組查詢、全文載入,或資料流無法滿足的篩選需求時,才需深入分析結構化檔案。
提出一個假設
找出問題確切發生的轉折點。若採用哪一項單一啟發式方法,本可避免此狀況?
在--browser-trace 模式下,假設必須引用 unified-events.jsonl中的特定事件(行號或時間戳記)——若您必須深入檢視某個檔案,則需指名該深入檢視檔案。此舉可確保更新基於實證,而非憑直覺。 若假設僅基於代理程式(agent)的指令,可能會說「點擊無效」; 若以統一事件流為依據,則可表述為:「unified-events.jsonl 第 47 行:在『browse open』之後,/api/checkout傳回Network.responseReceived狀態碼 403 — 請切換至--verified --proxies。」
範例:
- 「點擊下拉選單後,等待 1 秒 — 選項會先以動畫形式顯示,之後才可點擊」
- 「直接導航至
/pay-invoice/— 完全跳過登陸頁面」 - 「使用『
瀏覽填入 #field_3 值』而非『瀏覽類型』— 此欄位在獲得焦點時會清空」 - 「頁面在第 8 步顯示轉圈圖示 — 在擷取快照前,
為瀏覽操作新增2000 毫秒的等待超時」 - (搭配
--browser-trace參數) 「在 unified-events.jsonl 的第 47 行,瀏覽視窗開啟後,針對/api/availability的 3 個連續Network.responseReceived事件均返回 403 錯誤 — 該網站正在進行指紋識別;下一次迭代需使用--verified --proxies參數。」
更新 strategy.md
編輯./autobrowse/tasks/。保留所有運作正常的內容。修正特定的失敗情況。新增具體的啟發式規則。
良好的策略應具備:
- 快速路徑:直接網址或捷徑以跳過探索
- 逐步工作流程:附有時間註記的精確順序
- 網站特定知識:選取器 ID、表單欄位名稱、成功指標
- 失敗恢復:當 X 出錯時該如何處理
評估結果
閱讀新的摘要。是否通過?是否有明確的進展?
- 通過或有進展→ 保留,進入下一輪迭代
- 無進展或退步→ 將 strategy.md 還原至前一版本,並嘗試不同的假設
產生可執行的腳本(可選)
當任務收斂後,您可以透過scripts/codegen.mjs
在一個或多個框架中產生確定性且可執行的腳本。這代表每個框架僅進行一次
大型語言模型(LLM)呼叫,並根據內容雜湊值進行快取,同時可選配
「與最新會話進行驗證」及「失敗時重新撰寫」功能。
node ${CLAUDE_SKILL_DIR}/scripts/codegen.mjs \
--task \
--workspace ./autobrowse \
--frameworks playwright,stagehand \
--verify
每個框架都會在tasks/、之下
擁有各自的子目錄,其中包含生成的腳本以及一個自包含的基礎架構(package.json、
tsconfig.json)。 該目錄可透過以下指令獨立執行:
cd tasks/— 唯一的
執行時需求是BROWSERBASE_API_KEY(若針對 Stagehand 目標,則需額外提供
ANTHROPIC_API_KEY)。
內建框架:playwright、stagehand。若要新增自訂框架,請使用
--prompt-template(並提供您自己的執行程式
或傳入--no-verify)。
常用參數:
| 參數 | 用途 |
|---|---|
--frameworks a,b,... |
以逗號分隔;預設為playwright |
--verify/--no-verify |
在全新的 BB 工作階段中執行產出的腳本;預設為--verify |
--max-retries N |
驗證失敗時的重寫上限;預設值為 2 |
--僅快取 |
若快取未命中則報錯(適合持續整合) |
--force |
清除快取 |
--dry-run |
估算提示字串大小與運算成本;不呼叫大型語言模型 |
--run |
強制執行特定的運行-NNN(預設:最新通過的) |
輸出為每個框架一行 JSON 格式,輸出至標準輸出(stdout)。若任何
選定的框架之最終狀態為「false」,則返回非零退出狀態。
請參閱references/playwright-cdp-bridge.md,了解發出的腳本所遵循的標準
connectOverCDP模式。
所有迭代結束後 — 若已準備就緒則發布
若該任務在最近 3 次迭代中至少有 2 次通過,或已達到最大迭代次數限制,則將其安裝為 Claude Code 技能。請勿僅複製 strategy.md— 該技能必須是自包含的,且對從未接觸過此程式碼庫的人而言仍具實用價值。若在達到最大迭代次數時仍未完全通過測試,請記錄已知的失敗點,但仍需完整記錄所有學習成果。
透過在~/.claude/skills/ 寫入內容來安裝:
mkdir -p ~/.claude/skills/
SKILL.md 應採用以下結構:
---
name:
description:<1-2 sentences describing what this skill does and when to use it. Include trigger keywords.>
---
# — 瀏覽器技能
## 目的
<1-2 sentences: what this automates and why it exists.>
## 何時使用
## 瀏覽 CLI 參考手冊
內部代理程式使用 `browse` CLI。此任務的關鍵指令:
- `browse stop` — 終止現有工作階段(切換至遠端前務必執行)
- `browse open --remote` — 啟動全新的 Browserbase 雲端工作階段並進行瀏覽
- `browse open --local` — 啟動乾淨的本地瀏覽器並進行瀏覽
- `browse tab new` — 在新分頁中開啟網址
- `browse wait load` — 等待頁面載入完成
- `browse wait timeout` — 等待固定時間,以觀察旋轉圖示或動畫效果
- `browse wait selector ""` — 等待某個元素顯示出來
- `browse get title` — 驗證是否位於正確的頁面
- `browse get text body` — 擷取所有可見文字(建議用於內容擷取)
- `browse snapshot` — 取得輔助技術樹;每個節點皆有格式為 `[X-Y]` 的參考值(例如 `[0-5]`、`[2-147]`)
- `browse click [X-Y]` — 根據最新快照中的參考值點擊元素 (請包含方括號)
**切勿在 SKILL.md 中使用 `--session` 參數。** 命名會話是一種並行執行的變通方法 —— 它們會讓技能混入基礎架構相關的問題。技能必須在預設會話中獨立運作。
## 工作流程
### 步驟 1 — 啟動會話
### 步驟 2 — 導航
### 步驟 3 — 擷取
### 步驟 4 — 輸出
## 特定網站的注意事項
## 失敗恢復
## 預期輸出
```json
撰寫完 SKILL.md 後,請確認已安裝:
```bash
ls ~/.claude/skills//SKILL.md
該技能現已可在 Claude Code 中透過/存取。
最終報告(多任務模式)
當所有子代理程式執行完畢後,輸出一個 Markdown 表格:
| 任務 | 迭代次數 | 最終狀態 | 已畢業 | 成本 |
|---|---|---|---|---|
| google-flights | 5 | ✅ 通過 | 是 | 0.42 美元 |
| amazon-加入購物車 | 5 | ❌ 失敗 | 否 | $1.20 |
接著將持久性會話報告寫入./autobrowse/reports/,以便在工作區內保留此次執行的持久記錄:
mkdir -p ./autobrowse/reports
建立檔案./autobrowse/reports/YYYY-MM-DD-HH-MM-,內容如下:
#AutoBrowse 執行報告
**日期:**
**任務:**
**環境:** 遠端|本地
**總成本:** $X.XX
## 結果
| 任務 | 迭代次數 | 通過率 | 最終狀態 | 是否通過 | 成本 |
|------|-----------|-----------|--------------|-----------|------|
| ... | ... | X/5 | ✅/❌ | 是/否 | $X.XX |
## 各任務心得
###
- **關鍵洞見 1:**
- **關鍵洞見 2:**
- **已修正的失敗模式:**
## 迭代紀錄
###
| 迭代 | 輪次 | 成本 | 狀態 | 測試假設 |
|------|-------|------|--------|-------------------|
| 1 | 79 | $18.75 | ❌ 失敗 | 基準 |
| 2 | 9 | $0.26 | ✅ 通過 | 修正會話污染 |
| ... | ... | ... | ... | ... |
規則
- 僅編輯
strategy.md— 切勿修改task.md(除非是從範本建立該檔案)或evaluate.mjs - 請留在工作區內— 所有訓練寫入皆存至
./autobrowse/,絕不存至~/.claude/skills/autobrowse/。技能原始碼為唯讀。 - 每次迭代僅測試一個假設— 每次只測試一項變更
- 建立在成功之上— 保留有效的方法,並在此基礎上加以擴展
- 信任追蹤紀錄— 內部代理會精確顯示其所見與所為
- 晉級至
~/.claude/skills/—— 您在此目錄中唯一會寫入的檔案,是最終晉級的SKILL.md - 在進行二分法之前切勿發布—— 在
--browser-trace模式下,每次迭代結束時的順序不可更改:停止擷取→二分法分析 CDP→瀏覽 Cloud 會話並更新 REQUEST_RELEASE。二分法分析取決於追蹤停止時會話是否仍然存在。
---
name: autobrowse
description: Builds reliable browser automation skills through iterative experimentation, running an inner agent to browse sites and improving navigation instructions until tasks pass consistently.
license: MIT
---
# AutoBrowse — Self-Improving Browser Skill
Build reliable browser automation skills through iterative experimentation. An inner agent browses the site (`evaluate.ts`). You — the outer agent — read what happened and improve the instructions (`strategy.md`). Repeat until it passes consistently.
## Entry Points
Invocation is flexible — both explicit flags and free-form natural language work:
```
/autobrowse --task google-flights
/autobrowse --task google-flights --iterations 10 --env remote
/autobrowse --task google-flights --browser-trace
/autobrowse --tasks google-flights,amazon-add-to-cart
/autobrowse --all
# Also fine — parse freely:
/autobrowse https://flights.google.com/
/autobrowse book a flight on delta.com
/autobrowse fix the existing google-flights skill
```
`--browser-trace` (default off, remote-only): pairs each iteration with the sibling `browser-trace` skill — wraps the inner agent in a CDP capture for per-page network/console/page-lifecycle evidence. Implies `--env remote`; errors if combined with `--env local`. Requires the sibling `browser-trace` skill present at `${CLAUDE_SKILL_DIR}/../browser-trace/`, and the `BROWSERBASE_API_KEY` env var.
When the user drops a URL or free-form instruction instead of `--task <name>`:
- If an existing task in `${WORKSPACE}/tasks/` clearly matches the site/intent, use it.
- Otherwise, pick a short kebab-case name, create `${WORKSPACE}/tasks/<name>/task.md` from `${CLAUDE_SKILL_DIR}/references/example-task.md`, fill in the URL/goal based on what the user said, and proceed. Tell the user the chosen name in one line.
---
## How to run
### Step 1 — Parse arguments and orient
Check what was passed:
- `--task <name>` → single task mode
- `--tasks a,b,c` or `--all` → multi-task mode (spawn sub-agents)
- `--iterations N` → how many evaluate → improve cycles (default: 5)
- `--env local|remote` → browser environment (default: local; use remote for bot-protected sites)
- `--browser-trace` → opt in to the browser-trace integration (default off). Implies `--env remote`. If `--env local --browser-trace` are both passed explicitly, error with: `browser-trace requires Browserbase; drop --env local or drop --browser-trace.`
If the user passed free-form text instead, map it to one of the above before continuing.
### Step 2 — Set up the workspace
All training artifacts (task definitions, strategy iterations, traces, reports) live in a workspace directory in the **current working directory** — NOT inside `~/.claude/skills/`. This keeps the inner agent's file writes out of Claude's home dir and away from permission friction.
Default workspace: `${CWD}/autobrowse/`
```bash
mkdir -p ./autobrowse/tasks ./autobrowse/traces ./autobrowse/reports
```
If the task directory (`./autobrowse/tasks/<task>/task.md`) doesn't exist yet, scaffold it:
```bash
mkdir -p ./autobrowse/tasks/<task>
cp ${CLAUDE_SKILL_DIR}/references/example-task.md ./autobrowse/tasks/<task>/task.md
# Then edit task.md to describe the URL, inputs, steps, and expected JSON output
```
The skill source at `${CLAUDE_SKILL_DIR}` stays read-only — only `./autobrowse/` in CWD gets written to during training. Graduation (final step) writes a single file to `~/.claude/skills/<task>/SKILL.md`.
List available tasks:
```bash
ls ./autobrowse/tasks/
```
### Step 3 — Multi-task: spawn parallel sub-agents
If running multiple tasks, use the Agent tool to spawn one sub-agent per task simultaneously. Each sub-agent receives a self-contained prompt to run the full autobrowse loop for its task:
> "You are running the autobrowse skill for task `<name>`. Workspace: `<absolute-path-to-workspace>` (e.g. `/path/to/project/autobrowse`). Run `<N>` iterations of: evaluate → read trace → improve strategy.md → repeat. Use `--env <env>`. Pass `--workspace <workspace>` to every evaluate.mjs invocation. If the parent invocation used `--browser-trace`, you MUST use the traced-path block of the SKILL.md loop for every iteration (pre-create session, attach bb-capture, pass `--connect-url` to evaluate.mjs, stop+bisect, release) — do not fall back to the default single-command path. Follow the autobrowse loop instructions exactly.
>
> When graduating, install the skill to `~/.claude/skills/<task-name>/SKILL.md` with proper agentskills frontmatter (name + description). Do not just copy strategy.md — write a self-contained skill.
>
> At the end, output a structured summary with: task name, pass/fail on final run, total cumulative cost, iterations completed, per-iteration table (iter number, turns, cost, status, hypothesis tested), and 2-3 bullet key learnings."
Spawn all sub-agents in parallel, wait for all to complete, then collect their summaries and write the session report.
**For single task**, skip this step and run the loop directly below.
---
## The Loop (run this for each task)
### Iteration start
Check that `./autobrowse/tasks/<task>/task.md` exists (scaffold it from the template if not — see Step 2). `strategy.md` is auto-created empty by the harness on first run.
### Requirements
- `ANTHROPIC_API_KEY` must be in the environment (or in a `.env` file in CWD — `evaluate.mjs` auto-loads it). If missing, the harness prints a clear error and exits; don't hunt for keys in other paths.
### Run the inner agent
**Default path (no `--browser-trace`)** — single command, no orchestration:
```bash
node ${CLAUDE_SKILL_DIR}/scripts/evaluate.mjs --task <task-name> --workspace ./autobrowse
# or for bot-protected sites:
node ${CLAUDE_SKILL_DIR}/scripts/evaluate.mjs --task <task-name> --workspace ./autobrowse --env remote
```
This runs the browser session and writes a full trace to `./autobrowse/traces/<task>/latest/`.
**Traced path (`--browser-trace`, remote only)** — the outer harness pre-creates a Browserbase session, attaches `bb-capture` as a passive observer, and passes the session's `connectUrl` to `evaluate.mjs` so every inner `browse` call uses `--cdp $connectUrl --session autobrowse-main` (the canonical browser-trace pattern that gives observers full Network/Console events). Run this block once per iteration with `$N` set to the 1-indexed iteration number:
```bash
# Preflight — fail fast if browser-trace isn't installed alongside autobrowse.
BT_DIR="${CLAUDE_SKILL_DIR}/../browser-trace"
if [ ! -f "$BT_DIR/scripts/bb-capture.mjs" ]; then
echo "ERROR: --browser-trace requires the browser-trace skill at $BT_DIR." >&2
echo "Install it by cloning github.com/browserbase/skills and copying skills/browser-trace/" >&2
echo "into the same parent directory as autobrowse (e.g. ~/.claude/skills/browser-trace/)." >&2
exit 1
fi
# a. SESSION SETUP — pre-create the keep-alive session and derive its connectUrl
sid=$(browse cloud sessions create --keep-alive --verified --proxies \
| node -e "let s='';process.stdin.on('data',c=>s+=c).on('end',()=>process.stdout.write(JSON.parse(s).id))")
connect_url=$(browse cloud sessions get "$sid" \
| node -e "let s='';process.stdin.on('data',c=>s+=c).on('end',()=>process.stdout.write(JSON.parse(s).connectUrl))")
RUN_ID="run-$(printf '%03d' "$N")"
TRACE_ROOT="./autobrowse/traces/<task-name>/$RUN_ID"
mkdir -p "$TRACE_ROOT"
export O11Y_ROOT="$TRACE_ROOT/.o11y" # park browser-trace output inside the autobrowse run dir
export O11Y_RUN_ID="$RUN_ID" # tells the browse CLI which run dir to write descriptors.ndjson into
# b. ATTACH BROWSER-TRACE — passive observer; runs in background
node ${CLAUDE_SKILL_DIR}/../browser-trace/scripts/bb-capture.mjs "$sid" "$RUN_ID" &
sleep 2
# c. RUN AUTOBROWSE — connectUrl flag tells evaluate.mjs to inject --cdp/--session
# into every inner browse call. The inner agent never sees --remote.
node ${CLAUDE_SKILL_DIR}/scripts/evaluate.mjs \
--task <task-name> --workspace ./autobrowse --env remote \
--connect-url "$connect_url" --run-number "$N"
# d. STOP + BISECT + UNIFY — order matters; bisect needs the session to still
# exist, and unify-trace joins the bisect output with autobrowse's trace.json
# into a single time-ordered NDJSON the outer agent reads first each iter.
node ${CLAUDE_SKILL_DIR}/../browser-trace/scripts/stop-capture.mjs "$RUN_ID"
node ${CLAUDE_SKILL_DIR}/../browser-trace/scripts/bisect-cdp.mjs "$RUN_ID"
node ${CLAUDE_SKILL_DIR}/scripts/unify-trace.mjs \
--trace-dir "$TRACE_ROOT" \
--o11y-dir "$O11Y_ROOT/$RUN_ID"
# e. RELEASE
browse cloud sessions update "$sid" --status REQUEST_RELEASE
```
This writes the inner-agent trace to `./autobrowse/traces/<task-name>/latest/` and the CDP bisect to `./autobrowse/traces/<task-name>/latest/.o11y/<run-id>/`. The traced `browse` CLI also emits per-command rich node descriptors to `.o11y/<run-id>/cdp/descriptors.ndjson` (one JSON object per page-driving call: target tag/id/role/accessibleName/attributes/xpath/bounding-rect). The descriptors file feeds downstream codegen; it is **not** required for hypothesis formation — skip it when reading the trace.
### Read the trace
```bash
cat ./autobrowse/traces/<task-name>/latest/summary.md
```
The summary has duration, cost, turns, the decision log, and the final JSON output.
If the agent failed or got stuck, look deeper:
- Read `./autobrowse/traces/<task-name>/latest/trace.json` — search for the failure turn
- Read screenshots around the failure point with the Read tool
**When `--browser-trace` was used — start with `unified-events.jsonl`.** The harness joins the agent's turn log and the browser's CDP firehose into one time-ordered NDJSON stream at the run root. One file, source-tagged (`source: "agent" | "browser"`), interleaved by wall-clock timestamp. Skim it top-to-bottom; the failure cause is usually one or two adjacent lines (the agent issued command X, the browser responded with Y).
```bash
cat ./autobrowse/traces/<task-name>/latest/unified-events.jsonl
```
The structured files (`trace.json`, `.o11y/<run-id>/cdp/*`) are **also agent-consumable as drill-downs** when the unified stream points at something you need more of:
| Need | Drill-down file or command |
|---|---|
| Per-page totals + timing (events, network counts, errors by page) | `.o11y/<run-id>/cdp/summary.json` |
| All failed network requests in one place | `.o11y/<run-id>/cdp/network/failed.jsonl` |
| Full console exception payloads (stacktraces, etc.) | `.o11y/<run-id>/cdp/console/exceptions.jsonl` |
| Per-page slice (only events on page N) | `.o11y/<run-id>/cdp/pages/<pid>/` |
| Full reasoning text / untruncated tool outputs for a specific turn | `trace.json` (filter by `turn === N`) |
| Ad-hoc grouped query (e.g. top hosts, errors-by-page) | `O11Y_ROOT=./autobrowse/traces/<task-name>/latest/.o11y node ${CLAUDE_SKILL_DIR}/../browser-trace/scripts/query.mjs <run-id> <cmd>` |
The unified stream is the default; drill into structured files only when you need a grouped query, a full-text payload, or filtering the stream can't give you.
### Form one hypothesis
Find the exact turn where things went wrong. What single heuristic would have prevented it?
Under `--browser-trace`, the hypothesis must cite a **specific event from `unified-events.jsonl`** (line number or timestamp) — or name the drill-down file if you had to descend into one. This keeps updates evidence-grounded rather than vibes-driven. A hypothesis based only on the agent's commands might say "the click didn't work"; grounded in the unified stream, it can say "line 47 of unified-events.jsonl: `browse open` was followed by `Network.responseReceived` status 403 on `/api/checkout` — switch to `--verified --proxies`."
Examples:
- "After clicking the dropdown, wait 1s — options animate in before they're clickable"
- "Navigate directly to `/pay-invoice/` — skip the landing page entirely"
- "Use `browse fill #field_3 value` not `browse type` — this field clears on focus"
- "The page shows a spinner at turn 8 — add `browse wait timeout 2000` before snapshot"
- (with `--browser-trace`) "At line 47 of unified-events.jsonl, 3 consecutive `Network.responseReceived` events on `/api/availability` returned 403 right after `browse open` — the site is fingerprinting; the next iter needs `--verified --proxies`."
### Update strategy.md
Edit `./autobrowse/tasks/<task-name>/strategy.md`. Keep everything that worked. Fix the specific failure. Add a concrete heuristic.
Good strategies have:
- **Fast path**: direct URL or shortcuts to skip exploration
- **Step-by-step workflow**: exact sequence with timing notes
- **Site-specific knowledge**: selector IDs, form field names, success indicators
- **Failure recovery**: what to do when X goes wrong
### Judge the result
Read the new summary. Did it pass? Make clear progress?
- **Pass or progress** → keep, next iteration
- **No progress or regression** → revert strategy.md to the previous version and try a different hypothesis
### Generate a runnable script (optional)
Once the task has converged, you can produce a deterministic, runnable script
in one or more frameworks via `scripts/codegen.mjs`. This is one shot of an
LLM call per framework, cached by content hash, with optional verify-against-
fresh-session and rewrite-on-failure.
```bash
node ${CLAUDE_SKILL_DIR}/scripts/codegen.mjs \
--task <name> \
--workspace ./autobrowse \
--frameworks playwright,stagehand \
--verify
```
Each framework gets its own subdirectory under `tasks/<name>/<framework>/`
with the emitted script and a self-contained scaffold (`package.json`,
`tsconfig.json`). The directory is runnable standalone with
`cd tasks/<name>/playwright && npm install && npx tsx <name>.ts` — the only
runtime requirement is `BROWSERBASE_API_KEY` (plus `ANTHROPIC_API_KEY` for
the Stagehand target).
Builtin frameworks: `playwright`, `stagehand`. Add a custom framework with
`--prompt-template <path> --frameworks custom` (and provide your own runner
or pass `--no-verify`).
Common flags:
| Flag | Purpose |
|---|---|
| `--frameworks a,b,...` | Comma-separated; default `playwright` |
| `--verify` / `--no-verify` | Run the produced script against a fresh BB session; default `--verify` |
| `--max-retries N` | Rewrite-on-verify-failure cap; default 2 |
| `--cache-only` | Error if cache miss (CI-friendly) |
| `--force` | Bust the cache |
| `--dry-run` | Estimate prompt size + cost; don't call the LLM |
| `--run <id>` | Force a specific `run-NNN` (default: latest passing) |
Output is one JSON line per framework on stdout. Non-zero exit if any
selected framework's final state is `passed: false`.
See `references/playwright-cdp-bridge.md` for the canonical
`connectOverCDP` patterns the emitted scripts follow.
### After all iterations — publish if ready
If the task passed on 2+ of the last 3 iterations **or has reached the max iteration limit**, install it as a Claude Code skill. **Do not just copy strategy.md** — the skill must be self-contained and useful to someone who has never seen this codebase. If graduating at max iterations without a clean pass, note the known failure point but still document everything learned.
Install by writing to `~/.claude/skills/<task-name>/SKILL.md`:
```bash
mkdir -p ~/.claude/skills/<task-name>
```
Use this structure for the SKILL.md:
```markdown
---
name: <task-name>
description: <1-2 sentences describing what this skill does and when to use it. Include trigger keywords.>
---
# <Task Title> — Browser Skill
## Purpose
<1-2 sentences: what this automates and why it exists.>
## When to Use
<When should someone reach for this skill.>
## Browse CLI Reference
The inner agent uses the `browse` CLI. Key commands for this task:
- `browse stop` — kill existing session (always run before switching to remote)
- `browse open <url> --remote` — start a fresh Browserbase cloud session and navigate
- `browse open <url> --local` — start a clean local browser and navigate
- `browse tab new <url>` — open URL in a new tab
- `browse wait load` — wait for page to finish loading
- `browse wait timeout <ms>` — wait a fixed amount of time for spinners or animations
- `browse wait selector "<selector>"` — wait for an element to become visible
- `browse get title` — verify you're on the right page
- `browse get text body` — extract all visible text (preferred for content extraction)
- `browse snapshot` — get accessibility tree; each node has a ref in `[X-Y]` format (e.g. `[0-5]`, `[2-147]`)
- `browse click [X-Y]` — click element by ref from the latest snapshot (include the brackets)
**Never use `--session <name>` flags in SKILL.md.** Named sessions are a parallel-run workaround — they contaminate skills with infrastructure concerns. Skills must work in isolation with the default session.
## Workflow
### Step 1 — Start session
<exact browse commands in order>
### Step 2 — Navigate
<exact URL and verification steps>
### Step 3 — Extract
<exact extraction commands>
### Step 4 — Output
<what JSON to emit, referencing the schema below>
## Site-Specific Gotchas
<Bullet list of every hard-won heuristic from the iterations. This is the core value of the skill.>
## Failure Recovery
<What to do when navigation fails, session is contaminated, or extraction returns garbage>
## Expected Output
```json
<paste the exact expected output schema from task.md>
```
```
After writing the SKILL.md, confirm it's installed:
```bash
ls ~/.claude/skills/<task-name>/SKILL.md
```
The skill is now available as `/<task-name>` in Claude Code.
---
## Final report (multi-task mode)
After all sub-agents complete, print a markdown table:
| Task | Iterations | Final Status | Graduated | Cost |
|------|-----------|--------------|-----------|------|
| google-flights | 5 | ✅ pass | yes | $0.42 |
| amazon-add-to-cart | 5 | ❌ fail | no | $1.20 |
Then write a persistent session report to `./autobrowse/reports/` so there's a durable record of the run inside the workspace:
```bash
mkdir -p ./autobrowse/reports
```
Write the file `./autobrowse/reports/YYYY-MM-DD-HH-MM-<tasks>.md` with:
```markdown
# AutoBrowse Session Report
**Date:** <ISO date>
**Tasks:** <comma-separated list>
**Environment:** remote|local
**Total cost:** $X.XX
## Results
| Task | Iterations | Pass Rate | Final Status | Graduated | Cost |
|------|-----------|-----------|--------------|-----------|------|
| ... | ... | X/5 | ✅/❌ | yes/no | $X.XX |
## Per-Task Learnings
### <task-name>
- **Key insight 1:** <what the agent learned>
- **Key insight 2:** <another heuristic>
- **Failure mode fixed:** <what was failing and how it was resolved>
## Iteration Log
### <task-name>
| Iter | Turns | Cost | Status | Hypothesis tested |
|------|-------|------|--------|-------------------|
| 1 | 79 | $18.75 | ❌ fail | baseline |
| 2 | 9 | $0.26 | ✅ pass | session contamination fix |
| ... | ... | ... | ... | ... |
```
---
## Rules
- **Only edit `strategy.md`** — never touch `task.md` (unless creating it from the template) or `evaluate.mjs`
- **Stay in the workspace** — all training writes go to `./autobrowse/`, never to `~/.claude/skills/autobrowse/`. The skill source is read-only.
- **One hypothesis per iteration** — test one change at a time
- **Build on wins** — keep what worked, add to it
- **Trust the trace** — the inner agent shows exactly what it saw and did
- **Graduate to `~/.claude/skills/`** — the only file you write there is the final graduated `SKILL.md`
- **Don't release before bisecting** — under `--browser-trace`, the order at the end of each iteration is non-negotiable: `stop-capture` → `bisect-cdp` → `browse cloud sessions update REQUEST_RELEASE`. Bisect depends on the session still existing when the trace stops.





首頁
