autobrowse
browserbase/skills
反復的な実験を通じて、サイト閲覧を行う内部エージェントを稼働させ、タスクが安定して成功するようになるまでナビゲーションの指示を改良していくことで、信頼性の高いブラウザ自動化スキルを身につけます。
...すべて拡張しますAutoBrowse — 自己改善型ブラウザスキル
反復的な実験を通じて、信頼性の高いブラウザ自動化スキルを身につけましょう。内部エージェントがサイトを閲覧します(evaluate.ts)。あなた(外部エージェント)は、何が起きたかを確認し、指示(strategy.md)を改善します。一貫して成功するまでこのプロセスを繰り返します。
エントリポイント
呼び出し方法は柔軟で、明示的なフラグも自由形式の自然言語も使用可能です:
/autobrowse --task google-flights
/autobrowse --task google-flights --iterations 10 --env remote
/autobrowse --task google-flights --browser-trace
/autobrowse --tasks google-flights,amazon-add-to-cart
/autobrowse --all
# これも問題ありません — 自由に解析してください:
/autobrowse https://flights.google.com/
/autobrowse delta.comでフライトを予約する
/autobrowse 既存のgoogle-flightsスキルを修正する
--browser-trace(デフォルトはオフ、リモート専用):各反復処理を同階層のbrowser-traceスキルとペアにします。内部のエージェントを CDP キャプチャでラップし、ページごとのネットワーク/コンソール/ページライフサイクルのエビデンスを取得します。--env remote を暗黙的に指定します。--env local と組み合わせるとエラーになります。${CLAUDE_SKILL_DIR}/../browser-trace/ に同階層のbrowser-traceスキルが存在すること、および環境変数BROWSERBASE_API_KEYが設定されていることが必要です。
ユーザーが--task の代わりに URL や自由形式の指示を入力した場合:
${WORKSPACE}/tasks/に存在するタスクがサイトや意図と明確に一致する場合は、それを使用します。- そうでない場合は、短いケバブケース形式の名前を付け、
${CLAUDE_SKILL_DIR}/references/example-task.mdから${WORKSPACE}/tasks/を作成し/task.md 、ユーザーの入力内容に基づいて URL や目標を記入してから、処理を続行します。 選択した名前を1行でユーザーに伝えてください。
実行方法
ステップ 1 — 引数を解析し、状況を把握する
渡された内容をチェックします:
--task→ シングルタスクモード--tasks a,b,cまたは--all→ マルチタスクモード(サブエージェントを起動)--iterations N→ 評価・改善サイクル数 (デフォルト: 5)--env local|remote→ ブラウザ環境(デフォルト: local; ボット対策が施されたサイトでは remote を使用)--browser-trace→ ブラウザトレース機能の統合を有効にする (デフォルトはオフ)。--env remoteが暗黙的に指定されます。もし--env localと--browser-traceの両方が明示的に指定された場合、「browser-trace には Browserbase が必要です。--env local または --browser-trace のいずれかを削除してください」というエラーが発生します。
ユーザーが代わりに自由形式のテキストを指定した場合は、処理を続行する前に、それを上記のいずれかにマッピングしてください。
ステップ 2 — ワークスペースの設定
すべてのトレーニング成果物(タスク定義、戦略の反復、トレース、レポート)は、現在の作業ディレクトリ内のワークスペースディレクトリに保存されます。~/.claude/skills/ 内には保存されません。これにより、内部エージェントによるファイル書き込みが Claude のホームディレクトリ外で行われ、権限に関するトラブルを回避できます。
デフォルトのワークスペース:${CWD}/autobrowse/
mkdir -p ./autobrowse/tasks ./autobrowse/traces ./autobrowse/reports
タスクディレクトリ(./autobrowse/tasks/)がまだ存在しない場合は、次のように作成します:
mkdir -p ./autobrowse/tasks/
cp ${CLAUDE_SKILL_DIR}/references/example-task.md ./autobrowse/tasks//task.md
# 次に、task.md を編集して、URL、入力、ステップ、および期待される JSON 出力を記述します
${CLAUDE_SKILL_DIR}にあるスキルソースは読み取り専用です。トレーニング中は、CWD 内の./autobrowse/へのみ書き込みが行われます。卒業(最終ステップ)では、~/.claude/skills/ に単一のファイルが書き込まれます。
利用可能なタスクの一覧を表示するには:
ls ./autobrowse/tasks/
ステップ 3 — マルチタスク: 並列サブエージェントの起動
複数のタスクを実行する場合は、Agentツールを使用して、タスクごとに1つのサブエージェントを同時に起動します。各サブエージェントは、そのタスクのautobrowse ループ全体を実行するための独立したプロンプトを受け取ります:
autobrowse "あなたはタスク
の[email protected]スキルを実行しています。ワークスペース:(例:/path/to/project/autobrowse)。以下の処理を反復を実行してください:evaluate → read trace → improve strategy.md → repeat。--envを使用してください。すべての evaluate.mjs 呼び出しに対して--workspaceを渡してください。親の呼び出しで--browser-trace が使用されていた場合、すべての反復(セッションの事前作成、bb-capture のアタッチ、--connect-urlを evaluate.mjs に渡す、停止+バイセクト、リリース)——デフォルトの単一コマンドパスにフォールバックしないでください。autobrowse のループ手順を厳密に遵守してください。卒業時には、適切な agentskills フロントマター(名前 + 説明)を指定して、スキルを
~/.claude/skills/インストールしてください。単に strategy.md をコピーするのではなく、独立したスキルを作成してください。/SKILL.md に 最後に、以下の項目を含む構造化された要約を出力してください:タスク名、最終実行の結果(合格/不合格)、累積総コスト、完了した反復回数、反復ごとの表(反復番号、ターン数、コスト、ステータス、検証された仮説)、および2~3項目の重要な学び。
すべてのサブエージェントを並列で起動し、すべてが完了するのを待ってから、それぞれの要約を収集してセッションレポートを作成してください。
単一タスクの場合は、この手順をスキップし、直下のループを実行してください。
ループ(各タスクごとに実行)
反復開始
./autobrowse/tasks/が存在することを確認する(存在しない場合は、テンプレートから骨組みを作成する — ステップ2を参照)。strategy.md は、初回実行時にハーネスによって自動的に空のファイルとして作成される。
要件
ANTHROPIC_API_KEY が環境変数に設定されている必要があります(または、CWD 内の.envファイルに記述されている必要があります —evaluate.mjsがこれを自動読み込みします)。これが欠落している場合、ハネスは明確なエラーを出力して終了します。他のパスでキーを探そうとしないでください。
内部エージェントの実行
デフォルトのパス(--browser-trace なし)— 単一のコマンド、オーケストレーションなし:
node ${CLAUDE_SKILL_DIR}/scripts/evaluate.mjs --task --workspace ./autobrowse
# または、ボットによる保護が施されたサイトの場合は:
node ${CLAUDE_SKILL_DIR}/scripts/evaluate.mjs --task --workspace ./autobrowse --env remote
これにより、ブラウザセッションが実行され、完全なトレースが./autobrowse/traces/ に書き込まれます。
トレースされたパス (--browser-trace、リモートのみ)— 外側のハーネスが Browserbase セッションを事前に作成し、bb-capture をパッシブオブザーバーとしてアタッチし、セッションのconnectUrlをevaluate.mjsに渡すため、すべての内部browse呼び出しで--cdp $connectUrl が使用されます。--sessionautobrowse-main(オブザーバーに完全な Network/Console イベントを提供する標準的な browser-trace パターン)。このブロックを、$N を1 を基点とする反復番号に設定して、反復ごとに 1 回実行します:
# 事前チェック —autobrowse と一緒に browser-trace がインストールされていない場合は、速やかに失敗する。
BT_DIR="${CLAUDE_SKILL_DIR}/../browser-trace"
if [ ! -f "$BT_DIR/scripts/bb-capture.mjs" ]; then
echo "ERROR: --browser-trace を実行するには、$BT_DIR に browser-trace スキルが必要です。" >&2
echo "github.com/browserbase/skills をクローンし、skills/browser-trace/" >&2
echo "をautobrowse と同じ親ディレクトリ(例: ~/.claude/skills/browser-trace/)にコピーしてインストールしてください。" >&2
exit 1
fi
# a. セッションの設定 — キープアライブセッションを事前に作成し、その connectUrl を導出する
sid=$(browse cloud sessions create --keep-alive --verified --proxies \
| node -e "let s='';process.stdin.on('data',c=>s+=c).on('end',()=>process.stdout.write(JSON.parse(s).id))")
connect_url=$(browse cloud sessions get "$sid" \
| node -e "let s='';process.stdin.on('data',c=>s+=c).on('end',()=>process.stdout.write(JSON.parse(s).connectUrl))")
RUN_ID="run-$(printf '%03d' "$N")"
TRACE_ROOT="./autobrowse/traces//$RUN_ID"
mkdir -p "$TRACE_ROOT"
export O11Y_ROOT="$TRACE_ROOT/.o11y" # browser-traceの出力をautobrowse の実行ディレクトリ内に保存
export O11Y_RUN_ID="$RUN_ID" # browse CLIに対し、descriptors.ndjsonを書き込む実行ディレクトリを指定
# b. ブラウザトレースの接続 — パッシブオブザーバー;バックグラウンドで実行
node ${CLAUDE_SKILL_DIR}/../browser-trace/scripts/bb-capture.mjs "$sid" "$RUN_ID" &
sleep 2
# c. `AUTOBROWSE ` を実行 — `connectUrl` フラグは、`evaluate.mjs` に対し、すべての内部 `browse` 呼び出しに `--cdp/--session` を挿入するよう指示します。
# 内部エージェントは `--remote` を認識しません。
node ${CLAUDE_SKILL_DIR}/scripts/evaluate.mjs \
--task --workspace ./autobrowse --env remote \
--connect-url "$connect_url" --run-number "$N"
# d. STOP + BISECT + UNIFY — 順序が重要である。bisect を実行するにはセッションがまだ
# 存在している必要があり、unify-trace は bisect の出力をautobrowse の trace.json と結合する
# これらを、外側のエージェントが各反復で最初に読み込む、時刻順に並べられた単一の NDJSON に統合します。
node ${CLAUDE_SKILL_DIR}/../browser-trace/scripts/stop-capture.mjs "$RUN_ID"
node ${CLAUDE_SKILL_DIR}/../browser-trace/scripts/bisect-cdp.mjs "$RUN_ID"
node ${CLAUDE_SKILL_DIR}/scripts/unify-trace.mjs \
--trace-dir "$TRACE_ROOT" \
--o11y-dir "$O11Y_ROOT/$RUN_ID"
# e. リリース
browse cloud sessions update "$sid" --status REQUEST_RELEASE
これにより、エージェント内部のトレースが./autobrowse/traces/に、CDP バイセクトが./autobrowse/traces/ に書き込まれます。 トレースされたブラウズCLI は、コマンドごとの詳細なノード記述子も.o11y/に出力します(ページを駆動する呼び出しごとに 1 つの JSON オブジェクト:target tag/id/role/accessibleName/attributes/xpath/bounding-rect)。 この記述子ファイルは下流のコード生成に利用されますが、仮説の構築には必須ではありません。トレースを読み取る際は、このファイルをスキップしてください。
トレースの読み方
cat ./autobrowse/traces//latest/summary.md
サマリーには、所要時間、コスト、ターン数、決定ログ、および最終的な JSON 出力が含まれています。
エージェントが失敗したり、行き詰まったりした場合は、さらに詳しく調べてください:
./autobrowse/traces/を読み込み、失敗したターンを探してください/latest/trace.json - Readツールを使用して、失敗時点前後のスクリーンショットを確認してください
--browser-trace オプションが使用された場合 —unified-events.jsonl から調査を開始してください。ハーネスは、エージェントのターンログとブラウザの CDP ファイアホースを、実行ルートで時間順に並べられた単一の NDJSON ストリームに結合します。 1つのファイルで、ソースタグ(source: "agent" | "browser")が付与され、実時間のタイムスタンプでインターリーブされています。上から下までざっと目を通してください。障害の原因は通常、隣接する1行または2行にあります(エージェントがコマンドXを発行し、ブラウザがYで応答したなど)。
cat ./autobrowse/traces//latest/unified-events.jsonl
構造化されたファイル(trace.json、.o11y/)も、統合ストリームからさらに詳細が必要な項目が特定された場合に、エージェントがドリルダウンとして利用可能です:
| 必要 | ドリルダウン用のファイルまたはコマンド |
|---|---|
| ページごとの集計値とタイミング(イベント数、ネットワーク接続数、ページごとのエラー数) | .o11y/ |
| すべての失敗したネットワークリクエストを一か所に集約 | .o11y/ |
| コンソールの例外ペイロード全文(スタックトレースなど) | .o11y/ |
| ページごとのスライス(ページ N でのイベントのみ) | .o11y/ |
| 特定のターンに関する完全な推論テキスト/切り捨てられていないツール出力 | trace.json(ターン === N でフィルタリング) |
| アドホックなグループ化クエリ(例:上位ホスト、ページごとのエラー) | O11Y_ROOT=./autobrowse/traces/ |
統合ストリームがデフォルトです。グループ化されたクエリ、全文ペイロード、またはストリームでは得られないフィルタリングが必要な場合にのみ、構造化ファイルにドリルダウンしてください。
仮説を立てる
問題が発生した正確な時点を特定する。それを防ぐことができた単一のヒューリスティックは何か?
--browser-trace オプションを使用する場合、仮説には unified-events.jsonl内の特定のイベント(行番号またはタイムスタンプ)を引用する必要があります。あるいは、ドリルダウンが必要だった場合は、その対象ファイル名を明記してください。これにより、更新内容が「直感」ではなく「証拠」に基づいたものになります。 エージェントのコマンドのみに基づく仮説は、「クリックが機能しなかった」といった内容になるかもしれません。 統一ストリームに基づいた仮説であれば、「unified-events.jsonlの47行目:browse openの直後に、/api/checkoutで Network.responseReceivedステータス403が発生した —--verified --proxiesに切り替える」といった形になります。
例:
- 「ドロップダウンをクリックした後、1秒待機する — オプションはクリック可能になる前にアニメーションで表示される」
- 「
/pay-invoice/に直接移動 — ランディングページを完全にスキップ」 - 「
browse fill #field_3 valueを使用し、browse typeは使用しない — このフィールドはフォーカスが当たるとクリアされる」 - "ターン8でページにスピナーが表示される — スナップショット前に
ブラウズ待機タイムアウト2000を追加" - (
--browser-traceオプション使用時) 「unified-events.jsonlの47行目において、/api/availabilityへの3回連続のNetwork.responseReceivedイベントが、ブラウズ開始直後に403を返しました。サイトはフィンガープリントを取得しています。次の反復では--verified --proxies を指定する必要があります。」
strategy.md を更新
./autobrowse/tasks/ を編集する。正常に動作していた部分はすべて維持する。特定の失敗箇所を修正する。具体的なヒューリスティックを追加する。
優れた戦略には以下が含まれます:
- 高速パス:探索をスキップするための直接URLやショートカット
- 段階的なワークフロー:タイミングの注記を含む正確な手順
- サイト固有の知識:セレクタID、フォームフィールド名、成功の指標
- 失敗時の対応:Xがうまくいかなかった場合の対処法
結果を評価する
新しい要約を読む。合格したか?明確な進捗があったか?
- 合格または進展あり→ 維持し、次の反復へ
- 進展がない、または後退している→ strategy.mdを前のバージョンにロールバックし、別の仮説を試す
実行可能なスクリプトを生成する(オプション)
タスクが収束したら、scripts/codegen.mjs を使用して、1つ以上のフレームワーク向けの
決定論的で実行可能なスクリプトを生成できます。これは、フレームワークごとに
1回のLLM呼び出しを行い、コンテンツハッシュでキャッシュされるもので、
「最新のセッションとの照合」や「失敗時の書き換え」をオプションで指定できます。
node ${CLAUDE_SKILL_DIR}/scripts/codegen.mjs \
--task \
--workspace ./autobrowse \
--frameworks playwright,stagehand \
--verify
各フレームワークには、tasks/の下に、
生成されたスクリプトと独立したスケルトン(package.json、
tsconfig.json)を含む独自のサブディレクトリが作成されます。 このディレクトリは、
cd tasks/と実行することでスタンドアロンで動作します。必要な
実行環境はBROWSERBASE_API_KEYのみです(Stagehand をターゲットとする場合は、
ANTHROPIC_API_KEYも必要です)。
組み込みフレームワーク:playwright、stagehand。カスタムフレームワークを追加するには、
--prompt-templateを実行します(独自のランナーを
指定するか、--no-verify を指定してください)。
一般的なフラグ:
| フラグ | 目的 |
|---|---|
--frameworks a,b,... |
カンマ区切り。デフォルトはplaywright |
--verify/--no-verify |
生成されたスクリプトを新しいBBセッションで実行する。デフォルトは--verify |
--max-retries N |
検証失敗時の書き換え上限; デフォルトは 2 |
--cache-only |
キャッシュミスの場合はエラー(CI 対応) |
--force |
キャッシュを破棄する |
--dry-run |
プロンプトのサイズとコストを見積もる。LLMを呼び出さない |
--run |
特定の実行-NNNを強制(デフォルト:最新の合格例) |
出力は、フレームワークごとに1行のJSONが標準出力(stdout)に出力されます。選択されたフレームワークの最終状態が「false」と判定された場合、終了コードは0以外になります。
参照: references/playwright-cdp-bridge.md を参照し、生成されたスクリプトが従う標準的な
生成されたスクリプトが従う標準的な
connectOverCDPパターンについては、references/playwright-cdp-bridge.md を参照してください。
すべての反復処理終了後 — 準備が整っていれば公開する
直近3回の反復のうち2回以上でタスクが合格した場合、または最大反復回数に達した場合は、それをClaude Codeスキルとしてインストールします。単に strategy.md をコピーするだけではいけません。スキルは独立したものであり、このコードベースを見たことがない人にとっても有用でなければなりません。最大反復回数に達しても完全に合格できなかった場合は、既知の失敗点を記録しつつ、学んだことはすべて文書化してください。
~/.claude/skills/ に記述してインストールします:
mkdir -p ~/.claude/skills/
SKILL.md には以下の構造を使用してください:
---
name:
description:<1-2 sentences describing what this skill does and when to use it. Include trigger keywords.>
---
# — ブラウザスキル
## 目的
<1-2 sentences: what this automates and why it exists.>
## 使用場面
## Browse CLI リファレンス
内部エージェントは `browse` CLI を使用します。このタスクにおける主要なコマンドは以下の通りです:
- `browse stop` — 既存のセッションを終了する(リモートに切り替える前には必ず実行すること)
- `browse open --remote` — 新しい Browserbase クラウドセッションを開始し、ナビゲートする
- `browse open --local` — クリーンなローカルブラウザを起動し、ナビゲートする
- `browse tab new` — URLを新しいタブで開く
- `browse wait load` — ページの読み込みが完了するまで待機する
- `browse wait timeout` — ローディングアイコンやアニメーションが表示されるまで、指定した時間待機する
- `browse wait selector ""` — 要素が表示されるのを待つ
- `browse get title` — 正しいページにアクセスしていることを確認する
- `browse get text body` — 表示されているテキストをすべて抽出する(コンテンツ抽出に推奨)
- `browse snapshot` — アクセシビリティツリーを取得します。各ノードには `[X-Y]` 形式(例: `[0-5]`、`[2-147]`)の ref が割り当てられています
- `browse click [X-Y]` — 最新のスナップショットから ref に基づいて要素をクリックします (角括弧を含める)
**SKILL.md では、`--session` フラグを絶対に使用しないでください。** 名前付きセッションは並列実行のための回避策であり、スキルにインフラストラクチャ上の問題を混入させてしまいます。スキルは、デフォルトのセッションで独立して動作する必要があります。
## ワークフロー
### ステップ 1 — セッションの開始
### ステップ 2 — ナビゲート
### ステップ 3 — 抽出
### ステップ 4 — 出力
## サイト固有の注意点
## 障害復旧
## 期待される出力
```json
SKILL.md を作成したら、インストールされていることを確認してください:
```bash
ls ~/.claude/skills//SKILL.md
これで、Claude Code上で/としてスキルが利用可能になりました。
最終レポート(マルチタスクモード)
すべてのサブエージェントが完了したら、Markdown形式の表を出力します:
| タスク | 反復回数 | 最終ステータス | 修了 | コスト |
|---|---|---|---|---|
| google-flights | 5 | ✅ 合格 | はい | 0.42ドル |
| amazon-add-to-cart | 5 | ❌ 失敗 | いいえ | $1.20 |
次に、ワークスペース内に実行の永続的な記録が残るように、./autobrowse/reports/に永続的なセッションレポートを書き込みます:
mkdir -p ./autobrowse/reports
ファイル./autobrowse/reports/YYYY-MM-DD-HH-MM-に以下のように記述します:
#AutoBrowse セッションレポート
**日付:**
**タスク:**
**環境:** リモート|ローカル
**総コスト:** $X.XX
## 結果
| タスク | 反復回数 | 合格率 | 最終ステータス | 修了 | コスト |
|------|-----------|-----------|--------------|-----------|------|
| ... | ... | X/5 | ✅/❌ | はい/いいえ | $X.XX |
## タスクごとの学び
###
- **主な知見 1:**
- **主な知見 2:**
- **修正された不具合:**
## 反復ログ
###
| 反復 | ターン数 | コスト | ステータス | 検証した仮説 |
|------|-------|------|--------|-------------------|
| 1 | 79 | $18.75 | ❌ 失敗 | ベースライン |
| 2 | 9 | $0.26 | ✅ 成功 | セッション汚染の修正 |
| ... | ... | ... | ... | ... |
ルール
-
strategy.mdのみを編集してください。task.md(テンプレートから作成する場合を除く)やevaluate.mjsには絶対に手を加えないでください - ワークスペース内に留まること— すべてのトレーニング書き込みは
./autobrowse/へ行い、~/.claude/skills/autobrowse/には決して書き込まないこと。スキルのソースは読み取り専用です。 - 1回の反復につき1つの仮説— 変更は1回につき1つずつテストしてください
- 成功を積み重ねる— うまくいった部分は維持し、それに追加する
- トレースを信頼する— 内部エージェントは、自分が何を見て何をしたかを正確に示しています
-
~/.claude/skills/へ移行する— そこに書き込む唯一のファイルは、最終的な「卒業済み」SKILL.mdファイルである - バイセクトを行うまではリリースしない—
--browser-trace オプションを有効にしている場合、各反復の終了時の順序は絶対条件です:stop-capture→bisect-cdp→クラウドセッションを閲覧し、REQUEST_RELEASE を更新。バイセクトは、トレースが停止した時点でセッションがまだ存在していることを前提としています。
---
name: autobrowse
description: Builds reliable browser automation skills through iterative experimentation, running an inner agent to browse sites and improving navigation instructions until tasks pass consistently.
license: MIT
---
# AutoBrowse — Self-Improving Browser Skill
Build reliable browser automation skills through iterative experimentation. An inner agent browses the site (`evaluate.ts`). You — the outer agent — read what happened and improve the instructions (`strategy.md`). Repeat until it passes consistently.
## Entry Points
Invocation is flexible — both explicit flags and free-form natural language work:
```
/autobrowse --task google-flights
/autobrowse --task google-flights --iterations 10 --env remote
/autobrowse --task google-flights --browser-trace
/autobrowse --tasks google-flights,amazon-add-to-cart
/autobrowse --all
# Also fine — parse freely:
/autobrowse https://flights.google.com/
/autobrowse book a flight on delta.com
/autobrowse fix the existing google-flights skill
```
`--browser-trace` (default off, remote-only): pairs each iteration with the sibling `browser-trace` skill — wraps the inner agent in a CDP capture for per-page network/console/page-lifecycle evidence. Implies `--env remote`; errors if combined with `--env local`. Requires the sibling `browser-trace` skill present at `${CLAUDE_SKILL_DIR}/../browser-trace/`, and the `BROWSERBASE_API_KEY` env var.
When the user drops a URL or free-form instruction instead of `--task <name>`:
- If an existing task in `${WORKSPACE}/tasks/` clearly matches the site/intent, use it.
- Otherwise, pick a short kebab-case name, create `${WORKSPACE}/tasks/<name>/task.md` from `${CLAUDE_SKILL_DIR}/references/example-task.md`, fill in the URL/goal based on what the user said, and proceed. Tell the user the chosen name in one line.
---
## How to run
### Step 1 — Parse arguments and orient
Check what was passed:
- `--task <name>` → single task mode
- `--tasks a,b,c` or `--all` → multi-task mode (spawn sub-agents)
- `--iterations N` → how many evaluate → improve cycles (default: 5)
- `--env local|remote` → browser environment (default: local; use remote for bot-protected sites)
- `--browser-trace` → opt in to the browser-trace integration (default off). Implies `--env remote`. If `--env local --browser-trace` are both passed explicitly, error with: `browser-trace requires Browserbase; drop --env local or drop --browser-trace.`
If the user passed free-form text instead, map it to one of the above before continuing.
### Step 2 — Set up the workspace
All training artifacts (task definitions, strategy iterations, traces, reports) live in a workspace directory in the **current working directory** — NOT inside `~/.claude/skills/`. This keeps the inner agent's file writes out of Claude's home dir and away from permission friction.
Default workspace: `${CWD}/autobrowse/`
```bash
mkdir -p ./autobrowse/tasks ./autobrowse/traces ./autobrowse/reports
```
If the task directory (`./autobrowse/tasks/<task>/task.md`) doesn't exist yet, scaffold it:
```bash
mkdir -p ./autobrowse/tasks/<task>
cp ${CLAUDE_SKILL_DIR}/references/example-task.md ./autobrowse/tasks/<task>/task.md
# Then edit task.md to describe the URL, inputs, steps, and expected JSON output
```
The skill source at `${CLAUDE_SKILL_DIR}` stays read-only — only `./autobrowse/` in CWD gets written to during training. Graduation (final step) writes a single file to `~/.claude/skills/<task>/SKILL.md`.
List available tasks:
```bash
ls ./autobrowse/tasks/
```
### Step 3 — Multi-task: spawn parallel sub-agents
If running multiple tasks, use the Agent tool to spawn one sub-agent per task simultaneously. Each sub-agent receives a self-contained prompt to run the full autobrowse loop for its task:
> "You are running the autobrowse skill for task `<name>`. Workspace: `<absolute-path-to-workspace>` (e.g. `/path/to/project/autobrowse`). Run `<N>` iterations of: evaluate → read trace → improve strategy.md → repeat. Use `--env <env>`. Pass `--workspace <workspace>` to every evaluate.mjs invocation. If the parent invocation used `--browser-trace`, you MUST use the traced-path block of the SKILL.md loop for every iteration (pre-create session, attach bb-capture, pass `--connect-url` to evaluate.mjs, stop+bisect, release) — do not fall back to the default single-command path. Follow the autobrowse loop instructions exactly.
>
> When graduating, install the skill to `~/.claude/skills/<task-name>/SKILL.md` with proper agentskills frontmatter (name + description). Do not just copy strategy.md — write a self-contained skill.
>
> At the end, output a structured summary with: task name, pass/fail on final run, total cumulative cost, iterations completed, per-iteration table (iter number, turns, cost, status, hypothesis tested), and 2-3 bullet key learnings."
Spawn all sub-agents in parallel, wait for all to complete, then collect their summaries and write the session report.
**For single task**, skip this step and run the loop directly below.
---
## The Loop (run this for each task)
### Iteration start
Check that `./autobrowse/tasks/<task>/task.md` exists (scaffold it from the template if not — see Step 2). `strategy.md` is auto-created empty by the harness on first run.
### Requirements
- `ANTHROPIC_API_KEY` must be in the environment (or in a `.env` file in CWD — `evaluate.mjs` auto-loads it). If missing, the harness prints a clear error and exits; don't hunt for keys in other paths.
### Run the inner agent
**Default path (no `--browser-trace`)** — single command, no orchestration:
```bash
node ${CLAUDE_SKILL_DIR}/scripts/evaluate.mjs --task <task-name> --workspace ./autobrowse
# or for bot-protected sites:
node ${CLAUDE_SKILL_DIR}/scripts/evaluate.mjs --task <task-name> --workspace ./autobrowse --env remote
```
This runs the browser session and writes a full trace to `./autobrowse/traces/<task>/latest/`.
**Traced path (`--browser-trace`, remote only)** — the outer harness pre-creates a Browserbase session, attaches `bb-capture` as a passive observer, and passes the session's `connectUrl` to `evaluate.mjs` so every inner `browse` call uses `--cdp $connectUrl --session autobrowse-main` (the canonical browser-trace pattern that gives observers full Network/Console events). Run this block once per iteration with `$N` set to the 1-indexed iteration number:
```bash
# Preflight — fail fast if browser-trace isn't installed alongside autobrowse.
BT_DIR="${CLAUDE_SKILL_DIR}/../browser-trace"
if [ ! -f "$BT_DIR/scripts/bb-capture.mjs" ]; then
echo "ERROR: --browser-trace requires the browser-trace skill at $BT_DIR." >&2
echo "Install it by cloning github.com/browserbase/skills and copying skills/browser-trace/" >&2
echo "into the same parent directory as autobrowse (e.g. ~/.claude/skills/browser-trace/)." >&2
exit 1
fi
# a. SESSION SETUP — pre-create the keep-alive session and derive its connectUrl
sid=$(browse cloud sessions create --keep-alive --verified --proxies \
| node -e "let s='';process.stdin.on('data',c=>s+=c).on('end',()=>process.stdout.write(JSON.parse(s).id))")
connect_url=$(browse cloud sessions get "$sid" \
| node -e "let s='';process.stdin.on('data',c=>s+=c).on('end',()=>process.stdout.write(JSON.parse(s).connectUrl))")
RUN_ID="run-$(printf '%03d' "$N")"
TRACE_ROOT="./autobrowse/traces/<task-name>/$RUN_ID"
mkdir -p "$TRACE_ROOT"
export O11Y_ROOT="$TRACE_ROOT/.o11y" # park browser-trace output inside the autobrowse run dir
export O11Y_RUN_ID="$RUN_ID" # tells the browse CLI which run dir to write descriptors.ndjson into
# b. ATTACH BROWSER-TRACE — passive observer; runs in background
node ${CLAUDE_SKILL_DIR}/../browser-trace/scripts/bb-capture.mjs "$sid" "$RUN_ID" &
sleep 2
# c. RUN AUTOBROWSE — connectUrl flag tells evaluate.mjs to inject --cdp/--session
# into every inner browse call. The inner agent never sees --remote.
node ${CLAUDE_SKILL_DIR}/scripts/evaluate.mjs \
--task <task-name> --workspace ./autobrowse --env remote \
--connect-url "$connect_url" --run-number "$N"
# d. STOP + BISECT + UNIFY — order matters; bisect needs the session to still
# exist, and unify-trace joins the bisect output with autobrowse's trace.json
# into a single time-ordered NDJSON the outer agent reads first each iter.
node ${CLAUDE_SKILL_DIR}/../browser-trace/scripts/stop-capture.mjs "$RUN_ID"
node ${CLAUDE_SKILL_DIR}/../browser-trace/scripts/bisect-cdp.mjs "$RUN_ID"
node ${CLAUDE_SKILL_DIR}/scripts/unify-trace.mjs \
--trace-dir "$TRACE_ROOT" \
--o11y-dir "$O11Y_ROOT/$RUN_ID"
# e. RELEASE
browse cloud sessions update "$sid" --status REQUEST_RELEASE
```
This writes the inner-agent trace to `./autobrowse/traces/<task-name>/latest/` and the CDP bisect to `./autobrowse/traces/<task-name>/latest/.o11y/<run-id>/`. The traced `browse` CLI also emits per-command rich node descriptors to `.o11y/<run-id>/cdp/descriptors.ndjson` (one JSON object per page-driving call: target tag/id/role/accessibleName/attributes/xpath/bounding-rect). The descriptors file feeds downstream codegen; it is **not** required for hypothesis formation — skip it when reading the trace.
### Read the trace
```bash
cat ./autobrowse/traces/<task-name>/latest/summary.md
```
The summary has duration, cost, turns, the decision log, and the final JSON output.
If the agent failed or got stuck, look deeper:
- Read `./autobrowse/traces/<task-name>/latest/trace.json` — search for the failure turn
- Read screenshots around the failure point with the Read tool
**When `--browser-trace` was used — start with `unified-events.jsonl`.** The harness joins the agent's turn log and the browser's CDP firehose into one time-ordered NDJSON stream at the run root. One file, source-tagged (`source: "agent" | "browser"`), interleaved by wall-clock timestamp. Skim it top-to-bottom; the failure cause is usually one or two adjacent lines (the agent issued command X, the browser responded with Y).
```bash
cat ./autobrowse/traces/<task-name>/latest/unified-events.jsonl
```
The structured files (`trace.json`, `.o11y/<run-id>/cdp/*`) are **also agent-consumable as drill-downs** when the unified stream points at something you need more of:
| Need | Drill-down file or command |
|---|---|
| Per-page totals + timing (events, network counts, errors by page) | `.o11y/<run-id>/cdp/summary.json` |
| All failed network requests in one place | `.o11y/<run-id>/cdp/network/failed.jsonl` |
| Full console exception payloads (stacktraces, etc.) | `.o11y/<run-id>/cdp/console/exceptions.jsonl` |
| Per-page slice (only events on page N) | `.o11y/<run-id>/cdp/pages/<pid>/` |
| Full reasoning text / untruncated tool outputs for a specific turn | `trace.json` (filter by `turn === N`) |
| Ad-hoc grouped query (e.g. top hosts, errors-by-page) | `O11Y_ROOT=./autobrowse/traces/<task-name>/latest/.o11y node ${CLAUDE_SKILL_DIR}/../browser-trace/scripts/query.mjs <run-id> <cmd>` |
The unified stream is the default; drill into structured files only when you need a grouped query, a full-text payload, or filtering the stream can't give you.
### Form one hypothesis
Find the exact turn where things went wrong. What single heuristic would have prevented it?
Under `--browser-trace`, the hypothesis must cite a **specific event from `unified-events.jsonl`** (line number or timestamp) — or name the drill-down file if you had to descend into one. This keeps updates evidence-grounded rather than vibes-driven. A hypothesis based only on the agent's commands might say "the click didn't work"; grounded in the unified stream, it can say "line 47 of unified-events.jsonl: `browse open` was followed by `Network.responseReceived` status 403 on `/api/checkout` — switch to `--verified --proxies`."
Examples:
- "After clicking the dropdown, wait 1s — options animate in before they're clickable"
- "Navigate directly to `/pay-invoice/` — skip the landing page entirely"
- "Use `browse fill #field_3 value` not `browse type` — this field clears on focus"
- "The page shows a spinner at turn 8 — add `browse wait timeout 2000` before snapshot"
- (with `--browser-trace`) "At line 47 of unified-events.jsonl, 3 consecutive `Network.responseReceived` events on `/api/availability` returned 403 right after `browse open` — the site is fingerprinting; the next iter needs `--verified --proxies`."
### Update strategy.md
Edit `./autobrowse/tasks/<task-name>/strategy.md`. Keep everything that worked. Fix the specific failure. Add a concrete heuristic.
Good strategies have:
- **Fast path**: direct URL or shortcuts to skip exploration
- **Step-by-step workflow**: exact sequence with timing notes
- **Site-specific knowledge**: selector IDs, form field names, success indicators
- **Failure recovery**: what to do when X goes wrong
### Judge the result
Read the new summary. Did it pass? Make clear progress?
- **Pass or progress** → keep, next iteration
- **No progress or regression** → revert strategy.md to the previous version and try a different hypothesis
### Generate a runnable script (optional)
Once the task has converged, you can produce a deterministic, runnable script
in one or more frameworks via `scripts/codegen.mjs`. This is one shot of an
LLM call per framework, cached by content hash, with optional verify-against-
fresh-session and rewrite-on-failure.
```bash
node ${CLAUDE_SKILL_DIR}/scripts/codegen.mjs \
--task <name> \
--workspace ./autobrowse \
--frameworks playwright,stagehand \
--verify
```
Each framework gets its own subdirectory under `tasks/<name>/<framework>/`
with the emitted script and a self-contained scaffold (`package.json`,
`tsconfig.json`). The directory is runnable standalone with
`cd tasks/<name>/playwright && npm install && npx tsx <name>.ts` — the only
runtime requirement is `BROWSERBASE_API_KEY` (plus `ANTHROPIC_API_KEY` for
the Stagehand target).
Builtin frameworks: `playwright`, `stagehand`. Add a custom framework with
`--prompt-template <path> --frameworks custom` (and provide your own runner
or pass `--no-verify`).
Common flags:
| Flag | Purpose |
|---|---|
| `--frameworks a,b,...` | Comma-separated; default `playwright` |
| `--verify` / `--no-verify` | Run the produced script against a fresh BB session; default `--verify` |
| `--max-retries N` | Rewrite-on-verify-failure cap; default 2 |
| `--cache-only` | Error if cache miss (CI-friendly) |
| `--force` | Bust the cache |
| `--dry-run` | Estimate prompt size + cost; don't call the LLM |
| `--run <id>` | Force a specific `run-NNN` (default: latest passing) |
Output is one JSON line per framework on stdout. Non-zero exit if any
selected framework's final state is `passed: false`.
See `references/playwright-cdp-bridge.md` for the canonical
`connectOverCDP` patterns the emitted scripts follow.
### After all iterations — publish if ready
If the task passed on 2+ of the last 3 iterations **or has reached the max iteration limit**, install it as a Claude Code skill. **Do not just copy strategy.md** — the skill must be self-contained and useful to someone who has never seen this codebase. If graduating at max iterations without a clean pass, note the known failure point but still document everything learned.
Install by writing to `~/.claude/skills/<task-name>/SKILL.md`:
```bash
mkdir -p ~/.claude/skills/<task-name>
```
Use this structure for the SKILL.md:
```markdown
---
name: <task-name>
description: <1-2 sentences describing what this skill does and when to use it. Include trigger keywords.>
---
# <Task Title> — Browser Skill
## Purpose
<1-2 sentences: what this automates and why it exists.>
## When to Use
<When should someone reach for this skill.>
## Browse CLI Reference
The inner agent uses the `browse` CLI. Key commands for this task:
- `browse stop` — kill existing session (always run before switching to remote)
- `browse open <url> --remote` — start a fresh Browserbase cloud session and navigate
- `browse open <url> --local` — start a clean local browser and navigate
- `browse tab new <url>` — open URL in a new tab
- `browse wait load` — wait for page to finish loading
- `browse wait timeout <ms>` — wait a fixed amount of time for spinners or animations
- `browse wait selector "<selector>"` — wait for an element to become visible
- `browse get title` — verify you're on the right page
- `browse get text body` — extract all visible text (preferred for content extraction)
- `browse snapshot` — get accessibility tree; each node has a ref in `[X-Y]` format (e.g. `[0-5]`, `[2-147]`)
- `browse click [X-Y]` — click element by ref from the latest snapshot (include the brackets)
**Never use `--session <name>` flags in SKILL.md.** Named sessions are a parallel-run workaround — they contaminate skills with infrastructure concerns. Skills must work in isolation with the default session.
## Workflow
### Step 1 — Start session
<exact browse commands in order>
### Step 2 — Navigate
<exact URL and verification steps>
### Step 3 — Extract
<exact extraction commands>
### Step 4 — Output
<what JSON to emit, referencing the schema below>
## Site-Specific Gotchas
<Bullet list of every hard-won heuristic from the iterations. This is the core value of the skill.>
## Failure Recovery
<What to do when navigation fails, session is contaminated, or extraction returns garbage>
## Expected Output
```json
<paste the exact expected output schema from task.md>
```
```
After writing the SKILL.md, confirm it's installed:
```bash
ls ~/.claude/skills/<task-name>/SKILL.md
```
The skill is now available as `/<task-name>` in Claude Code.
---
## Final report (multi-task mode)
After all sub-agents complete, print a markdown table:
| Task | Iterations | Final Status | Graduated | Cost |
|------|-----------|--------------|-----------|------|
| google-flights | 5 | ✅ pass | yes | $0.42 |
| amazon-add-to-cart | 5 | ❌ fail | no | $1.20 |
Then write a persistent session report to `./autobrowse/reports/` so there's a durable record of the run inside the workspace:
```bash
mkdir -p ./autobrowse/reports
```
Write the file `./autobrowse/reports/YYYY-MM-DD-HH-MM-<tasks>.md` with:
```markdown
# AutoBrowse Session Report
**Date:** <ISO date>
**Tasks:** <comma-separated list>
**Environment:** remote|local
**Total cost:** $X.XX
## Results
| Task | Iterations | Pass Rate | Final Status | Graduated | Cost |
|------|-----------|-----------|--------------|-----------|------|
| ... | ... | X/5 | ✅/❌ | yes/no | $X.XX |
## Per-Task Learnings
### <task-name>
- **Key insight 1:** <what the agent learned>
- **Key insight 2:** <another heuristic>
- **Failure mode fixed:** <what was failing and how it was resolved>
## Iteration Log
### <task-name>
| Iter | Turns | Cost | Status | Hypothesis tested |
|------|-------|------|--------|-------------------|
| 1 | 79 | $18.75 | ❌ fail | baseline |
| 2 | 9 | $0.26 | ✅ pass | session contamination fix |
| ... | ... | ... | ... | ... |
```
---
## Rules
- **Only edit `strategy.md`** — never touch `task.md` (unless creating it from the template) or `evaluate.mjs`
- **Stay in the workspace** — all training writes go to `./autobrowse/`, never to `~/.claude/skills/autobrowse/`. The skill source is read-only.
- **One hypothesis per iteration** — test one change at a time
- **Build on wins** — keep what worked, add to it
- **Trust the trace** — the inner agent shows exactly what it saw and did
- **Graduate to `~/.claude/skills/`** — the only file you write there is the final graduated `SKILL.md`
- **Don't release before bisecting** — under `--browser-trace`, the order at the end of each iteration is non-negotiable: `stop-capture` → `bisect-cdp` → `browse cloud sessions update REQUEST_RELEASE`. Bisect depends on the session still existing when the trace stops.
すべてのファイル
21件のファイルautobrowseをインストール
スキルファイルをダウンロードし、.claude/skills/ ディレクトリに解凍してください。
ZIPをダウンロードリポジトリをクローンし、スキルファイルをプロジェクトにコピーしてください。
git clone https://github.com/browserbase/skills/tree/main/skills/autobrowse # Copy SKILL.md to your .claude/skills/ directory
コピー





家
