オプション
家家 Skill API開発 browser-to-api

browser-to-api

browserbase/skills browserbase/skills

観測されたHTTPトラフィックを分析し、URLをテンプレート化し、リクエスト/レスポンスのサンプルからJSONスキーマを推測することで、ブラウザトレースのキャプチャデータからOpenAPI 3.1仕様を生成します。

...すべて拡張します
0
更新された時間 2026年9月30日

ブラウザからAPIへ

リプレイ駆動型のAPI検出。ブラウザトレースのキャプチャを処理し、CDPのリクエスト/レスポンスイベントを照合し、観測されたURLをテンプレート化し、サンプルからJSONスキーマを推論し、OpenAPI 3.1ドキュメントと人間が読み取れるカバレッジレポートを出力します。

このスキルはトラフィックをキャプチャしません。これは、browser-traceの cdp/network/*.jsonlバケットを基にした、純粋なオフライン後処理です。2つのスキルは次のように組み合わされます:

browser-trace    →  .o11y//cdp/network/{requests,responses}.jsonl
browser-to-api  →  .o11y//api-spec/index.html + openapi.yaml + client.mjs

使用すべき場面

  • ユーザーが、サードパーティ製またはドキュメント化されていないWebサイトAPIのOpenAPIドキュメントを必要としている場合。
  • ユーザーがブラウザトレースを実行しており、そこからエンドポイントとスキーマを抽出したい場合。
  • ユーザーが、仕様を公開していないサイト向けのクライアントやSDKを構築している場合。
  • 仕様を拡張できるフローを示すカバレッジレポートが必要な場合。

トラフィックをキャプチャしたい場合は、まずブラウザトレースを利用するよう案内してください。

2段階のワークフロー

1.browser-traceを使用してトラフィックをキャプチャする(必要に応じて、「ネットワークを閲覧」を有効にしてボディもキャプチャする)

# 既存のデバッグ可能なChromeターゲットに対するローカル例
TARGET=9222

node ../browser-trace/scripts/start-capture.mjs "$TARGET" my-site
browse open about:blank --cdp "$TARGET"
browse network on                                    # リクエスト/レスポンスのボディをキャプチャ
browse open https://example.com
# ...キャプチャ対象とするフローを任意に実行...

# キャプチャをオフにする前に、ボディディレクトリのスナップショットを作成します(一時ディレクトリは
# セッションごとに共有されるため、この手順を省略すると、以降の `browse network on` の実行時に、
# ボディデータが将来のキャプチャで書き込まれるデータと混在してしまいます)。
cp -r "$(browse network path | jq -r .path)" .o11y/my-site/cdp/network/bodies/
browse network off

node ../browser-trace/scripts/stop-capture.mjs my-site
node ../browser-trace/scripts/bisect-cdp.mjs my-site

browse network on はオプションですが、強く推奨されます。これがないと、仕様にはレスポンスボディのスキーマが含まれません(browse cdpが使用する CDP ファイアホースにはボディが埋め込まれていないため)。これを実行すると、リクエストボディ(すでに CDP によってキャプチャ済み)とレスポンスボディの両方が、CDP のrequestId によってトレースに結合されます。

2. 仕様書を生成する

node scripts/discover.mjs --run .o11y/my-site
# → .o11y/my-site/api-spec/index.html          ← これを開く
#   .o11y/my-site/api-spec/client.mjs
#   .o11y/my-site/api-spec/openapi.yaml
#   .o11y/my-site/api-spec/openapi.json
#   .o11y/my-site/api-spec/report.md
#   .o11y/my-site/api-spec/confidence.json
#   .o11y/my-site/api-spec/samples/*.json
#   .o11y/my-site/api-spec/intermediate/*.jsonl

discover.mjs は 、/cdp/network/bodies/ を自動検出します。他の場所からのボディキャプチャを使用する場合(例:スナップショットを撮っていない、ライブブラウズのネットワークディレクトリを使用したい場合など)、--bodiesオプションに を明示的に指定してください。

3. HTMLレポートを開く

discover.mjsの処理が完了したら、必ず生成された HTML レポートを開いてください:

open .o11y/my-site/api-spec/index.html

このレポートは、サーバーを必要としない独立したHTMLファイルであり、検出された各操作が展開可能なカードとして表示されます。各カードには、変数、クライアントでの使用例、リクエスト/レスポンスの例、および下部に生成されたclient.mjsスニペットが含まれています。これが主要な成果物ですので、必ずユーザーに表示してください。

CLI フラグ

オプション 必須 意味
--run yes ブラウザトレースの実行ディレクトリへのパス
--out no 出力ディレクトリ。デフォルトは/api-spec/
--bodies なし トレースに追加するネットワークキャプチャディレクトリを参照(存在する場合、/api-spec/cdp/network/bodies/から自動検出されます)
--include なし 正規表現に一致するURLのみを含める(繰り返し指定可能)
--exclude なし 正規表現に一致するURLを除外する(繰り返し指定可能。デフォルト設定に加えて)
--origins なし コンマ区切りのオリジン許可リスト(例:api.example.com,example.com)
--format なし 出力形式。デフォルトはboth
--title なし OpenAPIのinfo.title。デフォルトはプライマリオリジンから派生
--redact なし 削除対象の追加ヘッダー名/JSONキー(カンマ区切り)
--min-samples なし エンドポイントごとに含めるサンプルの最小数。デフォルトは1
--stage なし 1つのステージのみを実行します:ロード、フィルタリング、正規化、推論、出力

出力レイアウト

/api-spec/
├── index.html                ビジュアルレポート — これを開いてください(スタンドアロン、サーバー不要)
├── client.mjs                操作ごとに型指定された関数を持つ、依存関係ゼロのフェッチクライアント
├── openapi.yaml              機械可読な仕様書
├── openapi.json              ミラー
├── report.md                 マークダウン形式の概要 + curl サンプル
├── confidence.json           エンドポイントごとの信頼度 + 正規化フラグ
├── samples/                  一部非公開のリクエスト/レスポンスのサンプル
│   └──__.json
└── intermediate/             パイプラインの副産物(ペア化/フィルタリング済み/エンドポイントの jsonl)

「browse cdp」および「browse network」から得られるもの

2つの相互に補完し合うキャプチャソース:

ソース 提供内容 制限事項
browse cdp(browser-traceで使用) リクエストメソッド/URL/ヘッダー/POSTデータ、レスポンスステータス/ヘッダー/MIMEタイプ、完全なイベントタイミング レスポンス本文は埋め込まれません。本文はNetwork.getResponseBody を使用して取得する必要がありますが、firehose ではこの処理は行われません。
browse network on(別のコマンド) ディスク上のリクエスト本文およびレスポンス本文(CDPリクエストIDをキーとして保存) キャプチャディレクトリはブラウズセッションごとに共有されます。別の `browse network on` を実行する前のスナップショットは上書きされます。

discover.mjsは、`--bodies` オプションを指定して` `を渡すと、browse networkのディレクトリからレスポンス本体を取得します(または、自動検出される `/cdp/network/bodies/` に保存することも可能です)。照合は`requestId`によって行われます。browse networkはこれを各`request.json`に`id` として書き込み、当ツールはそれを直接結合します。

ボディが存在する場合の変更点:

  • ✅ パステンプレーティング、クエリパラメータスキーマ、ステータスコード、コンテンツタイプ — どちらの場合でも同じです。
  • ✅ リクエスト・ボディのスキーマ — CDPからのpostDataがあれば十分です。postData以外の場合にボディディレクトリがあるのは便利ですが、必須ではありません。
  • ✅レスポンスボディのスキーマ— 実際のサンプルから完全に推論されます。ボディがない場合は、{ description, content: }という骨格が得られます。

このレポートでは、レスポンスボディのサンプルがないエンドポイントをすべてフラグ付けします。

ノイズの自動フィルタリング

正規化段階では、インフラストラクチャのノイズを自動的に分類して除外します:

  • トラッキング/分析—/track、/pixel、/beacon、/impression、/pageview、/dag/v*を含むパス
  • ボット防御— Akamai (/akam/)、フィンガープリントペイロード (sensor_data)、難読化されたマルチセグメントパス
  • セッション関連—/session、/authenticate/start、クッキー同意、A/Bテストエンドポイント
  • HTMLページのレンダリング—text/htmlを返すGETリクエスト(APIではなく、レンダリングされたページ)

これにより、通常、キャプチャされたトラフィックの60~80%が除外されます。--includeフラグを使用することで、誤検知を回避できます。

GraphQL / マルチプレックス化されたエンドポイントの分解

単一のエンドポイント(例:/dapi/fe/gql)が異なるoperationName値で呼び出された場合、本スキルはそれを自動的に個別の論理操作に分割します。それぞれに以下が割り当てられます:

  • OpenAPIパスエントリ(例:/dapi/fe/gql [オートコンプリート])
  • その操作のサンプルのみから推論されたリクエスト/レスポンススキーマ
  • レポート内のCurl例および変数テーブル

検出は、ボディフィールド(operationName、method、action)およびクエリパラメータ(opname、op)に基づいて行われます。これにより、GraphQL(APQおよびインライン)、JSON-RPC、および類似のディスパッチパターンがカバーされます。

制限事項

  • 検出範囲はキャプチャされたフローに限定されます。トレース内で実行されなかったエンドポイントは表示されません。本スキルは完全性を保証できません。
  • スキーマは帰納的であり、契約に基づくものではありません。すべてのサンプルにフィールドが含まれていたとしても、サーバー側ではそのフィールドがオプションである可能性があります。
  • 認証は観測されるものであり、指定されるものではありません。このスキルは、x-observed-auth拡張子に認証形式のヘッダーを記録しますが、セキュリティスキームを保証するものではありません。
  • パステンプレート化はヒューリスティックです。数値、UUID、16進数、スラッグのパターンはセグメントごとに検出されます。曖昧なURLはconfidence.json 内でフラグが立てられます。
  • 情報隠蔽はベストエフォートです。デフォルトの隠蔽処理では一般的な認証情報が対象となりますが、アプリ固有のシークレットが漏れる可能性があります。既知のカスタムヘッダーやキーについては、`--redact`オプションを使用してください。

ベストプラクティス

  1. 記録したいフローを意図的に誘導してください。ブラウザトレースが詳細であればあるほど、仕様も詳細になります。
  2. ノイズの多いサイトには--originsオプションを使用してください。マーケティングページは数十ものアナリティクスホストにアクセスすることがありますが、関心のある API オリジンに制限してください。
  3. まずreport.mdを確認してください。ここには、検出されたすべての操作に対する curl ですぐに使える例やレスポンスのサンプルが記載されています。
  4. 最終ドキュメントに確実に特定されたエンドポイントのみを含めたい場合は、 `--min-samples` を2 以上に設定してください。ロングテールのデータは除外されます。
  5. レスポンス本体のスキーマが重要な場合は、「ネットワークの閲覧」を有効にして併用してください。CDPのファイアホースだけでは、リクエスト本体は取得できますが、レスポンス本体は取得できません。

パイプラインの内部構造やファイル形式のリファレンスについては、REFERENCE.mdを参照してください。

GitHubで見る
---
name: browser-to-api
description: Generate an OpenAPI 3.1 specification from a browser-trace capture by analyzing observed HTTP traffic, templating URLs, and inferring JSON schemas from request/response samples.
license: MIT
---

# Browser to API

Replay-driven API discovery. Consume a `browser-trace` capture, pair its CDP request / response events, templatize observed URLs, infer JSON schemas from samples, and emit an **OpenAPI 3.1** document plus a human-readable coverage report.

This skill **does not capture traffic**. It is purely offline post-processing on top of `browser-trace`'s `cdp/network/*.jsonl` buckets. The two skills compose:

```
browser-trace    →  .o11y/<run>/cdp/network/{requests,responses}.jsonl
browser-to-api   →  .o11y/<run>/api-spec/index.html + openapi.yaml + client.mjs
```

## When to use

- The user wants an OpenAPI document for a third-party or undocumented website API.
- The user has a `browser-trace` run and wants endpoints + schemas extracted from it.
- The user is building a client/SDK against a site that doesn't publish a spec.
- The user wants a coverage report showing which flows would broaden the spec.

If the user wants to **capture** traffic, send them to `browser-trace` first.

## Two-step workflow

### 1. Capture with `browser-trace` (and optionally bodies via `browse network on`)

```bash
# Local example against an existing debuggable Chrome target
TARGET=9222

node ../browser-trace/scripts/start-capture.mjs "$TARGET" my-site
browse open about:blank --cdp "$TARGET"
browse network on                                    # capture request/response bodies
browse open https://example.com
# ...drive whatever flows you want covered...

# Snapshot the bodies dir BEFORE turning capture off (the temp dir is shared
# per-session, so subsequent `browse network on` runs would mix your bodies
# with whatever a future capture writes if you skip this step).
cp -r "$(browse network path | jq -r .path)" .o11y/my-site/cdp/network/bodies/
browse network off

node ../browser-trace/scripts/stop-capture.mjs my-site
node ../browser-trace/scripts/bisect-cdp.mjs my-site
```

`browse network on` is **optional but strongly recommended** — without it, the spec has no response-body schemas (the CDP firehose used by `browse cdp` does not embed bodies). With it, both request bodies (already captured by CDP) *and* response bodies are joined into the trace by CDP `requestId`.

### 2. Generate the spec

```bash
node scripts/discover.mjs --run .o11y/my-site
# → .o11y/my-site/api-spec/index.html          ← open this
#   .o11y/my-site/api-spec/client.mjs
#   .o11y/my-site/api-spec/openapi.yaml
#   .o11y/my-site/api-spec/openapi.json
#   .o11y/my-site/api-spec/report.md
#   .o11y/my-site/api-spec/confidence.json
#   .o11y/my-site/api-spec/samples/*.json
#   .o11y/my-site/api-spec/intermediate/*.jsonl
```

`discover.mjs` auto-detects `<run>/cdp/network/bodies/`. To use a body capture from elsewhere (e.g. didn't snapshot, want the live `browse network` dir), pass `--bodies <path>` explicitly.

### 3. Open the HTML report

After `discover.mjs` finishes, **always open the generated HTML report**:

```bash
open .o11y/my-site/api-spec/index.html
```

The report is a self-contained HTML file (no server needed) that shows each discovered operation as an expandable card with variables, client usage, request/response examples, and a generated `client.mjs` snippet at the bottom. This is the primary deliverable — always open it for the user.

## CLI flags

| Flag | Required | Meaning |
|---|---|---|
| `--run <path>` | yes | Path to a `browser-trace` run directory |
| `--out <path>` | no | Output dir; default `<run>/api-spec/` |
| `--bodies <path>` | no | `browse network` capture dir to join into the trace (auto-detected from `<run>/cdp/network/bodies/` when present) |
| `--include <regex>` | no | Only include URLs matching regex (repeatable) |
| `--exclude <regex>` | no | Exclude URLs matching regex (repeatable; in addition to defaults) |
| `--origins <list>` | no | Comma-separated origin allow-list (e.g. `api.example.com,example.com`) |
| `--format <yaml\|json\|both>` | no | Output format. Default `both` |
| `--title <string>` | no | OpenAPI `info.title`. Default derived from primary origin |
| `--redact <list>` | no | Extra header names / JSON keys to redact (comma-separated) |
| `--min-samples <n>` | no | Minimum samples per endpoint to include. Default `1` |
| `--stage <name>` | no | Run only one stage: `load`, `filter`, `normalize`, `infer`, `emit` |


## Output layout

```
<run>/api-spec/
├── index.html                visual report — open this (self-contained, no server)
├── client.mjs                zero-dep fetch client with typed functions per operation
├── openapi.yaml              machine-readable spec
├── openapi.json              mirror
├── report.md                 markdown summary + curl examples
├── confidence.json           per-endpoint confidence + normalization flags
├── samples/                  redacted request/response examples
│   └── <method>__<path-hash>.json
└── intermediate/             pipeline byproducts (paired/filtered/endpoints jsonl)
```

## What you get from `browse cdp` and `browse network`

Two complementary capture sources:

| Source | Provides | Limitation |
|---|---|---|
| `browse cdp` (used by `browser-trace`) | request method/URL/headers/`postData`, response status/headers/mimeType, full event timing | **Does not embed response bodies.** Bodies must be pulled with `Network.getResponseBody`, which the firehose doesn't do. |
| `browse network on` (separate command) | request bodies AND response bodies on disk, keyed by CDP `requestId` | Capture dir is shared per `browse` session; snapshot before another `browse network on` overwrites it. |

`discover.mjs` will pull bodies from a `browse network` dir if you pass `--bodies <path>` (or stash them under `<run>/cdp/network/bodies/`, which is auto-detected). The matching is by `requestId` — `browse network` writes that into each `request.json` as `id`, and we join directly.

What changes when bodies are present:

- ✅ Path templating, query-param schemas, status codes, content-types — same either way.
- ✅ Request-body schemas — `postData` from CDP is enough; bodies dir is a nice-to-have for non-`postData` cases.
- ✅ **Response-body schemas** — fully inferred from real samples. Without bodies you get `{ description, content: <mimeType> }` skeletons.

The report flags every endpoint that has no response-body sample.

## Automatic noise filtering

The normalize stage automatically classifies and drops infrastructure noise:

- **Tracking / analytics** — paths containing `/track`, `/pixel`, `/beacon`, `/impression`, `/pageview`, `/dag/v*`
- **Bot defense** — Akamai (`/akam/`), fingerprint payloads (`sensor_data`), obfuscated multi-segment paths
- **Session plumbing** — `/session`, `/authenticate/start`, cookie consent, A/B experiment endpoints
- **HTML page renders** — `GET` requests returning `text/html` (the rendered page, not the API)

This typically drops 60-80% of captured traffic. The `--include` flag can rescue a false positive.

## GraphQL / multiplexed endpoint decomposition

When a single endpoint (like `/dapi/fe/gql`) is called with different `operationName` values, the skill automatically splits it into separate logical operations. Each gets its own:
- OpenAPI path entry (e.g. `/dapi/fe/gql [Autocomplete]`)
- Request/response schema inferred from only that operation's samples
- Curl example and variables table in the report

Detection works on body fields (`operationName`, `method`, `action`) and query params (`opname`, `op`). This covers GraphQL (APQ and inline), JSON-RPC, and similar dispatch patterns.

## Limitations

- **Coverage is bounded by the captured flow.** Endpoints not exercised in the trace will not appear. The skill cannot prove completeness.
- **Schemas are inductive, not contractual.** A field might be optional on the server even if every sample contained it.
- **Auth is observed, not specified.** The skill records auth-shaped headers in an `x-observed-auth` extension but won't claim a security scheme.
- **Path templating is heuristic.** Numeric / UUID / hex / slug patterns are detected per segment. Ambiguous URLs are flagged in `confidence.json`.
- **Redaction is best-effort.** Default redactions cover common credentials, but app-specific secrets may slip through; use `--redact` for known custom headers/keys.

## Best practices

1. **Drive the flows you want documented.** The richer the browser-trace, the richer the spec.
2. **Use `--origins` for noisy sites.** A marketing page hits dozens of analytics hosts; restrict to the API origin you care about.
3. **Inspect `report.md` first.** It has curl-ready examples and response samples for every discovered operation.
4. **Bump `--min-samples` to 2+** when you want only confidently-shaped endpoints in the final doc — drop the long tail.
5. **Pair with `browse network on`** when response-body schemas matter. The CDP firehose alone has request bodies but not response bodies.

For pipeline internals and the file format reference, see [REFERENCE.md](REFERENCE.md).

browser-to-apiをインストール

スキルファイルをダウンロードし、.claude/skills/ ディレクトリに解凍してください。

ZIPをダウンロード

リポジトリをクローンし、スキルファイルをプロジェクトにコピーしてください。

git clone https://github.com/browserbase/skills/tree/main/skills/browser-to-api # Copy SKILL.md to your .claude/skills/ directory

コピー コピー
クイックセットアップ: スキルフォルダを .claude/skills/ にコピーしてください。 Claude が自動的にスキルを検出して使用します。
リポジトリ browserbase/skills

関連スキル

agentwallet
更新された時間 2026年7月7日
brightdata-cli
更新された時間 2026年6月29日
humanize
更新された時間 2026年7月7日
korean-stock-search
更新された時間 2026年7月8日
OR