browser-to-api
browserbase/skills
관측된 HTTP 트래픽을 분석하고, URL을 템플릿화하며, 요청/응답 샘플로부터 JSON 스키마를 추론하여 브라우저 추적 데이터를 기반으로 OpenAPI 3.1 사양을 생성합니다.
...모든 것을 확장하십시오브라우저에서 API로
리플레이 기반 API 탐색. 브라우저 트레이스 캡처를 분석하여 CDP 요청/응답 이벤트를 매칭하고, 관찰된 URL을 템플릿화하며, 샘플로부터 JSON 스키마를 추론한 후, OpenAPI 3.1 문서와 사람이 읽기 쉬운 커버리지 보고서를 생성합니다.
이 스킬은 트래픽을 캡처하지 않습니다. 이는 브라우저 트레이스의 cdp/network/*.jsonl 버킷을 기반으로 하는 순수한 오프라인 후처리입니다. 두 스킬은 다음과 같이 조합됩니다:
browser-trace → .o11y//cdp/network/{requests,responses}.jsonl
browser-to-api → .o11y//api-spec/index.html + openapi.yaml + client.mjs
사용 시점
- 사용자가 타사 또는 문서화되지 않은 웹사이트 API에 대한 OpenAPI 문서를 원할 때.
- 사용자가
브라우저 추적결과를 확보한 상태에서, 해당 결과에서 엔드포인트와 스키마를 추출하고자 할 때. - 사용자가 사양을 공개하지 않는 사이트를 대상으로 클라이언트/SDK를 구축하고 있는 경우.
- 사용자가 어떤 흐름이 사양을 확장할 수 있는지 보여주는 커버리지 보고서를 원할 때.
사용자가 트래픽을 캡처하고자 한다면, 먼저 브라우저 트레이스로 안내하십시오.
2단계 워크플로
1. browser-trace를 사용하여 캡처합니다(선택 사항: ‘네트워크 탐색’을 켜고 본문 데이터도 캡처).
# 디버깅 가능한 기존 Chrome 타깃을 대상으로 한 로컬 예시
TARGET=9222
node ../browser-trace/scripts/start-capture.mjs "$TARGET" my-site
browse open about:blank --cdp "$TARGET"
browse network on # 요청/응답 본문 캡처
browse open https://example.com
# ...캡처하고 싶은 트래픽 흐름을 생성합니다...
# 캡처를 끄기 전에 본문 디렉터리를 스냅샷으로 저장합니다(임시 디렉터리는
# 세션별로 공유되므로, 이 단계를 건너뛰면 이후 `browse network on` 실행 시
# 본문 데이터가 향후 캡처에서 기록되는 내용과 혼합될 수 있습니다).
cp -r "$(browse network path | jq -r .path)" .o11y/my-site/cdp/network/bodies/
browse network off
node ../browser-trace/scripts/stop-capture.mjs my-site
node ../browser-trace/scripts/bisect-cdp.mjs my-site
browse network on은 선택 사항이지만 강력히 권장됩니다. 이 명령을 생략하면 사양에 응답 본문 스키마가 포함되지 않습니다( browse cdp에서 사용하는 CDP 파이어호스는 본문을 포함하지 않음). 이 명령을 사용하면 (이미 CDP에 캡처된) 요청 본문과 응답 본문이 모두 CDP requestId를 기준으로 트레이스에 병합됩니다.
2. 사양 생성
node scripts/discover.mjs --run .o11y/my-site
# → .o11y/my-site/api-spec/index.html ← 이 파일을 엽니다
# .o11y/my-site/api-spec/client.mjs
# .o11y/my-site/api-spec/openapi.yaml
# .o11y/my-site/api-spec/openapi.json
# .o11y/my-site/api-spec/report.md
# .o11y/my-site/api-spec/confidence.json
# .o11y/my-site/api-spec/samples/*.json
# .o11y/my-site/api-spec/intermediate/*.jsonl
discover.mjs는 자동으로 감지합니다 . 다른 위치의 바디 캡처를 사용하려면(예: 스냅샷을 생성하지 않았거나, 라이브 브라우징 네트워크 디렉터리를 사용하려는 경우), --bodies 명시적으로 전달하십시오.
3. HTML 보고서 열기
discover.mjs가 완료된 후에는 항상 생성된 HTML 보고서를 열어주세요:
.o11y/my-site/api-spec/index.html 열기
이 보고서는 독립형 HTML 파일(서버 불필요)로, 탐지된 각 연산을 확장 가능한 카드 형태로 표시하며, 각 카드에는 변수, 클라이언트 사용법, 요청/응답 예시, 그리고 하단에 생성된 client.mjs 스니펫이 포함되어 있습니다. 이것이 주요 결과물이므로, 사용자에게 항상 이 보고서를 열어 보여주어야 합니다.
CLI 플래그
| 옵션 | 필수 | 의미 |
|---|---|---|
--run |
yes | 브라우저 추적 실행 디렉터리의 경로 |
--out |
아니요 | 출력 디렉터리; 기본값 |
--bodies |
없음 | 추적에 포함할네트워크 캡처 디렉터리를탐색합니다 (해당 디렉터리가 존재할 경우 /cdp/network/bodies/에서 자동으로 감지됨) |
--include |
아니요 | 정규 표현식과 일치하는 URL만 포함합니다(반복 가능). |
--exclude |
아니요 | 정규 표현식과 일치하는 URL을 제외합니다(반복 가능; 기본값에 추가로 적용됨) |
--origins |
아니요 | 쉼표로 구분된 허용 원본 목록 (예: api.example.com,example.com) |
--format |
없음 | 출력 형식. 기본값: both |
--title |
없음 | OpenAPI info.title. 기본값은 주 출처에서 파생됨 |
--redact |
아니요 | 가려야 할 추가 헤더 이름/JSON 키 (쉼표로 구분) |
--min-samples |
없음 | 엔드포인트당 포함할 최소 샘플 수. 기본값 1 |
--stage |
없음 | 단일 단계만 실행: 로드, 필터링, 정규화, 추론, 출력 |
출력 레이아웃
/api-spec/
├── index.html 시각적 보고서 — 이 파일을 열어보세요 (자립형, 서버 불필요)
├── client.mjs 작업별 타입 지정된 함수를 사용하는 zero-dep Fetch 클라이언트
├── openapi.yaml 기계 가독형 사양
├── openapi.json 미러
├── report.md 마크다운 요약 + curl 예제
├── confidence.json 엔드포인트별 신뢰도 + 정규화 플래그
├── samples/ 정보가 삭제된 요청/응답 예제
│ └── __.json
└── intermediate/ 파이프라인 부산물 (페어링/필터링된 엔드포인트 jsonl)
'browse cdp ' 및 'browse network'를 통해 얻을 수 있는 정보
서로 보완하는 두 가지 수집 소스:
| 소스 | 제공 내용 | 제한 사항 |
|---|---|---|
browse cdp ( browser-trace에서 사용) |
요청메서드/URL/헤더/POST 데이터, 응답 상태/헤더/MIME 유형, 전체 이벤트 타이밍 |
응답 본문은 포함되지 않습니다. 본문은 Network.getResponseBody를 사용하여 가져와야 하며, 파이어호스는 이를 수행하지 않습니다. |
네트워크 보기 켜기 (별도의 명령) |
디스크에 저장된 요청 본문 및 응답 본문(CDP requestId를 키로 사용) |
캡처 디렉터리는 브라우즈 세션별로 공유되며, 다음 'browse network on' 명령을 실행하기 전의 스냅샷이 이를 덮어씁니다. |
discover.mjs는 --bodies 전달하면 '네트워크 탐색' 디렉터리의 본문을 가져옵니다(또는 저장해 두면 자동으로 감지됩니다). 일치 여부는 requestId를 기준으로 판단되며, '네트워크 탐색'은 이를 각 request.json 파일에 id로 기록하고, 이를 직접 조인합니다.
본문이 존재할 때 변경되는 사항:
- ✅ 경로 템플릿, 쿼리 매개변수 스키마, 상태 코드, 콘텐츠 유형 — 두 경우 모두 동일합니다.
- ✅ 요청 본문 스키마 — CDP의
postData만으로도충분하며, 본문 디렉터리는postData가 아닌경우에 있으면 좋은 기능입니다. - ✅ 응답 본문 스키마 — 실제 샘플로부터 완전히 추론됩니다. 본문이 없는 경우
{ description, content:형태의 기본 구조가 제공됩니다.}
이 보고서는 응답 본문 샘플이 없는 모든 엔드포인트를 표시합니다.
자동 노이즈 필터링
정규화 단계에서는 인프라 노이즈를 자동으로 분류하고 제거합니다:
- 추적/분석 —
/track,/pixel,/beacon,/impression,/pageview,/dag/v*가포함된 경로 - 봇 방어 — Akamai (
/akam/), 지문 페이로드 (sensor_data), 난독화된 다중 세그먼트 경로 - 세션 관련 —
/session,/authenticate/start, 쿠키 동의, A/B 테스트 엔드포인트 - HTML 페이지 렌더링 —
text/html을반환하는GET요청(API가 아닌 렌더링된 페이지)
이 필터링으로 일반적으로 캡처된 트래픽의 60~80%가 제외됩니다. --include 플래그를 사용하면 오탐을 해결할 수 있습니다.
GraphQL / 다중화 엔드포인트 분해
단일 엔드포인트(예: /dapi/fe/gql)가 서로 다른 operationName 값으로 호출되면, 이 기능은 이를 자동으로 별도의 논리적 작업으로 분할합니다. 각 작업은 다음과 같이 구성됩니다:
- OpenAPI 경로 항목(예:
/dapi/fe/gql [자동 완성]) - 해당 작업의 샘플만을 기반으로 추론된 요청/응답 스키마
- 보고서 내의 curl 예시 및 변수 표
감지 기능은 본문 필드(operationName, method, action)와 쿼리 매개변수(opname, op)를 기반으로 작동합니다. 이는 GraphQL(APQ 및 인라인), JSON-RPC 및 유사한 디스패치 패턴을 모두 포함합니다.
제한 사항
- 범위는 캡처된 흐름에 의해 제한됩니다. 추적에서 실행되지 않은 엔드포인트는 표시되지 않습니다. 이 기능은 완전성을 보장할 수 없습니다.
- 스키마는 계약적(contractual)이 아닌 귀납적입니다. 모든 샘플에 특정 필드가 포함되어 있더라도 서버 측에서는 해당 필드가 선택적일 수 있습니다.
- 인증 정보는 명시된 것이 아니라 관찰된 것입니다. 이 스킬은
x-observed-auth확장자에 인증 형식의 헤더를 기록하지만, 보안 체계를 보장하지는 않습니다. - 경로 템플릿화는 휴리스틱 방식입니다. 숫자/UUID/16진수/슬러그 패턴은 세그먼트별로 감지됩니다. 모호한 URL은
confidence.json파일에 표시됩니다. - 정보 마스킹은 최선을 다해 수행됩니다. 기본 마스킹은 일반적인 자격 증명을 포함하지만, 앱별 비밀 정보는 누락될 수 있습니다. 알려진 사용자 정의 헤더/키의 경우
--redact 옵션을사용하십시오.
모범 사례
- 문서화하고 싶은 흐름을 유도하십시오. 브라우저 추적 정보가 풍부할수록 사양도 더 풍부해집니다.
- 잡음이 많은 사이트의 경우
--origins 옵션을사용하십시오. 마케팅 페이지는 수십 개의 분석 호스트를 호출하므로, 관심 있는 API 오리진으로만 제한하십시오. - 먼저
report.md를확인하십시오. 여기에는 발견된 모든 작업에 대한 curl 실행 가능한 예제와 응답 샘플이 포함되어 있습니다. - 최종 문서에 형태가 확실한 엔드포인트만 포함하려면
--min-samples 값을2 이상으로 높여 긴 꼬리 부분을 제외하십시오. - 응답 본문 스키마가 중요한
경우에는 네트워크 탐색 기능을함께 사용하십시오. CDP 파이어호스만으로는 요청 본문은 제공되지만 응답 본문은 제공되지 않습니다.
파이프라인 내부 구조 및 파일 형식 참조는 REFERENCE.md를 참조하십시오.
---
name: browser-to-api
description: Generate an OpenAPI 3.1 specification from a browser-trace capture by analyzing observed HTTP traffic, templating URLs, and inferring JSON schemas from request/response samples.
license: MIT
---
# Browser to API
Replay-driven API discovery. Consume a `browser-trace` capture, pair its CDP request / response events, templatize observed URLs, infer JSON schemas from samples, and emit an **OpenAPI 3.1** document plus a human-readable coverage report.
This skill **does not capture traffic**. It is purely offline post-processing on top of `browser-trace`'s `cdp/network/*.jsonl` buckets. The two skills compose:
```
browser-trace → .o11y/<run>/cdp/network/{requests,responses}.jsonl
browser-to-api → .o11y/<run>/api-spec/index.html + openapi.yaml + client.mjs
```
## When to use
- The user wants an OpenAPI document for a third-party or undocumented website API.
- The user has a `browser-trace` run and wants endpoints + schemas extracted from it.
- The user is building a client/SDK against a site that doesn't publish a spec.
- The user wants a coverage report showing which flows would broaden the spec.
If the user wants to **capture** traffic, send them to `browser-trace` first.
## Two-step workflow
### 1. Capture with `browser-trace` (and optionally bodies via `browse network on`)
```bash
# Local example against an existing debuggable Chrome target
TARGET=9222
node ../browser-trace/scripts/start-capture.mjs "$TARGET" my-site
browse open about:blank --cdp "$TARGET"
browse network on # capture request/response bodies
browse open https://example.com
# ...drive whatever flows you want covered...
# Snapshot the bodies dir BEFORE turning capture off (the temp dir is shared
# per-session, so subsequent `browse network on` runs would mix your bodies
# with whatever a future capture writes if you skip this step).
cp -r "$(browse network path | jq -r .path)" .o11y/my-site/cdp/network/bodies/
browse network off
node ../browser-trace/scripts/stop-capture.mjs my-site
node ../browser-trace/scripts/bisect-cdp.mjs my-site
```
`browse network on` is **optional but strongly recommended** — without it, the spec has no response-body schemas (the CDP firehose used by `browse cdp` does not embed bodies). With it, both request bodies (already captured by CDP) *and* response bodies are joined into the trace by CDP `requestId`.
### 2. Generate the spec
```bash
node scripts/discover.mjs --run .o11y/my-site
# → .o11y/my-site/api-spec/index.html ← open this
# .o11y/my-site/api-spec/client.mjs
# .o11y/my-site/api-spec/openapi.yaml
# .o11y/my-site/api-spec/openapi.json
# .o11y/my-site/api-spec/report.md
# .o11y/my-site/api-spec/confidence.json
# .o11y/my-site/api-spec/samples/*.json
# .o11y/my-site/api-spec/intermediate/*.jsonl
```
`discover.mjs` auto-detects `<run>/cdp/network/bodies/`. To use a body capture from elsewhere (e.g. didn't snapshot, want the live `browse network` dir), pass `--bodies <path>` explicitly.
### 3. Open the HTML report
After `discover.mjs` finishes, **always open the generated HTML report**:
```bash
open .o11y/my-site/api-spec/index.html
```
The report is a self-contained HTML file (no server needed) that shows each discovered operation as an expandable card with variables, client usage, request/response examples, and a generated `client.mjs` snippet at the bottom. This is the primary deliverable — always open it for the user.
## CLI flags
| Flag | Required | Meaning |
|---|---|---|
| `--run <path>` | yes | Path to a `browser-trace` run directory |
| `--out <path>` | no | Output dir; default `<run>/api-spec/` |
| `--bodies <path>` | no | `browse network` capture dir to join into the trace (auto-detected from `<run>/cdp/network/bodies/` when present) |
| `--include <regex>` | no | Only include URLs matching regex (repeatable) |
| `--exclude <regex>` | no | Exclude URLs matching regex (repeatable; in addition to defaults) |
| `--origins <list>` | no | Comma-separated origin allow-list (e.g. `api.example.com,example.com`) |
| `--format <yaml\|json\|both>` | no | Output format. Default `both` |
| `--title <string>` | no | OpenAPI `info.title`. Default derived from primary origin |
| `--redact <list>` | no | Extra header names / JSON keys to redact (comma-separated) |
| `--min-samples <n>` | no | Minimum samples per endpoint to include. Default `1` |
| `--stage <name>` | no | Run only one stage: `load`, `filter`, `normalize`, `infer`, `emit` |
## Output layout
```
<run>/api-spec/
├── index.html visual report — open this (self-contained, no server)
├── client.mjs zero-dep fetch client with typed functions per operation
├── openapi.yaml machine-readable spec
├── openapi.json mirror
├── report.md markdown summary + curl examples
├── confidence.json per-endpoint confidence + normalization flags
├── samples/ redacted request/response examples
│ └── <method>__<path-hash>.json
└── intermediate/ pipeline byproducts (paired/filtered/endpoints jsonl)
```
## What you get from `browse cdp` and `browse network`
Two complementary capture sources:
| Source | Provides | Limitation |
|---|---|---|
| `browse cdp` (used by `browser-trace`) | request method/URL/headers/`postData`, response status/headers/mimeType, full event timing | **Does not embed response bodies.** Bodies must be pulled with `Network.getResponseBody`, which the firehose doesn't do. |
| `browse network on` (separate command) | request bodies AND response bodies on disk, keyed by CDP `requestId` | Capture dir is shared per `browse` session; snapshot before another `browse network on` overwrites it. |
`discover.mjs` will pull bodies from a `browse network` dir if you pass `--bodies <path>` (or stash them under `<run>/cdp/network/bodies/`, which is auto-detected). The matching is by `requestId` — `browse network` writes that into each `request.json` as `id`, and we join directly.
What changes when bodies are present:
- ✅ Path templating, query-param schemas, status codes, content-types — same either way.
- ✅ Request-body schemas — `postData` from CDP is enough; bodies dir is a nice-to-have for non-`postData` cases.
- ✅ **Response-body schemas** — fully inferred from real samples. Without bodies you get `{ description, content: <mimeType> }` skeletons.
The report flags every endpoint that has no response-body sample.
## Automatic noise filtering
The normalize stage automatically classifies and drops infrastructure noise:
- **Tracking / analytics** — paths containing `/track`, `/pixel`, `/beacon`, `/impression`, `/pageview`, `/dag/v*`
- **Bot defense** — Akamai (`/akam/`), fingerprint payloads (`sensor_data`), obfuscated multi-segment paths
- **Session plumbing** — `/session`, `/authenticate/start`, cookie consent, A/B experiment endpoints
- **HTML page renders** — `GET` requests returning `text/html` (the rendered page, not the API)
This typically drops 60-80% of captured traffic. The `--include` flag can rescue a false positive.
## GraphQL / multiplexed endpoint decomposition
When a single endpoint (like `/dapi/fe/gql`) is called with different `operationName` values, the skill automatically splits it into separate logical operations. Each gets its own:
- OpenAPI path entry (e.g. `/dapi/fe/gql [Autocomplete]`)
- Request/response schema inferred from only that operation's samples
- Curl example and variables table in the report
Detection works on body fields (`operationName`, `method`, `action`) and query params (`opname`, `op`). This covers GraphQL (APQ and inline), JSON-RPC, and similar dispatch patterns.
## Limitations
- **Coverage is bounded by the captured flow.** Endpoints not exercised in the trace will not appear. The skill cannot prove completeness.
- **Schemas are inductive, not contractual.** A field might be optional on the server even if every sample contained it.
- **Auth is observed, not specified.** The skill records auth-shaped headers in an `x-observed-auth` extension but won't claim a security scheme.
- **Path templating is heuristic.** Numeric / UUID / hex / slug patterns are detected per segment. Ambiguous URLs are flagged in `confidence.json`.
- **Redaction is best-effort.** Default redactions cover common credentials, but app-specific secrets may slip through; use `--redact` for known custom headers/keys.
## Best practices
1. **Drive the flows you want documented.** The richer the browser-trace, the richer the spec.
2. **Use `--origins` for noisy sites.** A marketing page hits dozens of analytics hosts; restrict to the API origin you care about.
3. **Inspect `report.md` first.** It has curl-ready examples and response samples for every discovered operation.
4. **Bump `--min-samples` to 2+** when you want only confidently-shaped endpoints in the final doc — drop the long tail.
5. **Pair with `browse network on`** when response-body schemas matter. The CDP firehose alone has request bodies but not response bodies.
For pipeline internals and the file format reference, see [REFERENCE.md](REFERENCE.md).
모든 파일
15개 파일browser-to-api 설치
스킬 파일을 다운로드하여 .claude/skills/ 디렉터리에 압축을 풀어주세요.
ZIP 다운로드저장소를 클론하고 스킬 파일을 프로젝트에 복사하세요.
git clone https://github.com/browserbase/skills/tree/main/skills/browser-to-api # Copy SKILL.md to your .claude/skills/ directory
복사





집
