选项
首页首页 Skill 开发者工具 browser-use-to-stagehand

browser-use-to-stagehand

browserbase/skills browserbase/skills

将 Browserbase 上的浏览器操作(Python)自动化脚本转换为 Stagehand v3(TypeScript),并在可行的情况下,用确定性管道替换不透明的代理循环。

...展开全部
0
更新时间 2026-09-30

browser-use → Browserbase 上的 Stagehand (/browser-use-to-stagehand)

将一个 browser-use(Python)脚本转换为 Browserbase 上的符合 Stagehand v3 惯用风格的(TypeScript)脚本, 在每个步骤中选择合适的确定性级别,而非生成 完全一比一的代理式副本。

核心原则:browser-use 默认采用代理模式(由大型语言模型决定每项操作)。Stagehand 则允许您自主选择 AI 的介入程度。 一次成功的迁移应将不透明的代理循环替换为 可检查且基本确定性的处理流程——仅在页面确实 不可预测时才使用 AI。这是一次需要判断力的重构,而非简单的转译。

权威来源与版本。该技能的持久价值在于判断力——即确定性 的范围以及“分解”与“代理”之间的决策——而非 API 细节,后者在每次发布中都会发生变化。 此处的代码映射是针对 @browserbasehq/stagehand 3.6.x 和 browser-use 0.13.x(2026-06)验证的快照。如有冲突,以实时文档为准——在生成代码前,请务必对照 已安装的包和这些源文件进行验证:

  • Stagehand v3:https://docs.stagehand.dev/v3 · 已安装类型: node_modules/@browserbasehq/stagehand
  • Browserbase:https://docs.browserbase.com
  • browser-use:https://docs.browser-use.com

如果已安装的 Stagehand 主版本号不是 3,请将此技能视为仅供概念参考,并针对每个签名 遵循在线文档。

参考文件(按需阅读)

  • references/api-mapping.md — 技术层面的 browser-use → Stagehand 映射:变体检测、完整特性表、代码对比、Browserbase 平台 选项以及 v3 版本的注意事项。处理任何非平凡的构造时请阅读此文档。
  • references/determinism.md — 如何选择 agent() 与 act/extract/observe vs 缓存 observe→act。决策树。在决定 如何转换Agent(task=…)时请阅读此部分。
  • references/trace-assisted.md — 针对不透明/不稳定脚本的可选工作流:“在 Browserbase 上运行、阅读日志,然后重写”。
  • references/guide.md — 面向人类的迁移指南:理念转变、 功能映射、确定性光谱以及推荐的迁移路径。
  • references/prompt.md — 该技能的自包含、与工具无关的版本; 将其与浏览器使用脚本一起粘贴到任何 AI 助手中。
  • EXAMPLES.md — 脚本的“迁移前/迁移后”配对示例。

工作流

1. 获取源代码

获取浏览器使用脚本。如果用户仅描述了脚本,请索要相关文件。注意 目标格式:除非另有说明,否则应为基于 Browserbase 的 TypeScript Stagehand。

首先,判断适用范围——这是否可迁移?并非每个 browser-use 文件都是 Agent(task=…) 脚本。如果源文件是作为 MCP 服务器运行的 browser-use (uvx browser-use --mcp、 mcpServers 配置),则没有 Stagehand 的等效实现——将其标记为 超出范围,不要自行发明(参见 api-mapping §3.7b)。 如果 browser-use 调用嵌入在 更大的应用程序中(类/工具封装、Web 路由、队列任务),则仅转换 browser-use 的接口,并 保留周围的应用程序粘合代码——参见 api-mapping §3.8。

2. 检测 browser-use 的变体

区分旧版(0.12 之前)、稳定版和 Rust 测试版(仅当导入来自 browser_use.beta) ——参见 api-mapping §1。注意:经典的顶级 from browser_use import Agent, ChatBrowserUse 接口在 0.13.x 版本中依然有效—— ChatBrowserUse 仅凭这一点无法判断是否为测试版;只有 browser_use.beta 导入才算。所有变体翻译结果一致,因此若不确定,请采用 稳定版映射。翻译前请规范化旧版名称。注明您检测到的变体类型。

3. 梳理脚本

在编写任何 TypeScript 代码之前,请先提取结构化的清单:

  • 任务(Task)——即 task= 字符串;将每个任务拆分为其隐含的有序步骤。
  • 模型 — 由 Chat* 提供者 + 模型 ID。
  • 浏览器配置 — 本地与 cdp_url/Browserbase;无头模式;代理; user_data_dir/storage_state.
  • 结构化输出 — 任何 output_model_schema Pydantic 模型。
  • 密钥 — sensitive_data、环境变量使用、登录流程。
  • 防护措施 — allowed_domains, max_steps.
  • 自定义操作 — @tools.action / Controller 函数,以及每个函数是确定性的 副作用还是代理能力。
  • 设置 — initial_actions、辅助模型(page_extraction_llm, planner_llm).

4. 确定每个步骤的确定性级别

对于清单中的每个步骤,应用 determinism.md 中的决策树:

  • 导航至已知 URL → page.goto(url) 在 Stagehand 页面上(不使用 AI)。
  • 页面操作 → act("…");若需重复, observe() 执行一次后重放 act(action) (不调用LLM)。
  • 读取数据 → extract("…", zodSchema).
  • 真正开放式 → 保持 stagehand.agent().execute(...) (通过 maxSteps/systemPrompt).

当流程已知时默认采用分解;仅在 agent() 仅在未知时才进行分解。对于 首次“提升与移位”,忠实的 agent() 转换即可接受——请明确说明并注明 优化路径。

5. 生成 Stagehand v3 重写版本

首先,验证 API。在编写之前,请对照已安装的包( 类型)确认即将使用的确切签名,node_modules/@browserbasehq/stagehand 类型)或https://docs.stagehand.dev/v3中确认即将使用的确切签名。 下面的映射是3.6.x版本的快照;如果已安装版本中存在任何差异,则以已安装 版本为准。然后生成可运行的TypeScript代码。始终:

  • import { Stagehand } from "@browserbasehq/stagehand"; 并 import { z } from "zod"; 在提取时。
  • 通过 const page = stagehand.context.pages()[0];.
  • 调用实例上的 AI 方法: stagehand.act(...), stagehand.extract(...), stagehand.observe(...) — 切勿 page.act(...).
  • 将模型设置为 "provider/model" 字符串。
  • 默认值为 env: "BROWSERBASE";显示 env: "LOCAL" 作为开发选项。
  • 通过 variables 和 process.env,切勿硬编码。
  • await stagehand.init() 在开头, await stagehand.close() 在 finally.

包含项目配置以便其能够运行(参见下方的模板)。

6. 编写迁移摘要

在代码旁,编写一份简短摘要:

  • 检测到的变体以及所做的确定性选择(哪些步骤变成了确定性、AI 处理或代理处理),并附上理由。
  • 需要人工审核——任何未实现1:1映射的情况:丢失的 allowed_domains 安全防护措施、 自定义操作逻辑、次级模型意图、模棱两可的任务字符串。
  • 建议的下一步——使用 Browserbase Context 实现身份验证复用、生产环境缓存,或者 如果流程不透明,则采用 trace-assisted path。

7. 提供基于追踪的路径(仅在必要时)

如果源代码是一个庞大且不透明的 agent(task=…)、运行不稳定,或者无法确信重写后的代码能 准确映射,则提供基于跟踪的工作流(trace-assisted.md):在 Browserbase 上运行原始代码,提取 sessions.logs.list,并根据观察到的行为进行重写。未经用户许可,请勿执行任何操作。

输出模板

package.json

{
  "name": "stagehand-migration",
  "type": "module",
  "scripts": { "start": "tsx index.ts" },
  "dependencies": {
    "@browserbasehq/stagehand": "^3.0.0",
    "dotenv": "^16.0.0",
    "zod": "^3.25.0"
  },
  "devDependencies": { "tsx": "^4.0.0", "typescript": "^5.0.0" }
}

仅当 "ai": "^5.0.0" (Vercel AI SDK)仅当自定义浏览器使用操作映射到某个代理时才添加 tool时才添加。固定为 v5,而非 v4 — Stagehand 3.6.x 捆绑了 ai v5,且类型 agent({ tools }) 作为 v5 ToolSet,其中工具的模式字段为inputSchema。v4 tool() 辅助函数会输出 parameters ,且无法通过 Stagehand 的 v5 ToolSet。如果你无法 控制被提升的 ai ,请跳过 tool() 辅助函数,直接传递一个普通对象 { description, inputSchema: zodSchema, execute } ——它能满足 v5 ToolSet ,无论采用哪种 ai 主要解析结果如何。

.env

BROWSERBASE_API_KEY=...
BROWSERBASE_PROJECT_ID=...
ANTHROPIC_API_KEY=...   # or the provider matching your model string

index.ts 骨架(分解后的形式,推荐形式)

import "dotenv/config";
import { Stagehand } from "@browserbasehq/stagehand";
import { z } from "zod";

async function main() {
  const stagehand = new Stagehand({
    env: "BROWSERBASE",
    model: "anthropic/claude-sonnet-4-6",
  });
  await stagehand.init();
  try {
    const page = stagehand.context.pages()[0];

    await page.goto("https://example.com");          // deterministic skeleton
    await stagehand.act("…");                          // AI where the page varies
    const data = await stagehand.extract("…", z.object({ /* … */ }));  // structured reads

    console.log(data);
  } finally {
    await stagehand.close();
  }
}

main().catch((err) => { console.error(err); process.exit(1); });

验证清单(在宣布完成之前)

  • AI方法作用于实例(stagehand.act/extract/observe),而非页面上。
  • 通过以下方式获取页面 stagehand.context.pages()[0].
  • Model 是一个 "provider/model" 字符串;匹配的提供商键位于 .env.
  • extract 使用zod模式; zod 位于依赖项中。
  • 密钥使用 variables + process.env;没有任何硬编码内容。
  • init() / close() 存在; close() 在 finally.
  • 每个浏览器使用步骤均已考虑在内,并被刻意放置在确定性光谱上。
  • 迁移摘要列出了确定性选项以及“需要人工审核”的项目。

应避免的常见错误

  • 复制 v2 语法(page.act(), stagehand.page, modelName/modelClientOptions, enableCaching) 从旧博客文章中复制。请使用 v3 —— 参见 api-mapping 的“版本说明”。
  • 将每个步骤转换为act()——使用 page.goto 进行导航,并通过 observe→act;不要对每个操作都消耗一次 LLM 调用。
  • 将一切默认设置为agent()——这只是在 新框架中重现了浏览器使用的非确定性。在流程已知的情况下进行分解。
  • 静默舍弃allowed_domains——Stagehand没有领域防火墙;标记出来以供审查。
  • 自行发明 Browserbase/Stagehand 选项——若不确定某个字段,请查阅 https://docs.stagehand.dev/v3 / https://docs.browserbase.com,而非凭空猜测。
在 GitHub 上查看
---
name: browser-use-to-stagehand
description: Convert browser-use (Python) browser-automation scripts to Stagehand v3 (TypeScript) on Browserbase, replacing opaque agent loops with deterministic pipelines where possible.
license: MIT
---

# browser-use → Stagehand on Browserbase (`/browser-use-to-stagehand`)

Convert a browser-use (Python) script into an idiomatic **Stagehand v3 (TypeScript)** script on
**Browserbase**, choosing the right level of determinism at each step rather than producing a
one-to-one agentic copy.

**Core principle:** browser-use is agentic-by-default (the LLM decides every action). Stagehand
lets you choose how much AI to use. A good migration replaces opaque agent loops with an
inspectable, mostly-deterministic pipeline — using AI only where the page is genuinely
unpredictable. This is a refactor with judgment, not a transpile.

> **Source of truth & versions.** This skill's durable value is the *judgment* — the determinism
> spectrum and the decompose-vs-agent decision — not the API specifics, which drift every release.
> The code mappings here are a **snapshot validated against `@browserbasehq/stagehand` 3.6.x and
> browser-use 0.13.x (2026-06)**. On any conflict, the **live docs win** — always verify against the
> installed package and these sources before emitting code:
> - Stagehand v3: <https://docs.stagehand.dev/v3>  ·  installed types: `node_modules/@browserbasehq/stagehand`
> - Browserbase: <https://docs.browserbase.com>
> - browser-use: <https://docs.browser-use.com>
>
> If the installed Stagehand major is **not 3**, treat this skill as conceptual only and follow the
> live docs for every signature.

## Reference files (read as needed)

- [`references/api-mapping.md`](references/api-mapping.md) — the mechanical browser-use → Stagehand
  mapping: variant detection, the full feature table, before/after code, Browserbase platform
  options, and v3 version gotchas. **Read this for any non-trivial construct.**
- [`references/determinism.md`](references/determinism.md) — how to choose `agent()` vs
  `act`/`extract`/`observe` vs cached `observe`→`act`. The decision tree. **Read this when deciding
  how to translate an `Agent(task=…)`.**
- [`references/trace-assisted.md`](references/trace-assisted.md) — the optional "run it on
  Browserbase, read the logs, then rewrite" workflow for opaque/flaky scripts.
- [`references/guide.md`](references/guide.md) — the human migration guide: philosophy shift,
  feature mapping, the determinism spectrum, and a recommended migration path.
- [`references/prompt.md`](references/prompt.md) — a self-contained, tool-agnostic version of this
  skill; paste it into any AI assistant along with a browser-use script.
- [`EXAMPLES.md`](EXAMPLES.md) — before/after script pairs.

## Workflow

### 1. Get the source
Obtain the browser-use script(s). If the user only described a script, ask for the file(s). Note
the target: **TypeScript Stagehand on Browserbase** unless they say otherwise.

> **First, gate on scope — is this even migratable?** Not every browser-use file is an
> `Agent(task=…)` script. If the source is **browser-use running as an MCP server**
> (`uvx browser-use --mcp`, a `mcpServers` config) there is **no Stagehand equivalent** — flag it as
> out of scope, don't invent one (see api-mapping §3.7b). If the browser-use call is **embedded in a
> larger app** (a class/tool wrapper, web route, queue task), convert only the browser-use surface and
> preserve the surrounding app glue — see api-mapping §3.8.

### 2. Detect the browser-use variant
Identify legacy (pre-0.12) vs stable vs Rust beta (only when imports come from `browser_use.beta`)
— see api-mapping §1. Note: the classic top-level `from browser_use import Agent, ChatBrowserUse`
surface is alive and well in 0.13.x — `ChatBrowserUse` alone is **not** a beta tell; only a
`browser_use.beta` import is. All variants translate identically, so when unsure, proceed with the
stable mapping. Normalize legacy names before translating. State which variant you found.

### 3. Inventory the script
Extract a structured inventory before writing any TypeScript:
- **Task(s)** — the `task=` string(s); split each into its implied ordered steps.
- **Model** — the `Chat*` provider + model id.
- **Browser config** — local vs `cdp_url`/Browserbase; headless; proxies; `user_data_dir`/`storage_state`.
- **Structured output** — any `output_model_schema` Pydantic models.
- **Secrets** — `sensitive_data`, env-var usage, login flows.
- **Guardrails** — `allowed_domains`, `max_steps`.
- **Custom actions** — `@tools.action` / `Controller` functions, and whether each is a deterministic
  side-effect or an agent capability.
- **Setup** — `initial_actions`, secondary models (`page_extraction_llm`, `planner_llm`).

### 4. Decide the determinism level per step
For each step from the inventory, apply the decision tree in determinism.md:
- Navigate to a known URL → `page.goto(url)` on the Stagehand page (no AI).
- On-page action → `act("…")`; if it repeats, `observe()` once then replay `act(action)` (no LLM call).
- Reading data → `extract("…", zodSchema)`.
- Genuinely open-ended → keep `stagehand.agent().execute(...)` (tightened with `maxSteps`/`systemPrompt`).

Default to **decomposition** when the flow is known; keep `agent()` only where it isn't. For a
first lift-and-shift, a faithful `agent()` translation is acceptable — say so and note the
optimization path.

### 5. Produce the Stagehand v3 rewrite
**First, verify the API.** Before writing, confirm the exact signatures you're about to use against
the installed package (`node_modules/@browserbasehq/stagehand` types) or <https://docs.stagehand.dev/v3>.
The mappings below are a 3.6.x snapshot; if anything differs in the installed version, the installed
version wins. Then emit runnable TypeScript. Always:
- `import { Stagehand } from "@browserbasehq/stagehand";` and `import { z } from "zod";` when extracting.
- Get the page via `const page = stagehand.context.pages()[0];`.
- Call AI methods on the **instance**: `stagehand.act(...)`, `stagehand.extract(...)`,
  `stagehand.observe(...)` — **never** `page.act(...)`.
- Set the model as a `"provider/model"` string.
- Default to `env: "BROWSERBASE"`; show `env: "LOCAL"` as the dev option.
- Pass secrets via `variables` and `process.env`, never hardcoded.
- `await stagehand.init()` at the start, `await stagehand.close()` in a `finally`.

Include the project setup so it runs (see the templates below).

### 6. Write the migration summary
Alongside the code, produce a short summary:
- **Variant detected** and the determinism choices made (which steps became deterministic vs AI vs agent), with the reasoning.
- **Needs human review** — anything that didn't map 1:1: lost `allowed_domains` guardrails,
  custom-action logic, secondary-model intent, ambiguous task strings.
- **Recommended next step** — Browserbase Context for auth reuse, caching for production, or the
  trace-assisted path if the flow was opaque.

### 7. Offer the trace-assisted path (only if warranted)
If the source was one large opaque `agent(task=…)`, was flaky, or your rewrite can't be confidently
mapped, offer the trace-assisted workflow (trace-assisted.md): run the original on Browserbase, pull
`sessions.logs.list`, and rewrite from observed behavior. Don't run anything without the user's go-ahead.

## Output templates

**`package.json`**
```json
{
  "name": "stagehand-migration",
  "type": "module",
  "scripts": { "start": "tsx index.ts" },
  "dependencies": {
    "@browserbasehq/stagehand": "^3.0.0",
    "dotenv": "^16.0.0",
    "zod": "^3.25.0"
  },
  "devDependencies": { "tsx": "^4.0.0", "typescript": "^5.0.0" }
}
```
> Add `"ai": "^5.0.0"` (Vercel AI SDK) **only** if a custom browser-use action maps to an agent
> `tool`. **Pin v5, not v4** — Stagehand 3.6.x bundles `ai` v5 and types `agent({ tools })` as the v5
> `ToolSet`, where a tool's schema field is **`inputSchema`**. The v4 `tool()` helper emits
> `parameters` instead and will **fail to type-check** against Stagehand's v5 `ToolSet`. If you can't
> control the hoisted `ai` version, skip the `tool()` helper and pass a plain object
> `{ description, inputSchema: zodSchema, execute }` — it satisfies the v5 `ToolSet` regardless of which
> `ai` major resolves.

**`.env`**
```bash
BROWSERBASE_API_KEY=...
BROWSERBASE_PROJECT_ID=...
ANTHROPIC_API_KEY=...   # or the provider matching your model string
```

**`index.ts` skeleton** (decomposed, the preferred shape)
```typescript
import "dotenv/config";
import { Stagehand } from "@browserbasehq/stagehand";
import { z } from "zod";

async function main() {
  const stagehand = new Stagehand({
    env: "BROWSERBASE",
    model: "anthropic/claude-sonnet-4-6",
  });
  await stagehand.init();
  try {
    const page = stagehand.context.pages()[0];

    await page.goto("https://example.com");          // deterministic skeleton
    await stagehand.act("…");                          // AI where the page varies
    const data = await stagehand.extract("…", z.object({ /* … */ }));  // structured reads

    console.log(data);
  } finally {
    await stagehand.close();
  }
}

main().catch((err) => { console.error(err); process.exit(1); });
```

## Validation checklist (before declaring done)
- [ ] AI methods are on the **instance** (`stagehand.act/extract/observe`), not the page.
- [ ] Page obtained via `stagehand.context.pages()[0]`.
- [ ] Model is a `"provider/model"` string; the matching provider key is in `.env`.
- [ ] `extract` uses a zod schema; `zod` is in dependencies.
- [ ] Secrets use `variables` + `process.env`; nothing hardcoded.
- [ ] `init()` / `close()` present; `close()` in `finally`.
- [ ] Each browser-use step is accounted for, placed deliberately on the determinism spectrum.
- [ ] Migration summary lists determinism choices and "needs human review" items.

## Common mistakes to avoid
- **Copying v2 syntax** (`page.act()`, `stagehand.page`, `modelName`/`modelClientOptions`,
  `enableCaching`) from old blog posts. Use v3 — see api-mapping "Version notes".
- **Translating every step into `act()`** — navigate with `page.goto` and cache repeatable steps via `observe`→`act`; don't spend an LLM call on every action.
- **Defaulting everything to `agent()`** — that just reproduces browser-use's non-determinism in a
  new framework. Decompose where the flow is known.
- **Silently dropping `allowed_domains`** — Stagehand has no domain firewall; flag it for review.
- **Inventing Browserbase/Stagehand options** — if unsure of a field, check
  <https://docs.stagehand.dev/v3> / <https://docs.browserbase.com> rather than guessing.

安装 browser-use-to-stagehand

下载技能文件并将其解压到 .claude/skills/ 目录中。

下载ZIP

克隆仓库并复制技能文件到您的项目中。

git clone https://github.com/browserbase/skills/tree/main/skills/browser-use-to-stagehand # Copy SKILL.md to your .claude/skills/ directory

复制 复制
快速设置: 将技能文件夹复制到 .claude/skills/ Claude 会自动检测并使用该技能

相关技能

algorithmic-art
更新时间 2026-08-27
systematic-debugging
更新时间 2026-09-03
tech-debt-tracker
更新时间 2026-08-29
continual-learning
更新时间 2026-09-10
OR