選項

利用 browse CLI 在真實瀏覽器中執行對抗性 UI 測試,透過分析 git diff 僅測試變更區域,或全面探索整個應用程式,以找出功能、無障礙性、響應式佈局及使用者體驗方面的錯誤。

...展開全部
0
更新時間 2026-09-29

UI 測試 — Agentic UI 測試技能

在真實瀏覽器中測試 UI 變更。你的任務是設法找出問題,而非確認功能正常運作。

三種工作流程:

  • 差異驅動 — 分析 Git 差異,僅測試變更之處
  • 探索式 — 瀏覽應用程式,找出開發者未曾預見的錯誤
  • 並行測試 — 將獨立的測試組分散部署至多個 Browserbase 瀏覽器中

測試運作方式

主代理負責協調 — 它規劃測試策略、將任務委派給子代理,並整合結果。子代理則執行實際的瀏覽器測試。

規劃:多角度規劃,一次執行

您必須親自完成所有三輪規劃並輸出結果,方可啟動任何子代理。規劃作業須在您的回應中進行——絕不委派給子代理。請勿跳過規劃直接進入執行階段。

第一輪 — 功能性:核心使用者流程為何?哪些功能應能正常運作?請將每項測試以「操作 → 預期結果」的形式寫出。

第二輪 — 對抗性測試:重新閱讀第一輪內容。您遺漏了什麼?請思考:不同類型的使用者/角色、錯誤路徑、空狀態、競態條件、邊界輸入(空值、超大數值、特殊字元、快速點擊)。

第三輪 — 覆蓋缺口:重新閱讀第一至第二輪內容。請關注以下方面:無障礙功能(axe-core、僅鍵盤操作)、行動裝置視口、主控台錯誤,以及與應用程式其他部分的視覺一致性?

去重:將前三輪的測試合併為一份編號清單,移除重複項目,並將每個測試分配至不同組別(例如:A 組、B 組)。

接著執行一次 — 針對每個組別啟動一個子代理。每個子代理僅接收其專屬的測試清單以供執行,別無其他。子代理不會進行探索或規劃 — 它們僅執行指派的測試並回報結果。

在呼叫任何 Agent 工具之前,請於回覆中輸出這三輪測試、合併後的計畫以及分組分配結果。

工作拆分原則

  • 子代理執行指派的測試,而非進行開放式探索。主代理會將一份特定的編號測試清單交給每個子代理。子代理不會進行規劃、探索或決定測試內容——它們僅執行清單中的項目並停止。
  • 瓶頸取決於最慢的代理程式——應將工作拆分,確保沒有單一代理程式承擔不成比例的工作量。多個小型代理程式 > 少數大型代理程式。
  • 工作量應與變更規模相稱——單一元件的修復不需要太多代理程式或太多步驟,但整頁的重新設計則需要。讓 diff 的範圍來驅動計畫。
  • 遇到失敗時不要過早中止——在指派的測試範圍內盡可能找出更多錯誤。

為子代理設定步驟預算

主代理程式必須在每個子代理程式的提示中明確設定瀏覽步驟上限。子代理程式不會自行限制——除非另有指示,否則它們會一直運行直到完成為止。

作為粗略的經驗法則:針對少數特定檢查約需 25 個步驟;針對包含功能性、對抗性及無障礙性檢查的完整頁面約需 40 個步驟;針對多個頁面或廣泛類別則約需 75 個步驟。請根據指派測試的實際需求進行調整——這些僅為起點,而非硬性規定。

作為粗略的經驗法則:針對少數特定檢查約需 25 步;針對包含功能性、對抗性及無障礙性檢查的完整頁面約需 40 步;針對多個頁面或廣泛類別則約需 75 步。請根據指派測試的實際需求進行調整——這些僅為起點,並非硬性規定。

每個子代理的提示語必須包含:

You have a budget of N browse steps (each `browse` command = 1 step). Count your steps as you go. When you reach N, stop immediately and report:
- STEP_PASS/STEP_FAIL for every test you completed
- STEP_SKIP||budget reached for every test you didn't get to

Do not retry or continue after hitting the budget.
Run only these tests: [numbered list from the merged plan]
Do not explore beyond the assigned tests.
Do NOT generate an HTML report or write any files. Return only step markers and your findings as text.

主代理程式不應執行 browse 執行任何指令(僅限於驗證開發伺服器是否正常運作)。所有測試均在子代理程式中進行。

當子代理達到預算上限時,主代理應直接接受部分結果。請勿重新執行或重試子代理。請在最終報告中列出「已跳過」的測試項目,以便開發人員了解未涵蓋的範圍。

報告

每個子代理程式回報時應包含:

Tests: 8 | Passed: 5 | Failed: 2 | Skipped: 1 | Pages visited: 2

主代理程式會將這些內容彙整成最終報告,其中包含:

Tests: 20 | Passed: 14 | Failed: 4 | Skipped: 2 | Agents: 3 | Pass rate: 70%

請勿報告「已使用的步驟數」——瀏覽指令的計數屬於實作層面的細節,對審查者而言並非有意義的指標。

測試哲學

您是位「對抗性測試者」。您的目標是找出錯誤,而非證明正確性。

  • 試著讓每個測試的功能失效。不要只檢查「按鈕是否存在?」——試著快速點擊兩次、提交空表單、貼上 500 個字元、在流程中途按下 Escape 鍵。
  • 測試開發者未曾考慮到的情境。例如:空狀態、錯誤恢復、僅使用鍵盤導航、行動裝置的畫面溢出。
  • 每項斷言都必須有據可依。比較測試前後的快照。透過參考標識檢查特定元素。除非有來自無障礙樹或確定性檢查的具體證據,否則絕不要報告「通過」。
  • 回報失敗時,請提供足夠的細節以便重現問題。包含確切的操作步驟、預期結果、實際結果,以及建議的修正方案。

斷言規範

每個測試步驟都必須產生結構化的斷言。請勿撰寫自由形式的「這看起來沒問題。」

步驟標記

每個測試步驟必須產生且僅產生一個標記:

STEP_PASS||

或

STEP_FAIL|| → |
  • step-id:簡短識別碼,例如 homepage-cta, form-validation-error, modal-cancel
  • evidence:您觀察到的、能證明該步驟通過的資訊(元素參考、文字內容、URL、評估結果)
  • expected → actual:預期結果與實際結果的對比
  • screenshot-path:已儲存螢幕截圖的路徑(僅限失敗情況 — 請參閱下方的「螢幕截圖擷取」)

失敗情況的螢幕截圖擷取

每個 STEP_FAIL 都必須附帶一張螢幕截圖,以便開發人員能直觀地了解問題出在哪裡。

當測試步驟失敗時:

# 1. Take a screenshot immediately after observing the failure
browse screenshot --path .context/ui-test-screenshots/.png

# If --path is not supported, take the screenshot and save manually:
browse screenshot
# The browse CLI will output the screenshot path — move/copy it:
cp /tmp/browse-screenshot-*.png .context/ui-test-screenshots/.png

請在任何測試執行開始時設定截圖目錄:

mkdir -p .context/ui-test-screenshots

規則:

  • 檔案名稱 = 步驟編號(例如: double-submit.png, axe-audit.png, modal-focus-trap.png)
  • 儲存於 .context/ui-test-screenshots/ — 此目錄已加入 Git 忽略清單,且開發人員及其他代理程式均可存取
  • 若進行並行執行,請包含會話名稱: -.png (例如, signup-double-submit.png)
  • 請在發生錯誤的當下擷取螢幕截圖 — 捕捉故障狀態,而非恢復後的畫面
  • 若為視覺/版面配置錯誤,請一併擷取基準狀態(正常運作狀態)的螢幕截圖以供比較: -baseline.png

驗證方法(依嚴謹程度排序)

  1. 確定性檢查(最嚴謹)—— browse eval 會回傳可供檢視的結構化資料。範例:axe-core 違規次數、 document.title、表單欄位值、控制台錯誤陣列、元素數量。
  2. 快照元素匹配 — 輔助功能樹中存在具有特定角色與文字的特定元素。透過參考進行檢查: @0-12 button "Save"。元素在樹中要麼存在,要麼不存在。
  3. 前後比對 — 操作前的快照、執行操作、操作後的快照。驗證樹狀結構是否依預期方式變更(元素出現、消失、文字變更)。
  4. 螢幕截圖 + 視覺判斷(最不可靠)— 僅適用於無障礙樹無法擷取的純視覺屬性(顏色、間距、版面配置)。務必明確說明具體評估的項目。

前後比較模式

這是核心的驗證循環。請將其應用於每次互動:

# 1. BEFORE: capture state
browse snapshot
# Record: what elements exist, their text, their refs

# 2. ACT: perform the interaction
browse click @0-12

# 3. AFTER: capture new state
browse snapshot
# Compare: what changed? What appeared? What disappeared?

# 4. ASSERT: emit marker based on comparison
# If dialog appeared: STEP_PASS|modal-open|dialog "Confirm" appeared at @0-20
# If nothing changed:
browse screenshot --path .context/ui-test-screenshots/modal-open.png
# STEP_FAIL|modal-open|expected dialog to appear → snapshot unchanged|.context/ui-test-screenshots/modal-open.png

設定

which browse || npm install -g browse

避免權限倦怠

此技能會執行許多 browse 指令(快照、點擊、評估)。為避免必須逐一核准,請將 browse 至您的允許指令清單中:

將這兩種模式都加入 .claude/settings.json (專案層級) 或 ~/.claude/settings.json (使用者層級):

{
  "permissions": {
    "allow": [
      "Bash(browse:*)",
      "Bash(BROWSE_SESSION=*)"
    ]
  }
}

第一個模式涵蓋一般 browse 指令。第二個模式涵蓋並行工作階段(BROWSE_SESSION=signup browse open ...)。兩者皆不可或缺,以避免出現核准提示。

模式選擇

目標 模式 指令 驗證
localhost / 127.0.0.1 本地 browse open --local 無需(預設為乾淨且隔離的本地瀏覽器)
已部署/預備站 遠端 browse open --remote Browserbase 憑證;若受支援,請使用上下文

規則:若目標 URL 包含 localhost 或 127.0.0.1,則在第一個 browse open 上傳入 --local。

本地模式(localhost 的預設設定)

browse open http://localhost:3000 --local

browse open ... --local 預設會使用一個乾淨且隔離的本地瀏覽器,這對於可重現的 localhost 品質保證執行最為理想。

僅在必要時使用 local-mode 的變體:

  • browse open --auto-connect — 自動偵測現有的可除錯本地 Chrome 瀏覽器。僅當測試明確需要現有的本地登入資訊/Cookie/狀態時,才應使用此選項。
  • browse open --cdp — 連接至特定的 CDP 目標(明確的本地瀏覽器連接)。

遠端模式(透過 Cookie 同步部署的網站)

# Step 1: Sync cookies from local Chrome to Browserbase
node .claude/skills/cookie-sync/scripts/cookie-sync.mjs --domains your-app.com
# Output: Context ID: ctx_abc123

# Step 2: Open in remote mode with the synced context
SESSION_JSON="$(browse cloud sessions create --context-id ctx_abc123 --persist --keep-alive)"
SESSION_ID="$(echo "$SESSION_JSON" | jq -r .id)"
CONNECT_URL="$(echo "$SESSION_JSON" | jq -r .connectUrl)"

browse open https://staging.your-app.com --cdp "$CONNECT_URL"
browse snapshot
# ... run tests ...
browse stop
browse cloud sessions update "$SESSION_ID" --status REQUEST_RELEASE

Cookie-sync 標誌: --domains, --context, --verified, --proxy "City,ST,US"

工作流程 A:差異驅動測試

第一階段:分析差異

git diff --name-only HEAD~1          # or: git diff --name-only / git diff --name-only main...HEAD
git diff HEAD~1 --             # read actual changes

將變更的檔案分類:

檔案模式 對使用者介面的影響 測試項目
*.tsx, *.jsx, *.vue, *.svelte 元件 渲染、互動、狀態、邊界情況
pages/**, app/**, src/routes/** 路由/頁面 導航、頁面載入、內容、404 錯誤處理
*.css, *.scss, *.module.css 樣式 視覺呈現(螢幕截圖)、響應式設計
*form*, *input*, *field* 表單 驗證、提交、空欄輸入、長文字輸入、特殊字元
*modal*, *dialog*, *dropdown* 互動 展開/收合、跳出、焦點陷阱、取消與確認
*nav*, *menu*, *header* 導覽 連結、活躍狀態、路由、鍵盤導航
僅限非 UI 檔案 無 跳過 — 回報「無需 UI 測試」

第二階段:將檔案映射至 URL

偵測框架: cat package.json | grep -E '"(next|react|vue|nuxt|svelte|@sveltejs|angular|vite)"'

框架 預設埠 檔案 → URL 模式
Next.js 應用程式路由器 3000 app/dashboard/page.tsx → /dashboard
Next.js 頁面路由器 3000 pages/about.tsx → /about
Vite 5173 檢查路由器設定
Nuxt 3000 pages/index.vue → /
SvelteKit 5173 src/routes/+page.svelte → /
Angular 4200 檢查路由模組

第三階段:確保執行的是正確的程式碼

在測試前,請確認開發伺服器所提供的程式碼來自 diff —— 而非過期的分支。

若要測試拉取請求(PR)或特定分支:

# Check what branch is currently checked out
git branch --show-current

# If it's not the PR branch, switch to it
git fetch origin  && git checkout 

# Install deps — the lockfile may differ between branches
yarn install  # or npm install / pnpm install

若開發伺服器原本已在其他分支上運行,請在檢出後重新啟動它。

尋找正在運行的開發伺服器:

for port in 3000 3001 5173 4200 8080 8000 5000; do
  s=$(curl -s -o /dev/null -w "%{http_code}" "http://localhost:$port" 2>/dev/null)
  if [ "$s" != "000" ]; then echo "Dev server on port $port (HTTP $s)"; fi
done

若未找到:請告知使用者啟動其開發伺服器。

驗證其是否確實能渲染: 完成 browse open + browse snapshot,請確認可存取性樹中包含真實的頁面內容(導覽列、標題、互動元件)——而非僅有錯誤覆蓋層或空白的 main 區塊。 Next.js 開發伺服器可能會在顯示全螢幕建置錯誤對話方塊的同時,仍傳回 HTTP 200 狀態碼。若快照為空或主要內容被錯誤對話方塊佔據,表示伺服器已損壞——請在測試前先修復建置問題。

第四階段:制定測試計畫

針對每個變更區域,規劃「正常流程」與「對抗性測試」:

Test Plan (based on git diff)
=============================
Changed: src/components/SignupForm.tsx (added email validation)

1. [happy] Valid email submits successfully
   URL: http://localhost:3000/signup
   Steps: fill valid email → submit → verify success message appears

2. [adversarial] Invalid email shows error
   Steps: fill "not-an-email" → submit → verify error message appears

3. [adversarial] Empty form submission
   Steps: click submit without filling anything → verify error, no crash

4. [adversarial] XSS in email field
   Steps: fill "" → submit → verify sanitized/rejected

5. [adversarial] Rapid double-submit
   Steps: click submit twice quickly → verify no duplicate submission

6. [adversarial] Keyboard-only flow
   Steps: Tab to email → type → Tab to submit → Enter → verify success

第五階段:執行測試

browse stop 2>/dev/null
mkdir -p .context/ui-test-screenshots
# localhost/default QA → clean, reproducible local run
browse open http://localhost:3000 --local

針對每項測試,請遵循「測試前/測試後」的模式:

# Navigate
browse open http://localhost:3000/path --local
browse wait load

# BEFORE snapshot
browse snapshot
# Note the current state: elements, refs, text

# ACT
browse click @0-ref
# or: browse fill "selector" "value"
# or: browse type "text"
# or: browse press Enter

# AFTER snapshot
browse snapshot
# Compare against BEFORE: what changed?

# ASSERT with marker
# STEP_PASS|step-id|evidence  OR  STEP_FAIL|step-id|expected → actual

第 6 階段:彙報結果

## UI Test Results

### STEP_PASS|valid-email-submit|status "Thanks!" appeared at @0-42 after submit
- URL: http://localhost:3000/signup
- Before: form with email input @0-3, submit button @0-7
- Action: filled "[email protected]", clicked @0-7
- After: form replaced by status element with "Thanks! We'll be in touch."

### STEP_FAIL|double-submit|expected single submission → form submitted twice|.context/ui-test-screenshots/double-submit.png
- URL: http://localhost:3000/signup
- Before: form with submit button @0-7
- Action: clicked @0-7 twice rapidly
- After: two success toasts appeared, suggesting duplicate submission
- Screenshot: .context/ui-test-screenshots/double-submit.png
- Suggestion: disable submit button after first click, or debounce the handler

---
**Summary: 4/6 passed, 2 failed**
Failed: double-submit, xss-sanitization

Screenshots saved to `.context/ui-test-screenshots/` — open any failed step's screenshot to see the broken state.

完成後 browse stop 完成後。

第 7 階段:產生 HTML 報告

在產生文字報告後,請生成一份獨立的 HTML 報告,讓審查人員能直接在瀏覽器中開啟。該報告會將螢幕截圖內嵌其中(base64 格式),因此可作為單一檔案運作——無需任何外部依賴項。

原因:文字報告雖適用於代理程式對話,但審查者(專案經理、設計師、其他工程師)需要能開啟、瀏覽並分享的視覺化文件。內嵌的螢幕截圖能讓失敗情況一目瞭然。

如何產生

  1. 請參閱 references/report-template.html 中的 HTML 範本
  2. 透過將範本中的佔位符替換為實際測試資料來建置報告:
佔位符 值
{{TITLE}} 報告標題為 </code> tag (e.g., "UI Test: PR #1234 — OAuth Settings")</td> </tr> <tr> <td><code>{{TITLE_HTML}}</code></td> <td>Report title for the visible <code><h1></code>. If a PR URL is available, wrap the PR reference in an <code><a></code> tag so it's clickable (e.g., <code>UI Test: <a href="https://github.com/org/repo/pull/1234">PR #1234</a> — OAuth Settings</code>). If no URL, use plain text same as <code>{{TITLE}}</code>.</td> </tr> <tr> <td><code>{{META}}</code></td> <td>One-line context: date, app URL, user, branch</td> </tr> <tr> <td><code>{{TOTAL_TESTS}}</code></td> <td>Total STEP_PASS + STEP_FAIL count</td> </tr> <tr> <td><code>{{AGENT_COUNT}}</code></td> <td>Number of sub-agents that ran</td> </tr> <tr> <td><code>{{PASS_COUNT}}</code></td> <td>Number of STEP_PASS</td> </tr> <tr> <td><code>{{FAIL_COUNT}}</code></td> <td>Number of STEP_FAIL</td> </tr> <tr> <td><code>{{PASS_RATE}}</code></td> <td>Integer percentage (e.g., "92")</td> </tr> <tr> <td><code>{{RATE_CLASS}}</code></td> <td><code>good</code> (≥90%), <code>warn</code> (70–89%), <code>bad</code> (<70%)</td> </tr> <tr> <td><code>{{FAILURES_SECTION}}</code></td> <td>HTML for failed test cards (see below)</td> </tr> <tr> <td><code>{{PASSES_SECTION}}</code></td> <td>HTML for passed test cards (see below)</td> </tr> </tbody></table> <ol start="3"> <li>For each test result, generate a <code><details></code> card. Failed tests should be <strong>open by default</strong> so reviewers see them immediately:</li> </ol> <pre><code class="language-html"><!-- Failed test card (open by default) --> <div class="section"> <h2>Failures <span class="count">{{FAIL_COUNT}}</span></h2> <details class="test-card fail" open> <summary> <span class="badge fail">FAIL</span> <span class="step-id">step-id-here</span> <span class="evidence">expected → actual</span> </summary> <div class="body"> <dl> <dt>URL</dt><dd>http://localhost:3000/path</dd> <dt>Action</dt><dd>What was done</dd> <dt>Expected</dt><dd>What should have happened</dd> <dt>Actual</dt><dd>What happened instead</dd> </dl> <div class="suggestion">Fix: description of suggested fix</div> <div class="screenshot"> <img src="data:image/png;base64,..." alt="失敗的螢幕截圖"> <div class="caption">step-id.png — captured at moment of failure</div> </div> </div> </details> </div> <!-- Passed test card (collapsed by default) --> <div class="section"> <h2>Passed <span class="count">{{PASS_COUNT}}</span></h2> <details class="test-card pass"> <summary> <span class="badge pass">PASS</span> <span class="step-id">step-id-here</span> <span class="evidence">evidence summary</span> </summary> <div class="body"> <dl> <dt>URL</dt><dd>http://localhost:3000/path</dd> <dt>Evidence</dt><dd>What was observed</dd> </dl> </div> </details> </div> </code></pre> <ol start="4"> <li><strong>Embed screenshots as base64</strong> so the HTML is fully self-contained:</li> </ol> <pre><code class="language-bash"># Convert screenshot to base64 data URI base64 -i .context/<span class="notranslate">ui-test</span>-screenshots/step-id.png | tr -d '\n' # Use as: src="data:image/png;base64,<output>" </code></pre> <p>Read each screenshot file referenced in STEP_FAIL markers, base64-encode it, and embed it as an <code><img src="data:image/png;base64,..."></code> in the corresponding test card. For STEP_PASS, only embed a screenshot if one was explicitly taken (e.g., baseline screenshots).</p> <ol start="5"> <li>Write the final HTML to <code>.context/<span class="notranslate">ui-test</span>-report.html</code>:</li> </ol> <pre><code class="language-bash"># Write the generated HTML cat > .context/<span class="notranslate">ui-test</span>-report.html << 'REPORT_EOF' <!DOCTYPE html> ...generated report... REPORT_EOF # Open it for the reviewer open .context/<span class="notranslate">ui-test</span>-report.html # macOS # xdg-open .context/<span class="notranslate">ui-test</span>-report.html # Linux </code></pre> <ol start="6"> <li>Tell the user: <code>Report saved to .context/<span class="notranslate">ui-test</span>-report.html</code> and offer to open it.</li> </ol> <p><strong>Rules:</strong></p> <ul> <li>Failures section comes before passes — reviewers care about what's broken first</li> <li>Failed cards are <code>open</code> by default; passed cards are collapsed</li> <li>Every STEP_FAIL card MUST have an embedded screenshot — if the screenshot file is missing, note it in the card</li> <li>Include the suggestion/fix in each failure card if one was provided</li> <li>The report must work offline — no CDN links, no external assets</li> <li>Keep the HTML under 5MB — if screenshots push it over, reduce image quality or skip baseline screenshots for passes</li> </ul> <h2>Adversarial Test Patterns</h2> <p>Apply these to every interactive element you test. Read references/adversarial-patterns.md for the full pattern library (forms, modals, navigation, error states, keyboard accessibility).</p> <h2>Deterministic Checks</h2> <p>These produce structured data, not judgment calls. Use them as the strongest form of assertion.</p> <table> <thead> <tr> <th>Check</th> <th>What it catches</th> <th>Assertion</th> </tr> </thead> <tbody><tr> <td>axe-core</td> <td>WCAG violations</td> <td><code>violations.length === 0</code></td> </tr> <tr> <td>Console errors</td> <td>Runtime exceptions, failed requests</td> <td>empty error array</td> </tr> <tr> <td>Broken images</td> <td>Missing/failed image loads</td> <td>no images with <code>naturalWidth === 0</code></td> </tr> <tr> <td>Form labels</td> <td>Inputs without accessible labels</td> <td>every input has <code>hasLabel: true</code></td> </tr> </tbody></table> <p>For the exact <code>browse eval</code> recipes, read references/browser-recipes.md.</p> <h2>Workflow B: Exploratory Testing</h2> <p>No diff, no plan — just open the app and try to break it. Use this when the user says "test my app", "find bugs", or "QA this site."</p> <h3>Approach</h3> <ol> <li><strong>Discover the app</strong> — read <code>package.json</code> to detect the framework, then open the root URL and snapshot to see what's there</li> <li><strong>Navigate everything</strong> — click through nav links, visit every reachable page, note what exists</li> <li><strong>Test what you find</strong> — for each page, apply the adversarial patterns below (forms, modals, navigation, keyboard, error states)</li> <li><strong>Run deterministic checks</strong> — axe-core, console errors, broken images, form labels on every page</li> <li><strong>Report findings</strong> — use STEP_PASS/STEP_FAIL markers, include reproduction steps for failures</li> </ol> <p>Don't try to be systematic about coverage. Just explore like a user would, but with the intent to break things. The agent is good at this — let it roam.</p> <h3>Tips for exploratory runs</h3> <ul> <li>Start with the homepage, then follow the navigation naturally</li> <li>Try the 404 page (<code>/does-not-exist</code>) — is it custom or default?</li> <li>Look for empty states (pages with no data)</li> <li>Test forms with garbage input before valid input</li> <li>Check mobile viewport (375px) on every page — does it overflow?</li> <li>If the app has auth, use cookie-sync first</li> </ul> <h2>Workflow C: Parallel Testing</h2> <p>Run independent test groups concurrently using named <code>browse</code> sessions (<code>BROWSE_SESSION=<name></code>). Each session gets its own browser. Works with both local and remote mode.</p> <p>Use when testing multiple pages or categories and you want faster wall clock time.</p> <p>Read references/parallel-testing.md for the full workflow: session setup, agent fan-out, cookie-sync for auth, and result merging.</p> <h2>Design Consistency</h2> <p>Check whether changed UI matches the rest of the app visually. Read references/design-consistency.md when doing visual or design checks.</p> <h2>Test Categories</h2> <table> <thead> <tr> <th>Category</th> <th>How</th> <th>Assertion type</th> </tr> </thead> <tbody><tr> <td>Accessibility</td> <td>axe-core + keyboard nav</td> <td>Deterministic (violation count)</td> </tr> <tr> <td>Visual Quality</td> <td>Screenshot + heuristic evaluation</td> <td>Visual judgment (weakest — note specifics)</td> </tr> <tr> <td>Responsive</td> <td>Viewport sweep + screenshots</td> <td>Visual + deterministic (overflow check)</td> </tr> <tr> <td>Console Health</td> <td>Console capture eval</td> <td>Deterministic (error count)</td> </tr> <tr> <td>UX Heuristics</td> <td>Snapshot + Laws of UX + Nielsen's</td> <td>Structured judgment (cite specific heuristic)</td> </tr> <tr> <td>Error States</td> <td>Navigate to empty/error states</td> <td>Before/after comparison</td> </tr> <tr> <td>Data Display</td> <td>Snapshot on tables/dashboards</td> <td>Element match (column count, formatting)</td> </tr> <tr> <td>Design Consistency</td> <td>Screenshot baseline + changed page comparison</td> <td>Visual judgment (cite specific property)</td> </tr> <tr> <td>Exploratory</td> <td>Free navigation + adversarial testing</td> <td>Before/after + judgment</td> </tr> </tbody></table> <p>Reference guides (load on demand):</p> <ul> <li><strong>Adversarial patterns</strong> — references/adversarial-patterns.md — load when testing forms, modals, navigation, or keyboard a11y</li> <li><strong>Browser recipes</strong> — references/browser-recipes.md — load when running deterministic checks (axe-core, console, images, form labels)</li> <li><strong>Exploratory testing</strong> — references/exploratory-testing.md — load for Workflow B (no diff, open exploration)</li> <li><strong>UX heuristics</strong> — references/ux-heuristics.md — load when evaluating UX quality or citing specific heuristics</li> <li><strong>Design system</strong> — references/design-system.example.md — template for users to customize</li> <li><strong>Design consistency</strong> — references/design-consistency.md — load when doing visual consistency checks</li> <li><strong>Parallel testing</strong> — references/parallel-testing.md — load for Workflow C (concurrent sessions)</li> <li><strong>Report template</strong> — references/report-template.html — HTML template for Phase 7 report generation</li> </ul> <p>For worked examples with exact commands, read EXAMPLES.md if you need to see the assertion protocol in action.</p> <h2>Best Practices</h2> <ol> <li><strong>Be adversarial</strong> — try to break things, don't just confirm they work</li> <li><strong>Every assertion needs evidence</strong> — snapshot ref, eval result, or before/after diff</li> <li><strong>Before/after for every interaction</strong> — snapshot, act, snapshot, compare</li> <li><strong>Screenshot every failure</strong> — <code>browse screenshot</code> immediately on STEP_FAIL, save to <code>.context/<span class="notranslate">ui-test</span>-screenshots/<step-id>.png</code></li> <li><strong>Deterministic checks first</strong> — axe-core, console errors, form labels before visual judgment</li> <li><strong>For localhost, start with clean local mode</strong> — pass <code>--local</code> on the first <code>browse open</code> for reproducible runs; use <code>--auto-connect</code> only when existing local state is required</li> <li><strong>Always <code>browse stop</code> when done</strong> — for parallel runs, stop every named session</li> <li><strong>Report failures with reproduction steps</strong> — action, expected, actual, screenshot path, suggestion</li> <li><strong>Parallelize independent tests</strong> — use Workflow C with named sessions when testing multiple pages or categories on a deployed site</li> </ol> <h2>Troubleshooting</h2> <ul> <li><strong>"No active page"</strong>: <code>browse stop</code>, retry. For zombies: <code>pkill -f "browse.*daemon"</code></li> <li><strong>Dev server not responding</strong>: <code>curl http://localhost:<port></code> — ask user to start it</li> <li><strong><code>browse eval</code> with <code>await</code> fails</strong>: Use <code>.then()</code> instead — <code>browse eval</code> doesn't support top-level await</li> <li><strong>Element ref not found</strong>: <code>browse snapshot</code> again — refs change on page update</li> <li><strong>Blank snapshot</strong>: <code>browse wait load</code> or <code>browse wait selector ".expected"</code> before snapshotting</li> <li><strong>SPA deep links 404</strong>: Navigate to <code>/</code> first, then click through</li> <li><strong>Remote auth fails</strong>: Re-run cookie-sync with <code>--context <id></code>, try <code>--verified</code></li> <li><strong>Parallel session conflicts</strong>: Ensure every <code>browse</code> command uses <code>BROWSE_SESSION=<name></code> — without it, commands go to the default session</li> <li><strong>Session not stopping</strong>: <code>BROWSE_SESSION=<name> browse stop</code>. For zombies: <code>pkill -f "browse.*<name>.*daemon"</code></li> </ul>
在 GitHub 上查看
---
name: ui-test
description: Runs adversarial UI tests in a real browser using the browse CLI, analyzing git diffs to test only changed areas or exploring the full app to find bugs in functionality, accessibility, responsive layout, and UX.
license: MIT
---

# UI Test — Agentic UI Testing Skill

Test UI changes in a real browser. Your job is to **try to break things**, not confirm they work.

Three workflows:
- **Diff-driven** — analyze a git diff, test only what changed
- **Exploratory** — navigate the app, find bugs the developer didn't think about
- **Parallel** — fan out independent test groups across multiple Browserbase browsers

## How Testing Works

The main agent **coordinates** — it plans test strategy, delegates to sub-agents, and merges results. Sub-agents do the actual browser testing.

### Planning: multiple angles, then execute once

**You MUST complete all three planning rounds yourself and output them before launching any sub-agents.** Planning happens in your own response — it is NOT delegated to sub-agents. Do not skip ahead to execution.

**Round 1 — Functional:** What are the core user flows? What should work? Write out each test as: action → expected result.

**Round 2 — Adversarial:** Re-read Round 1. What did you miss? Think about: different user types/roles, error paths, empty states, race conditions, edge inputs (empty, huge, special chars, rapid clicks).

**Round 3 — Coverage gaps:** Re-read Rounds 1–2. What about: accessibility (axe-core, keyboard-only), mobile viewports, console errors, visual consistency with the rest of the app?

**Deduplicate:** Merge all three rounds into one numbered list of tests. Remove overlaps. Assign each test to a group (e.g. Group A, Group B).

**Then execute once** — launch one sub-agent per group. Each sub-agent receives its specific list of tests to run, nothing more. Sub-agents do not explore or plan — they execute assigned tests and report results.

Output the three rounds, the merged plan, and the group assignments in your response before calling any Agent tool.

### Principles for splitting work

- **Sub-agents run assigned tests, not open exploration.** The main agent hands each sub-agent a specific numbered list of tests. Sub-agents do not plan, explore, or decide what to test — they execute the list and stop.
- **The bottleneck is the slowest agent** — split work so no single agent has a disproportionate share. Many small agents > few large ones.
- **Size the effort to the change** — a single component fix doesn't need many agents or many steps. A full-page redesign does. Let the scope of the diff drive the plan.
- **No early stopping on failures** — find as many bugs as possible within the assigned tests.

### Giving sub-agents a step budget

**The main agent MUST include an explicit browse step limit in every sub-agent prompt.** Sub-agents do not self-limit — they will run until done unless told otherwise.

As a rough heuristic: ~25 steps for a few targeted checks, ~40 for a full page with functional + adversarial + a11y, ~75 for multiple pages or a broad category. **Adjust based on what the assigned tests actually require** — these are starting points, not rules.

As a rough heuristic: ~25 steps for a few targeted checks, ~40 for a full page with functional + adversarial + a11y, ~75 for multiple pages or a broad category. **Adjust based on what the assigned tests actually require** — these are starting points, not rules.

Every sub-agent prompt must include:
```
You have a budget of N browse steps (each `browse` command = 1 step). Count your steps as you go. When you reach N, stop immediately and report:
- STEP_PASS/STEP_FAIL for every test you completed
- STEP_SKIP|<test-id>|budget reached for every test you didn't get to

Do not retry or continue after hitting the budget.
Run only these tests: [numbered list from the merged plan]
Do not explore beyond the assigned tests.
Do NOT generate an HTML report or write any files. Return only step markers and your findings as text.
```

The main agent should NOT run `browse` commands itself (except to verify the dev server is up). All testing happens in sub-agents.

**When a sub-agent hits its budget, the main agent accepts the partial results as-is.** Do not re-run or retry the sub-agent. Include SKIPPED tests in the final report so the developer knows what wasn't covered.

### Reporting

**Every sub-agent reports back with:**
```
Tests: 8 | Passed: 5 | Failed: 2 | Skipped: 1 | Pages visited: 2
```

**The main agent merges into a final report with:**
```
Tests: 20 | Passed: 14 | Failed: 4 | Skipped: 2 | Agents: 3 | Pass rate: 70%
```

Do not report "steps used" — browse command counts are implementation plumbing, not a meaningful metric for reviewers.

## Testing Philosophy

**You are an adversarial tester.** Your goal is to find bugs, not prove correctness.

- **Try to break every feature you test.** Don't just check "does the button exist?" — click it twice rapidly, submit empty forms, paste 500 characters, press Escape mid-flow.
- **Test what the developer didn't think about.** Empty states, error recovery, keyboard-only navigation, mobile overflow.
- **Every assertion must be evidence-based.** Compare before/after snapshots. Check specific elements by ref. Never report PASS without concrete evidence from the accessibility tree or a deterministic check.
- **Report failures with enough detail to reproduce.** Include the exact action, what you expected, what you got, and a suggested fix.

## Assertion Protocol

Every test step MUST produce a structured assertion. Do not write freeform "this looks good."

### Step markers

For each test step, emit exactly one marker:

```
STEP_PASS|<step-id>|<evidence>
```
or
```
STEP_FAIL|<step-id>|<expected> → <actual>|<screenshot-path>
```

- `step-id`: short identifier like `homepage-cta`, `form-validation-error`, `modal-cancel`
- `evidence`: what you observed that proves the step passed (element ref, text content, URL, eval result)
- `expected → actual`: what you expected vs what you got
- `screenshot-path`: path to the saved screenshot (failures only — see Screenshot Capture below)

### Screenshot Capture for Failures

**Every STEP_FAIL MUST have an accompanying screenshot** so the developer can see what went wrong visually.

When a test step fails:

```bash
# 1. Take a screenshot immediately after observing the failure
browse screenshot --path .context/ui-test-screenshots/<step-id>.png

# If --path is not supported, take the screenshot and save manually:
browse screenshot
# The browse CLI will output the screenshot path — move/copy it:
cp /tmp/browse-screenshot-*.png .context/ui-test-screenshots/<step-id>.png
```

Setup the screenshot directory at the start of any test run:

```bash
mkdir -p .context/ui-test-screenshots
```

**Rules:**
- File name = step-id (e.g., `double-submit.png`, `axe-audit.png`, `modal-focus-trap.png`)
- Store in `.context/ui-test-screenshots/` — this directory is gitignored and accessible to the developer and other agents
- For parallel runs, include the session name: `<session>-<step-id>.png` (e.g., `signup-double-submit.png`)
- Take the screenshot at the moment of failure — capture the broken state, not after recovery
- For visual/layout bugs, also screenshot the baseline (working state) for comparison: `<step-id>-baseline.png`

### How to verify (in order of rigor)

1. **Deterministic check** (strongest) — `browse eval` returns structured data you can inspect. Examples: axe-core violation count, `document.title`, form field value, console error array, element count.
2. **Snapshot element match** — a specific element with a specific role and text exists in the accessibility tree. Check by ref: `@0-12 button "Save"`. An element either exists in the tree or it doesn't.
3. **Before/after comparison** — snapshot before action, act, snapshot after. Verify the tree changed in the expected way (element appeared, disappeared, text changed).
4. **Screenshot + visual judgment** (weakest) — only for visual-only properties (color, spacing, layout) that the accessibility tree cannot capture. Always accompany with what specifically you're evaluating.

### Before/after comparison pattern

This is the core verification loop. Use it for every interaction:

```bash
# 1. BEFORE: capture state
browse snapshot
# Record: what elements exist, their text, their refs

# 2. ACT: perform the interaction
browse click @0-12

# 3. AFTER: capture new state
browse snapshot
# Compare: what changed? What appeared? What disappeared?

# 4. ASSERT: emit marker based on comparison
# If dialog appeared: STEP_PASS|modal-open|dialog "Confirm" appeared at @0-20
# If nothing changed:
browse screenshot --path .context/ui-test-screenshots/modal-open.png
# STEP_FAIL|modal-open|expected dialog to appear → snapshot unchanged|.context/ui-test-screenshots/modal-open.png
```

## Setup

```bash
which browse || npm install -g browse
```

### Avoid permission fatigue

This skill runs many `browse` commands (snapshots, clicks, evals). To avoid approving each one, add `browse` to your allowed commands:

Add both patterns to `.claude/settings.json` (project-level) or `~/.claude/settings.json` (user-level):
```json
{
  "permissions": {
    "allow": [
      "Bash(browse:*)",
      "Bash(BROWSE_SESSION=*)"
    ]
  }
}
```

The first pattern covers plain `browse` commands. The second covers parallel sessions (`BROWSE_SESSION=signup browse open ...`). Both are needed to avoid approval prompts.

## Mode Selection

| Target | Mode | Command | Auth |
|--------|------|---------|------|
| `localhost` / `127.0.0.1` | Local | `browse open <url> --local` | None needed (clean isolated local browser by default) |
| Deployed/staging site | Remote | `browse open <url> --remote` | Browserbase credentials; use contexts where supported |

**Rule: If the target URL contains `localhost` or `127.0.0.1`, pass `--local` on the first `browse open`.**

### Local Mode (default for localhost)

```bash
browse open http://localhost:3000 --local
```

`browse open ... --local` uses a clean isolated local browser by default, which is best for reproducible localhost QA runs.

Use local-mode variants only when needed:

- `browse open <url> --auto-connect` — auto-discover an existing debuggable local Chrome. Use this only when the test explicitly needs existing local login/cookies/state.
- `browse open <url> --cdp <port|url>` — attach to a specific CDP target (explicit local browser attach).

### Remote Mode (deployed sites via cookie-sync)

```bash
# Step 1: Sync cookies from local Chrome to Browserbase
node .claude/skills/cookie-sync/scripts/cookie-sync.mjs --domains your-app.com
# Output: Context ID: ctx_abc123

# Step 2: Open in remote mode with the synced context
SESSION_JSON="$(browse cloud sessions create --context-id ctx_abc123 --persist --keep-alive)"
SESSION_ID="$(echo "$SESSION_JSON" | jq -r .id)"
CONNECT_URL="$(echo "$SESSION_JSON" | jq -r .connectUrl)"

browse open https://staging.your-app.com --cdp "$CONNECT_URL"
browse snapshot
# ... run tests ...
browse stop
browse cloud sessions update "$SESSION_ID" --status REQUEST_RELEASE
```

Cookie-sync flags: `--domains`, `--context`, `--verified`, `--proxy "City,ST,US"`

## Workflow A: Diff-Driven Testing

### Phase 1: Analyze the diff

```bash
git diff --name-only HEAD~1          # or: git diff --name-only / git diff --name-only main...HEAD
git diff HEAD~1 -- <file>            # read actual changes
```

Categorize changed files:

| File pattern | UI impact | What to test |
|-------------|-----------|--------------|
| `*.tsx`, `*.jsx`, `*.vue`, `*.svelte` | Component | Render, interaction, state, edge cases |
| `pages/**`, `app/**`, `src/routes/**` | Route/page | Navigation, page load, content, 404 handling |
| `*.css`, `*.scss`, `*.module.css` | Style | Visual appearance (screenshot), responsive |
| `*form*`, `*input*`, `*field*` | Form | Validation, submission, empty input, long input, special chars |
| `*modal*`, `*dialog*`, `*dropdown*` | Interactive | Open/close, escape, focus trap, cancel vs confirm |
| `*nav*`, `*menu*`, `*header*` | Navigation | Links, active states, routing, keyboard nav |
| Non-UI files only | None | Skip — report "no UI tests needed" |

### Phase 2: Map files to URLs

Detect framework: `cat package.json | grep -E '"(next|react|vue|nuxt|svelte|@sveltejs|angular|vite)"'`

| Framework | Default port | File → URL pattern |
|-----------|-------------|-----|
| Next.js App Router | 3000 | `app/dashboard/page.tsx` → `/dashboard` |
| Next.js Pages Router | 3000 | `pages/about.tsx` → `/about` |
| Vite | 5173 | Check router config |
| Nuxt | 3000 | `pages/index.vue` → `/` |
| SvelteKit | 5173 | `src/routes/+page.svelte` → `/` |
| Angular | 4200 | Check routing module |

### Phase 3: Ensure the right code is running

Before testing, verify the dev server is serving the code from the diff — not a stale branch.

**If testing a PR or specific branch:**
```bash
# Check what branch is currently checked out
git branch --show-current

# If it's not the PR branch, switch to it
git fetch origin <branch> && git checkout <branch>

# Install deps — the lockfile may differ between branches
yarn install  # or npm install / pnpm install
```

If the dev server was already running on a different branch, restart it after checkout.

**Find a running dev server:**
```bash
for port in 3000 3001 5173 4200 8080 8000 5000; do
  s=$(curl -s -o /dev/null -w "%{http_code}" "http://localhost:$port" 2>/dev/null)
  if [ "$s" != "000" ]; then echo "Dev server on port $port (HTTP $s)"; fi
done
```

If nothing found: tell the user to start their dev server.

**Verify it actually renders:**
After `browse open` + `browse snapshot`, check that the accessibility tree contains real page content (navigation, headings, interactive elements) — not just an error overlay or empty body. Next.js dev servers can return HTTP 200 while showing a full-screen build error dialog. If the snapshot is empty or dominated by an error dialog, the server is broken — fix the build before testing.

### Phase 4: Generate test plan

For each changed area, plan **both happy path AND adversarial tests**:

```
Test Plan (based on git diff)
=============================
Changed: src/components/SignupForm.tsx (added email validation)

1. [happy] Valid email submits successfully
   URL: http://localhost:3000/signup
   Steps: fill valid email → submit → verify success message appears

2. [adversarial] Invalid email shows error
   Steps: fill "not-an-email" → submit → verify error message appears

3. [adversarial] Empty form submission
   Steps: click submit without filling anything → verify error, no crash

4. [adversarial] XSS in email field
   Steps: fill "<script>alert(1)</script>" → submit → verify sanitized/rejected

5. [adversarial] Rapid double-submit
   Steps: click submit twice quickly → verify no duplicate submission

6. [adversarial] Keyboard-only flow
   Steps: Tab to email → type → Tab to submit → Enter → verify success
```

### Phase 5: Execute tests

```bash
browse stop 2>/dev/null
mkdir -p .context/ui-test-screenshots
# localhost/default QA → clean, reproducible local run
browse open http://localhost:3000 --local
```

For each test, follow the **before/after pattern**:

```bash
# Navigate
browse open http://localhost:3000/path --local
browse wait load

# BEFORE snapshot
browse snapshot
# Note the current state: elements, refs, text

# ACT
browse click @0-ref
# or: browse fill "selector" "value"
# or: browse type "text"
# or: browse press Enter

# AFTER snapshot
browse snapshot
# Compare against BEFORE: what changed?

# ASSERT with marker
# STEP_PASS|step-id|evidence  OR  STEP_FAIL|step-id|expected → actual
```

### Phase 6: Report results

```
## UI Test Results

### STEP_PASS|valid-email-submit|status "Thanks!" appeared at @0-42 after submit
- URL: http://localhost:3000/signup
- Before: form with email input @0-3, submit button @0-7
- Action: filled "[email protected]", clicked @0-7
- After: form replaced by status element with "Thanks! We'll be in touch."

### STEP_FAIL|double-submit|expected single submission → form submitted twice|.context/ui-test-screenshots/double-submit.png
- URL: http://localhost:3000/signup
- Before: form with submit button @0-7
- Action: clicked @0-7 twice rapidly
- After: two success toasts appeared, suggesting duplicate submission
- Screenshot: .context/ui-test-screenshots/double-submit.png
- Suggestion: disable submit button after first click, or debounce the handler

---
**Summary: 4/6 passed, 2 failed**
Failed: double-submit, xss-sanitization

Screenshots saved to `.context/ui-test-screenshots/` — open any failed step's screenshot to see the broken state.
```

Always `browse stop` when done.

### Phase 7: Generate HTML report

After producing the text report, generate a standalone HTML report that a reviewer can open in a browser. The report embeds screenshots inline (base64) so it works as a single file — no external dependencies.

**Why:** Text reports are good for the agent conversation, but reviewers (PMs, designers, other engineers) want a visual artifact they can open, scan, and share. Screenshots inline make failures immediately obvious.

#### How to generate

1. Read the HTML template at [references/report-template.html](references/report-template.html)
2. Build the report by replacing the template placeholders with actual test data:

| Placeholder | Value |
|-------------|-------|
| `{{TITLE}}` | Report title for `<title>` tag (e.g., "UI Test: PR #1234 — OAuth Settings") |
| `{{TITLE_HTML}}` | Report title for the visible `<h1>`. If a PR URL is available, wrap the PR reference in an `<a>` tag so it's clickable (e.g., `UI Test: <a href="https://github.com/org/repo/pull/1234">PR #1234</a> — OAuth Settings`). If no URL, use plain text same as `{{TITLE}}`. |
| `{{META}}` | One-line context: date, app URL, user, branch |
| `{{TOTAL_TESTS}}` | Total STEP_PASS + STEP_FAIL count |
| `{{AGENT_COUNT}}` | Number of sub-agents that ran |
| `{{PASS_COUNT}}` | Number of STEP_PASS |
| `{{FAIL_COUNT}}` | Number of STEP_FAIL |
| `{{PASS_RATE}}` | Integer percentage (e.g., "92") |
| `{{RATE_CLASS}}` | `good` (≥90%), `warn` (70–89%), `bad` (<70%) |
| `{{FAILURES_SECTION}}` | HTML for failed test cards (see below) |
| `{{PASSES_SECTION}}` | HTML for passed test cards (see below) |

3. For each test result, generate a `<details>` card. Failed tests should be **open by default** so reviewers see them immediately:

```html
<!-- Failed test card (open by default) -->
<div class="section">
  <h2>Failures <span class="count">{{FAIL_COUNT}}</span></h2>
  <details class="test-card fail" open>
    <summary>
      <span class="badge fail">FAIL</span>
      <span class="step-id">step-id-here</span>
      <span class="evidence">expected → actual</span>
    </summary>
    <div class="body">
      <dl>
        <dt>URL</dt><dd>http://localhost:3000/path</dd>
        <dt>Action</dt><dd>What was done</dd>
        <dt>Expected</dt><dd>What should have happened</dd>
        <dt>Actual</dt><dd>What happened instead</dd>
      </dl>
      <div class="suggestion">Fix: description of suggested fix</div>
      <div class="screenshot">
        <img src="data:image/png;base64,..." alt="Screenshot of failure">
        <div class="caption">step-id.png — captured at moment of failure</div>
      </div>
    </div>
  </details>
</div>

<!-- Passed test card (collapsed by default) -->
<div class="section">
  <h2>Passed <span class="count">{{PASS_COUNT}}</span></h2>
  <details class="test-card pass">
    <summary>
      <span class="badge pass">PASS</span>
      <span class="step-id">step-id-here</span>
      <span class="evidence">evidence summary</span>
    </summary>
    <div class="body">
      <dl>
        <dt>URL</dt><dd>http://localhost:3000/path</dd>
        <dt>Evidence</dt><dd>What was observed</dd>
      </dl>
    </div>
  </details>
</div>
```

4. **Embed screenshots as base64** so the HTML is fully self-contained:

```bash
# Convert screenshot to base64 data URI
base64 -i .context/ui-test-screenshots/step-id.png | tr -d '\n'
# Use as: src="data:image/png;base64,<output>"
```

Read each screenshot file referenced in STEP_FAIL markers, base64-encode it, and embed it as an `<img src="data:image/png;base64,...">` in the corresponding test card. For STEP_PASS, only embed a screenshot if one was explicitly taken (e.g., baseline screenshots).

5. Write the final HTML to `.context/ui-test-report.html`:

```bash
# Write the generated HTML
cat > .context/ui-test-report.html << 'REPORT_EOF'
<!DOCTYPE html>
...generated report...
REPORT_EOF

# Open it for the reviewer
open .context/ui-test-report.html  # macOS
# xdg-open .context/ui-test-report.html  # Linux
```

6. Tell the user: `Report saved to .context/ui-test-report.html` and offer to open it.

**Rules:**
- Failures section comes before passes — reviewers care about what's broken first
- Failed cards are `open` by default; passed cards are collapsed
- Every STEP_FAIL card MUST have an embedded screenshot — if the screenshot file is missing, note it in the card
- Include the suggestion/fix in each failure card if one was provided
- The report must work offline — no CDN links, no external assets
- Keep the HTML under 5MB — if screenshots push it over, reduce image quality or skip baseline screenshots for passes

## Adversarial Test Patterns

Apply these to every interactive element you test. Read [references/adversarial-patterns.md](references/adversarial-patterns.md) for the full pattern library (forms, modals, navigation, error states, keyboard accessibility).

## Deterministic Checks

These produce structured data, not judgment calls. Use them as the strongest form of assertion.

| Check | What it catches | Assertion |
|-------|----------------|-----------|
| axe-core | WCAG violations | `violations.length === 0` |
| Console errors | Runtime exceptions, failed requests | empty error array |
| Broken images | Missing/failed image loads | no images with `naturalWidth === 0` |
| Form labels | Inputs without accessible labels | every input has `hasLabel: true` |

For the exact `browse eval` recipes, read [references/browser-recipes.md](references/browser-recipes.md).

## Workflow B: Exploratory Testing

No diff, no plan — just open the app and try to break it. Use this when the user says "test my app", "find bugs", or "QA this site."

### Approach

1. **Discover the app** — read `package.json` to detect the framework, then open the root URL and snapshot to see what's there
2. **Navigate everything** — click through nav links, visit every reachable page, note what exists
3. **Test what you find** — for each page, apply the adversarial patterns below (forms, modals, navigation, keyboard, error states)
4. **Run deterministic checks** — axe-core, console errors, broken images, form labels on every page
5. **Report findings** — use STEP_PASS/STEP_FAIL markers, include reproduction steps for failures

Don't try to be systematic about coverage. Just explore like a user would, but with the intent to break things. The agent is good at this — let it roam.

### Tips for exploratory runs

- Start with the homepage, then follow the navigation naturally
- Try the 404 page (`/does-not-exist`) — is it custom or default?
- Look for empty states (pages with no data)
- Test forms with garbage input before valid input
- Check mobile viewport (375px) on every page — does it overflow?
- If the app has auth, use cookie-sync first

## Workflow C: Parallel Testing

Run independent test groups concurrently using named `browse` sessions (`BROWSE_SESSION=<name>`). Each session gets its own browser. Works with both local and remote mode.

Use when testing multiple pages or categories and you want faster wall clock time.

Read [references/parallel-testing.md](references/parallel-testing.md) for the full workflow: session setup, agent fan-out, cookie-sync for auth, and result merging.

## Design Consistency

Check whether changed UI matches the rest of the app visually. Read [references/design-consistency.md](references/design-consistency.md) when doing visual or design checks.

## Test Categories

| Category | How | Assertion type |
|----------|-----|---------------|
| Accessibility | axe-core + keyboard nav | Deterministic (violation count) |
| Visual Quality | Screenshot + heuristic evaluation | Visual judgment (weakest — note specifics) |
| Responsive | Viewport sweep + screenshots | Visual + deterministic (overflow check) |
| Console Health | Console capture eval | Deterministic (error count) |
| UX Heuristics | Snapshot + Laws of UX + Nielsen's | Structured judgment (cite specific heuristic) |
| Error States | Navigate to empty/error states | Before/after comparison |
| Data Display | Snapshot on tables/dashboards | Element match (column count, formatting) |
| Design Consistency | Screenshot baseline + changed page comparison | Visual judgment (cite specific property) |
| Exploratory | Free navigation + adversarial testing | Before/after + judgment |

Reference guides (load on demand):
- **Adversarial patterns** — [references/adversarial-patterns.md](references/adversarial-patterns.md) — load when testing forms, modals, navigation, or keyboard a11y
- **Browser recipes** — [references/browser-recipes.md](references/browser-recipes.md) — load when running deterministic checks (axe-core, console, images, form labels)
- **Exploratory testing** — [references/exploratory-testing.md](references/exploratory-testing.md) — load for Workflow B (no diff, open exploration)
- **UX heuristics** — [references/ux-heuristics.md](references/ux-heuristics.md) — load when evaluating UX quality or citing specific heuristics
- **Design system** — [references/design-system.example.md](references/design-system.example.md) — template for users to customize
- **Design consistency** — [references/design-consistency.md](references/design-consistency.md) — load when doing visual consistency checks
- **Parallel testing** — [references/parallel-testing.md](references/parallel-testing.md) — load for Workflow C (concurrent sessions)
- **Report template** — [references/report-template.html](references/report-template.html) — HTML template for Phase 7 report generation

For worked examples with exact commands, read [EXAMPLES.md](EXAMPLES.md) if you need to see the assertion protocol in action.

## Best Practices

1. **Be adversarial** — try to break things, don't just confirm they work
2. **Every assertion needs evidence** — snapshot ref, eval result, or before/after diff
3. **Before/after for every interaction** — snapshot, act, snapshot, compare
4. **Screenshot every failure** — `browse screenshot` immediately on STEP_FAIL, save to `.context/ui-test-screenshots/<step-id>.png`
5. **Deterministic checks first** — axe-core, console errors, form labels before visual judgment
6. **For localhost, start with clean local mode** — pass `--local` on the first `browse open` for reproducible runs; use `--auto-connect` only when existing local state is required
7. **Always `browse stop` when done** — for parallel runs, stop every named session
8. **Report failures with reproduction steps** — action, expected, actual, screenshot path, suggestion
9. **Parallelize independent tests** — use Workflow C with named sessions when testing multiple pages or categories on a deployed site

## Troubleshooting

- **"No active page"**: `browse stop`, retry. For zombies: `pkill -f "browse.*daemon"`
- **Dev server not responding**: `curl http://localhost:<port>` — ask user to start it
- **`browse eval` with `await` fails**: Use `.then()` instead — `browse eval` doesn't support top-level await
- **Element ref not found**: `browse snapshot` again — refs change on page update
- **Blank snapshot**: `browse wait load` or `browse wait selector ".expected"` before snapshotting
- **SPA deep links 404**: Navigate to `/` first, then click through
- **Remote auth fails**: Re-run cookie-sync with `--context <id>`, try `--verified`
- **Parallel session conflicts**: Ensure every `browse` command uses `BROWSE_SESSION=<name>` — without it, commands go to the default session
- **Session not stopping**: `BROWSE_SESSION=<name> browse stop`. For zombies: `pkill -f "browse.*<name>.*daemon"`

安裝 ui-test

請下載並將技能檔案解壓縮至您的 .claude/skills/ 目錄中。

下載 ZIP

複製儲存庫並將技能檔案複製到您的專案中。

git clone https://github.com/browserbase/skills/tree/main/skills/ui-test # Copy SKILL.md to your .claude/skills/ directory

複製 複製
快速設定: 將技能資料夾複製到 .claude/skills/ Claude 會自動偵測並使用該技能
儲存庫 browserbase/skills

相關技能

github-code-search
更新時間 2026-06-29
drizzle-orm
更新時間 2026-06-29
prisma-client-api
更新時間 2026-06-29
clickhouse-io
更新時間 2026-06-29
OR