ui-test
browserbase/skills
browse CLI를 사용하여 실제 브라우저에서 대항적 UI 테스트를 실행하며, git diff를 분석해 변경된 부분만 테스트하거나 앱 전체를 탐색하여 기능, 접근성, 반응형 레이아웃 및 UX에서 발생하는 버그를 찾아냅니다.
...모든 것을 확장하십시오UI 테스트 — 에이전틱 UI 테스트 기술
실제 브라우저에서 UI 변경 사항을 테스트하세요. 여러분의 임무는 기능이 제대로 작동하는지 확인하는 것이 아니라, 오류가 발생하도록 시도하는 것입니다.
세 가지 워크플로:
- 차이점 기반 — git diff를 분석하여 변경된 부분만 테스트
- 탐색적 테스트 — 앱을 탐색하며 개발자가 미처 생각하지 못한 버그를 찾아냅니다
- 병렬 — 여러 Browserbase 브라우저에 독립적인 테스트 그룹을 분산시켜 실행
테스트 작동 방식
메인 에이전트가 조율합니다 — 테스트 전략을 수립하고, 하위 에이전트에 작업을 위임하며, 결과를 통합합니다. 하위 에이전트가 실제 브라우저 테스트를 수행합니다.
계획: 다양한 각도에서 검토한 후 한 번에 실행
하위 에이전트를 실행하기 전에 반드시 세 가지 계획 단계를 모두 직접 완료하고 결과를 출력해야 합니다. 계획 수립은 본인의 응답 내에서 이루어지며, 하위 에이전트에게 위임되지 않습니다. 실행 단계로 건너뛰지 마십시오.
1차 — 기능적: 핵심 사용자 흐름은 무엇인가? 무엇이 정상적으로 작동해야 하는가? 각 테스트를 ‘작업 → 예상 결과’ 형식으로 작성하십시오.
2차 — 적대적 테스트: 1차 내용을 다시 읽어보세요. 놓친 부분은 무엇인가요? 다음 사항을 고려하세요: 다양한 사용자 유형/역할, 오류 경로, 빈 상태, 경합 조건, 극한 입력(빈 값, 과도한 입력, 특수 문자, 빠른 클릭).
3차 — 커버리지 누락: 1~2차를 다시 읽어보세요. 접근성(axe-core, 키보드 전용), 모바일 뷰포트, 콘솔 오류, 앱의 나머지 부분과의 시각적 일관성 등은 어떻게 되나요?
중복 제거: 세 라운드의 테스트를 모두 하나의 번호가 매겨진 테스트 목록으로 통합하세요. 중복되는 항목을 제거하세요. 각 테스트를 그룹(예: 그룹 A, 그룹 B)에 할당하세요.
그런 다음 한 번 실행하세요 — 그룹당 하나의 하위 에이전트를 실행합니다. 각 하위 에이전트는 실행해야 할 특정 테스트 목록만 수신하며, 그 외에는 아무것도 수행하지 않습니다. 하위 에이전트는 탐색하거나 계획을 세우지 않으며, 할당된 테스트를 실행하고 결과를 보고합니다.
에이전트 도구를 호출하기 전에 응답에 세 번의 실행 라운드, 병합된 계획, 그룹 할당 내용을 출력하십시오.
작업 분할 원칙
- 서브 에이전트는 자유롭게 탐색하지 않고 할당된 테스트를 실행합니다. 메인 에이전트는 각 서브 에이전트에게 특정 번호가 매겨진 테스트 목록을 전달합니다. 서브 에이전트는 계획을 세우거나 탐색하거나 테스트 대상을 결정하지 않으며, 단순히 목록을 실행한 후 중단합니다.
- 병목 현상은 가장 느린 에이전트에서 발생합니다. 따라서 단일 에이전트가 불균형하게 많은 작업을 떠안지 않도록 작업을 분할하십시오. 소수의 대형 에이전트보다 다수의 소형 에이전트가 더 낫습니다.
- 변경 사항의 규모에 맞춰 작업량을 조정하십시오. 단일 컴포넌트 수정에는 많은 에이전트나 단계가 필요하지 않습니다. 반면, 전체 페이지 재설계에는 많은 에이전트와 단계가 필요합니다. diff의 범위에 따라 계획을 수립하십시오.
- 실패 시 조기에 중단하지 마십시오 — 할당된 테스트 범위 내에서 가능한 한 많은 버그를 찾아내야 합니다.
하위 에이전트에 단계 할당량 부여
메인 에이전트는 모든 하위 에이전트 프롬프트에 명시적인 탐색 단계 제한을 반드시 포함해야 합니다. 하위 에이전트는 자체적으로 제한을 두지 않습니다. 별도의 지시가 없는 한 작업이 완료될 때까지 실행됩니다.
대략적인 경험칙: 몇 가지 특정 점검의 경우 약 25단계, 기능적 + 적대적 + 접근성(a11y) 점검이 포함된 전체 페이지의 경우 약 40단계, 여러 페이지나 광범위한 카테고리의 경우 약 75단계입니다. 할당된 테스트가 실제로 요구하는 바에 따라 조정하십시오 — 이는 규칙이 아닌 출발점일 뿐입니다.
대략적인 지침으로: 몇 가지 특정 검사의 경우 약 25단계, 기능적 + 적대적 + 접근성(a11y) 검사가 포함된 전체 페이지의 경우 약 40단계, 여러 페이지나 광범위한 카테고리의 경우 약 75단계를 권장합니다. 할당된 테스트의 실제 요구 사항에 따라 조정하십시오. 이는 규칙이 아닌 시작점일 뿐입니다.
모든 하위 에이전트 프롬프트에는 다음이 포함되어야 합니다:
You have a budget of N browse steps (each `browse` command = 1 step). Count your steps as you go. When you reach N, stop immediately and report:
- STEP_PASS/STEP_FAIL for every test you completed
- STEP_SKIP||budget reached for every test you didn't get to
Do not retry or continue after hitting the budget.
Run only these tests: [numbered list from the merged plan]
Do not explore beyond the assigned tests.
Do NOT generate an HTML report or write any files. Return only step markers and your findings as text.
메인 에이전트는 browse (개발 서버가 가동 중인지 확인하는 경우를 제외하고) 명령어를 직접 실행해서는 안 됩니다. 모든 테스트는 서브 에이전트에서 수행됩니다.
서브 에이전트가 할당된 테스트 한도에 도달하면, 메인 에이전트는 부분 결과를 있는 그대로 수락해야 합니다. 서브 에이전트를 다시 실행하거나 재시도해서는 안 됩니다. 개발자가 어떤 테스트가 수행되지 않았는지 알 수 있도록 최종 보고서에 ‘SKIPPED’ 테스트 항목을 포함시켜야 합니다.
보고
모든 하위 에이전트는 다음 내용을 포함하여 결과를 보고합니다:
Tests: 8 | Passed: 5 | Failed: 2 | Skipped: 1 | Pages visited: 2
메인 에이전트는 다음 내용을 포함하여 최종 보고서를 생성합니다:
Tests: 20 | Passed: 14 | Failed: 4 | Skipped: 2 | Agents: 3 | Pass rate: 70%
"사용된 단계 수"는 보고하지 마십시오. 브라우즈 명령어 횟수는 구현상의 세부 사항일 뿐, 검토자에게 의미 있는 지표가 아닙니다.
테스트 철학
여러분은 적대적 테스터입니다. 여러분의 목표는 버그를 찾는 것이지, 올바름을 증명하는 것이 아닙니다.
- 테스트하는 모든 기능을 고장 나게 만들어 보십시오. 단순히 “버튼이 있는지” 확인하는 데 그치지 마십시오. — 버튼을 빠르게 두 번 클릭하고, 빈 양식을 제출하고, 500자 분량을 붙여넣고, 작업 도중 Esc 키를 눌러보십시오.
- 개발자가 생각하지 못한 부분을 테스트하십시오. 빈 상태, 오류 복구, 키보드만 사용한 탐색, 모바일 화면 넘침 현상 등을 확인하십시오.
- 모든 어설션은 증거에 기반해야 합니다. 테스트 전후의 스냅샷을 비교하십시오. 참조(ref)를 통해 특정 요소를 확인하십시오. 접근성 트리나 결정론적 검사를 통한 구체적인 증거 없이는 절대 ‘PASS’를 보고하지 마십시오.
- 재현이 가능할 만큼 충분한 세부 정보를 포함하여 실패 사항을 보고하십시오. 정확한 동작, 예상 결과, 실제 결과, 그리고 제안된 수정 사항을 포함하십시오.
주장 프로토콜
모든 테스트 단계는 반드시 구조화된 어설션을 생성해야 합니다. “이건 괜찮아 보인다”와 같은 자유 형식의 문구는 작성하지 마십시오.
단계 마커
각 테스트 단계마다 정확히 하나의 마커를 출력하십시오:
STEP_PASS||
또는
STEP_FAIL|| → |
step-id: 다음과 같은 짧은 식별자homepage-cta,form-validation-error,modal-cancelevidence: 해당 단계가 통과되었음을 입증하는 관찰 결과 (요소 참조, 텍스트 내용, URL, 평가 결과)expected → actual: 예상 결과와 실제 결과의 비교screenshot-path: 저장된 스크린샷의 경로(실패 시에만 — 아래 ‘스크린샷 캡처’ 참조)
실패 시 스크린샷 캡처
모든 STEP_FAIL에는 개발자가 시각적으로 무엇이 잘못되었는지 확인할 수 있도록 반드시 스크린샷이 첨부되어야 합니다.
테스트 단계가 실패할 경우:
# 1. Take a screenshot immediately after observing the failure
browse screenshot --path .context/ui-test-screenshots/.png
# If --path is not supported, take the screenshot and save manually:
browse screenshot
# The browse CLI will output the screenshot path — move/copy it:
cp /tmp/browse-screenshot-*.png .context/ui-test-screenshots/.png
테스트 실행을 시작할 때 스크린샷 저장 디렉터리를 설정하십시오:
mkdir -p .context/ui-test-screenshots
규칙:
- 파일 이름 = 단계 ID (예:
double-submit.png,axe-audit.png,modal-focus-trap.png) - 저장 위치
.context/ui-test-screenshots/— 이 디렉터리는 gitignored 상태이며, 개발자 및 다른 에이전트가 접근할 수 있습니다 - 병렬 실행의 경우 세션 이름을 포함하십시오:
(예:- .png signup-double-submit.png) - 오류가 발생한 순간에 스크린샷을 찍으십시오 — 복구된 상태가 아닌 오류 발생 상태를 캡처하십시오
- 시각적/레이아웃 버그의 경우, 비교를 위해 기준 상태(정상 작동 상태)의 스크린샷도 함께 촬영하십시오:
-baseline.png
검증 방법 (엄격도 순)
- 결정론적 확인(가장 확실함) —
browse eval검토할 수 있는 구조화된 데이터를 반환합니다. 예: axe-core 위반 횟수,document.title, 양식 필드 값, 콘솔 오류 배열, 요소 수. - 스냅샷 요소 일치 — 특정 역할과 텍스트를 가진 특정 요소가 접근성 트리에 존재하는지 확인합니다. 참조를 통해 확인:
@0-12 button "Save". 요소는 트리에 존재하거나 존재하지 않습니다. - 전후 비교 — 작업 전 스냅샷, 작업 수행, 작업 후 스냅샷. 트리가 예상한 대로 변경되었는지(요소가 나타났는지, 사라졌는지, 텍스트가 변경되었는지) 확인합니다.
- 스크린샷 + 시각적 판단 (가장 신뢰도가 낮음) — 접근성 트리에서 포착할 수 없는 순수 시각적 속성(색상, 간격, 레이아웃)에만 적용됩니다. 항상 구체적으로 무엇을 평가하고 있는지 명시해야 합니다.
전후 비교 패턴
이것이 핵심 검증 루프입니다. 모든 상호작용에 이 방법을 적용하십시오:
# 1. BEFORE: capture state
browse snapshot
# Record: what elements exist, their text, their refs
# 2. ACT: perform the interaction
browse click @0-12
# 3. AFTER: capture new state
browse snapshot
# Compare: what changed? What appeared? What disappeared?
# 4. ASSERT: emit marker based on comparison
# If dialog appeared: STEP_PASS|modal-open|dialog "Confirm" appeared at @0-20
# If nothing changed:
browse screenshot --path .context/ui-test-screenshots/modal-open.png
# STEP_FAIL|modal-open|expected dialog to appear → snapshot unchanged|.context/ui-test-screenshots/modal-open.png
준비
which browse || npm install -g browse
권한 요청 과다를 피하십시오
이 스킬은 많은 browse 명령어(스냅샷, 클릭, eval)를 실행합니다. 매번 승인하는 번거로움을 피하려면 browse 허용된 명령어에 다음을 추가하세요:
두 패턴을 모두 .claude/settings.json (프로젝트 수준) 또는 ~/.claude/settings.json (사용자 수준)에 추가하세요:
{
"permissions": {
"allow": [
"Bash(browse:*)",
"Bash(BROWSE_SESSION=*)"
]
}
}
첫 번째 패턴은 일반 browse 명령을 다룹니다. 두 번째 패턴은 병렬 세션(BROWSE_SESSION=signup browse open ...)을 다룹니다. 승인 요청 메시지를 피하려면 두 가지 모두 필요합니다.
모드 선택
| 대상 | 모드 | 명령어 | 인증 |
|---|---|---|---|
localhost / 127.0.0.1 |
로컬 | browse open |
필요 없음 (기본적으로 깨끗하고 격리된 로컬 브라우저) |
| 배포/스테이징 사이트 | 원격 | browse open |
Browserbase 자격 증명; 지원되는 경우 컨텍스트 사용 |
규칙: 대상 URL에 localhost 또는 127.0.0.1가 포함된 경우, 첫 번째 browse open에 --local를 전달합니다.
로컬 모드(localhost의 기본값)
browse open http://localhost:3000 --local
browse open ... --local 기본적으로 깨끗하고 격리된 로컬 브라우저를 사용하며, 이는 재현 가능한 localhost QA 실행에 가장 적합합니다.
필요한 경우에만 로컬 모드 변형을 사용하십시오:
browse open— 디버깅이 가능한 기존 로컬 Chrome을 자동으로 탐지합니다. 테스트에서 기존 로컬 로그인 정보/쿠키/상태가 명시적으로 필요한 경우에만 이 옵션을 사용하십시오.--auto-connect browse open— 특정 CDP 대상에 연결(명시적인 로컬 브라우저 연결).--cdp
원격 모드 (쿠키 동기화를 통한 배포된 사이트)
# Step 1: Sync cookies from local Chrome to Browserbase
node .claude/skills/cookie-sync/scripts/cookie-sync.mjs --domains your-app.com
# Output: Context ID: ctx_abc123
# Step 2: Open in remote mode with the synced context
SESSION_JSON="$(browse cloud sessions create --context-id ctx_abc123 --persist --keep-alive)"
SESSION_ID="$(echo "$SESSION_JSON" | jq -r .id)"
CONNECT_URL="$(echo "$SESSION_JSON" | jq -r .connectUrl)"
browse open https://staging.your-app.com --cdp "$CONNECT_URL"
browse snapshot
# ... run tests ...
browse stop
browse cloud sessions update "$SESSION_ID" --status REQUEST_RELEASE
쿠키 동기화 플래그: --domains, --context, --verified, --proxy "City,ST,US"
워크플로 A: 차이점 기반 테스트
1단계: 차이점 분석
git diff --name-only HEAD~1 # or: git diff --name-only / git diff --name-only main...HEAD
git diff HEAD~1 -- # read actual changes
변경된 파일 분류:
| 파일 패턴 | UI 영향 | 테스트 대상 |
|---|---|---|
*.tsx, *.jsx, *.vue, *.svelte |
컴포넌트 | 렌더링, 상호작용, 상태, 경계 사례 |
pages/**, app/**, src/routes/** |
경로/페이지 | 네비게이션, 페이지 로딩, 콘텐츠, 404 오류 처리 |
*.css, *.scss, *.module.css |
스타일 | 시각적 외관(스크린샷), 반응형 디자인 |
*form*, *input*, *field* |
폼 | 유효성 검사, 제출, 빈 입력, 긴 입력, 특수 문자 |
*modal*, *dialog*, *dropdown* |
상호작용 | 열기/닫기, 이스케이프, 포커스 트랩, 취소 대 확인 |
*nav*, *menu*, *header* |
탐색 | 링크, 활성 상태, 라우팅, 키보드 탐색 |
| UI 관련 파일 제외 | 없음 | 건너뛰기 — "UI 테스트 불필요"로 보고 |
2단계: 파일을 URL에 매핑
프레임워크 감지: cat package.json | grep -E '"(next|react|vue|nuxt|svelte|@sveltejs|angular|vite)"'
| 프레임워크 | 기본 포트 | 파일 → URL 패턴 |
|---|---|---|
| Next.js 앱 라우터 | 3000 | app/dashboard/page.tsx → /dashboard |
| Next.js 페이지 라우터 | 3000 | pages/about.tsx → /about |
| Vite | 5173 | 라우터 구성 확인 |
| Nuxt | 3000 | pages/index.vue → / |
| SvelteKit | 5173 | src/routes/+page.svelte → / |
| Angular | 4200 | 라우팅 모듈 확인 |
3단계: 올바른 코드가 실행되고 있는지 확인
테스트를 진행하기 전에, 개발 서버가 오래된 브랜치가 아닌 diff에 포함된 코드를 제공하고 있는지 확인하십시오.
PR이나 특정 브랜치를 테스트하는 경우:
# Check what branch is currently checked out
git branch --show-current
# If it's not the PR branch, switch to it
git fetch origin && git checkout
# Install deps — the lockfile may differ between branches
yarn install # or npm install / pnpm install
개발 서버가 이미 다른 브랜치에서 실행 중이었다면, 체크아웃 후 서버를 다시 시작하십시오.
실행 중인 개발 서버 찾기:
for port in 3000 3001 5173 4200 8080 8000 5000; do
s=$(curl -s -o /dev/null -w "%{http_code}" "http://localhost:$port" 2>/dev/null)
if [ "$s" != "000" ]; then echo "Dev server on port $port (HTTP $s)"; fi
done
아무것도 발견되지 않은 경우: 사용자에게 개발 서버를 시작하도록 안내하십시오.
실제로 렌더링되는지 확인하십시오:
그 후 browse open + browse snapshot, 접근성 트리에 오류 오버레이나 빈 본문뿐만 아니라 실제 페이지 콘텐츠(네비게이션, 제목, 상호작용 요소)가 포함되어 있는지 확인하십시오. Next.js 개발 서버는 전체 화면 빌드 오류 대화 상자를 표시하면서도 HTTP 200을 반환할 수 있습니다. 스냅샷이 비어 있거나 오류 대화 상자로 가득 차 있다면 서버에 문제가 있는 것이므로, 테스트 전에 빌드를 수정하십시오.
4단계: 테스트 계획 수립
변경된 각 영역에 대해 정상 경로 테스트와 적대적 테스트를 모두 계획하십시오:
Test Plan (based on git diff)
=============================
Changed: src/components/SignupForm.tsx (added email validation)
1. [happy] Valid email submits successfully
URL: http://localhost:3000/signup
Steps: fill valid email → submit → verify success message appears
2. [adversarial] Invalid email shows error
Steps: fill "not-an-email" → submit → verify error message appears
3. [adversarial] Empty form submission
Steps: click submit without filling anything → verify error, no crash
4. [adversarial] XSS in email field
Steps: fill "" → submit → verify sanitized/rejected
5. [adversarial] Rapid double-submit
Steps: click submit twice quickly → verify no duplicate submission
6. [adversarial] Keyboard-only flow
Steps: Tab to email → type → Tab to submit → Enter → verify success
5단계: 테스트 실행
browse stop 2>/dev/null
mkdir -p .context/ui-test-screenshots
# localhost/default QA → clean, reproducible local run
browse open http://localhost:3000 --local
각 테스트에 대해 '전/후' 패턴을 따르세요:
# Navigate
browse open http://localhost:3000/path --local
browse wait load
# BEFORE snapshot
browse snapshot
# Note the current state: elements, refs, text
# ACT
browse click @0-ref
# or: browse fill "selector" "value"
# or: browse type "text"
# or: browse press Enter
# AFTER snapshot
browse snapshot
# Compare against BEFORE: what changed?
# ASSERT with marker
# STEP_PASS|step-id|evidence OR STEP_FAIL|step-id|expected → actual
6단계: 결과 보고
## UI Test Results
### STEP_PASS|valid-email-submit|status "Thanks!" appeared at @0-42 after submit
- URL: http://localhost:3000/signup
- Before: form with email input @0-3, submit button @0-7
- Action: filled "[email protected]", clicked @0-7
- After: form replaced by status element with "Thanks! We'll be in touch."
### STEP_FAIL|double-submit|expected single submission → form submitted twice|.context/ui-test-screenshots/double-submit.png
- URL: http://localhost:3000/signup
- Before: form with submit button @0-7
- Action: clicked @0-7 twice rapidly
- After: two success toasts appeared, suggesting duplicate submission
- Screenshot: .context/ui-test-screenshots/double-submit.png
- Suggestion: disable submit button after first click, or debounce the handler
---
**Summary: 4/6 passed, 2 failed**
Failed: double-submit, xss-sanitization
Screenshots saved to `.context/ui-test-screenshots/` — open any failed step's screenshot to see the broken state.
항상 browse stop 완료 시.
7단계: HTML 보고서 생성
텍스트 보고서를 생성한 후, 검토자가 브라우저에서 바로 열 수 있는 독립형 HTML 보고서를 생성하십시오. 이 보고서는 스크린샷을 인라인(base64)으로 삽입하므로 외부 의존성 없이 단일 파일로 작동합니다.
이유: 텍스트 보고서는 에이전트 대화에는 유용하지만, 검토자(PM, 디자이너, 다른 엔지니어)는 열어 보고, 훑어보고, 공유할 수 있는 시각적 자료를 원합니다. 인라인으로 삽입된 스크린샷을 통해 오류 사항을 즉시 파악할 수 있습니다.
생성 방법
- references/report-template.html에 있는 HTML 템플릿을 확인하세요.
- 템플릿의 자리 표시자를 실제 테스트 데이터로 대체하여 보고서를 생성하세요:
| 자리 표시자 | 값 |
|---|---|
{{TITLE}} |
보고서 제목 |
---
name: ui-test
description: Runs adversarial UI tests in a real browser using the browse CLI, analyzing git diffs to test only changed areas or exploring the full app to find bugs in functionality, accessibility, responsive layout, and UX.
license: MIT
---
# UI Test — Agentic UI Testing Skill
Test UI changes in a real browser. Your job is to **try to break things**, not confirm they work.
Three workflows:
- **Diff-driven** — analyze a git diff, test only what changed
- **Exploratory** — navigate the app, find bugs the developer didn't think about
- **Parallel** — fan out independent test groups across multiple Browserbase browsers
## How Testing Works
The main agent **coordinates** — it plans test strategy, delegates to sub-agents, and merges results. Sub-agents do the actual browser testing.
### Planning: multiple angles, then execute once
**You MUST complete all three planning rounds yourself and output them before launching any sub-agents.** Planning happens in your own response — it is NOT delegated to sub-agents. Do not skip ahead to execution.
**Round 1 — Functional:** What are the core user flows? What should work? Write out each test as: action → expected result.
**Round 2 — Adversarial:** Re-read Round 1. What did you miss? Think about: different user types/roles, error paths, empty states, race conditions, edge inputs (empty, huge, special chars, rapid clicks).
**Round 3 — Coverage gaps:** Re-read Rounds 1–2. What about: accessibility (axe-core, keyboard-only), mobile viewports, console errors, visual consistency with the rest of the app?
**Deduplicate:** Merge all three rounds into one numbered list of tests. Remove overlaps. Assign each test to a group (e.g. Group A, Group B).
**Then execute once** — launch one sub-agent per group. Each sub-agent receives its specific list of tests to run, nothing more. Sub-agents do not explore or plan — they execute assigned tests and report results.
Output the three rounds, the merged plan, and the group assignments in your response before calling any Agent tool.
### Principles for splitting work
- **Sub-agents run assigned tests, not open exploration.** The main agent hands each sub-agent a specific numbered list of tests. Sub-agents do not plan, explore, or decide what to test — they execute the list and stop.
- **The bottleneck is the slowest agent** — split work so no single agent has a disproportionate share. Many small agents > few large ones.
- **Size the effort to the change** — a single component fix doesn't need many agents or many steps. A full-page redesign does. Let the scope of the diff drive the plan.
- **No early stopping on failures** — find as many bugs as possible within the assigned tests.
### Giving sub-agents a step budget
**The main agent MUST include an explicit browse step limit in every sub-agent prompt.** Sub-agents do not self-limit — they will run until done unless told otherwise.
As a rough heuristic: ~25 steps for a few targeted checks, ~40 for a full page with functional + adversarial + a11y, ~75 for multiple pages or a broad category. **Adjust based on what the assigned tests actually require** — these are starting points, not rules.
As a rough heuristic: ~25 steps for a few targeted checks, ~40 for a full page with functional + adversarial + a11y, ~75 for multiple pages or a broad category. **Adjust based on what the assigned tests actually require** — these are starting points, not rules.
Every sub-agent prompt must include:
```
You have a budget of N browse steps (each `browse` command = 1 step). Count your steps as you go. When you reach N, stop immediately and report:
- STEP_PASS/STEP_FAIL for every test you completed
- STEP_SKIP|<test-id>|budget reached for every test you didn't get to
Do not retry or continue after hitting the budget.
Run only these tests: [numbered list from the merged plan]
Do not explore beyond the assigned tests.
Do NOT generate an HTML report or write any files. Return only step markers and your findings as text.
```
The main agent should NOT run `browse` commands itself (except to verify the dev server is up). All testing happens in sub-agents.
**When a sub-agent hits its budget, the main agent accepts the partial results as-is.** Do not re-run or retry the sub-agent. Include SKIPPED tests in the final report so the developer knows what wasn't covered.
### Reporting
**Every sub-agent reports back with:**
```
Tests: 8 | Passed: 5 | Failed: 2 | Skipped: 1 | Pages visited: 2
```
**The main agent merges into a final report with:**
```
Tests: 20 | Passed: 14 | Failed: 4 | Skipped: 2 | Agents: 3 | Pass rate: 70%
```
Do not report "steps used" — browse command counts are implementation plumbing, not a meaningful metric for reviewers.
## Testing Philosophy
**You are an adversarial tester.** Your goal is to find bugs, not prove correctness.
- **Try to break every feature you test.** Don't just check "does the button exist?" — click it twice rapidly, submit empty forms, paste 500 characters, press Escape mid-flow.
- **Test what the developer didn't think about.** Empty states, error recovery, keyboard-only navigation, mobile overflow.
- **Every assertion must be evidence-based.** Compare before/after snapshots. Check specific elements by ref. Never report PASS without concrete evidence from the accessibility tree or a deterministic check.
- **Report failures with enough detail to reproduce.** Include the exact action, what you expected, what you got, and a suggested fix.
## Assertion Protocol
Every test step MUST produce a structured assertion. Do not write freeform "this looks good."
### Step markers
For each test step, emit exactly one marker:
```
STEP_PASS|<step-id>|<evidence>
```
or
```
STEP_FAIL|<step-id>|<expected> → <actual>|<screenshot-path>
```
- `step-id`: short identifier like `homepage-cta`, `form-validation-error`, `modal-cancel`
- `evidence`: what you observed that proves the step passed (element ref, text content, URL, eval result)
- `expected → actual`: what you expected vs what you got
- `screenshot-path`: path to the saved screenshot (failures only — see Screenshot Capture below)
### Screenshot Capture for Failures
**Every STEP_FAIL MUST have an accompanying screenshot** so the developer can see what went wrong visually.
When a test step fails:
```bash
# 1. Take a screenshot immediately after observing the failure
browse screenshot --path .context/ui-test-screenshots/<step-id>.png
# If --path is not supported, take the screenshot and save manually:
browse screenshot
# The browse CLI will output the screenshot path — move/copy it:
cp /tmp/browse-screenshot-*.png .context/ui-test-screenshots/<step-id>.png
```
Setup the screenshot directory at the start of any test run:
```bash
mkdir -p .context/ui-test-screenshots
```
**Rules:**
- File name = step-id (e.g., `double-submit.png`, `axe-audit.png`, `modal-focus-trap.png`)
- Store in `.context/ui-test-screenshots/` — this directory is gitignored and accessible to the developer and other agents
- For parallel runs, include the session name: `<session>-<step-id>.png` (e.g., `signup-double-submit.png`)
- Take the screenshot at the moment of failure — capture the broken state, not after recovery
- For visual/layout bugs, also screenshot the baseline (working state) for comparison: `<step-id>-baseline.png`
### How to verify (in order of rigor)
1. **Deterministic check** (strongest) — `browse eval` returns structured data you can inspect. Examples: axe-core violation count, `document.title`, form field value, console error array, element count.
2. **Snapshot element match** — a specific element with a specific role and text exists in the accessibility tree. Check by ref: `@0-12 button "Save"`. An element either exists in the tree or it doesn't.
3. **Before/after comparison** — snapshot before action, act, snapshot after. Verify the tree changed in the expected way (element appeared, disappeared, text changed).
4. **Screenshot + visual judgment** (weakest) — only for visual-only properties (color, spacing, layout) that the accessibility tree cannot capture. Always accompany with what specifically you're evaluating.
### Before/after comparison pattern
This is the core verification loop. Use it for every interaction:
```bash
# 1. BEFORE: capture state
browse snapshot
# Record: what elements exist, their text, their refs
# 2. ACT: perform the interaction
browse click @0-12
# 3. AFTER: capture new state
browse snapshot
# Compare: what changed? What appeared? What disappeared?
# 4. ASSERT: emit marker based on comparison
# If dialog appeared: STEP_PASS|modal-open|dialog "Confirm" appeared at @0-20
# If nothing changed:
browse screenshot --path .context/ui-test-screenshots/modal-open.png
# STEP_FAIL|modal-open|expected dialog to appear → snapshot unchanged|.context/ui-test-screenshots/modal-open.png
```
## Setup
```bash
which browse || npm install -g browse
```
### Avoid permission fatigue
This skill runs many `browse` commands (snapshots, clicks, evals). To avoid approving each one, add `browse` to your allowed commands:
Add both patterns to `.claude/settings.json` (project-level) or `~/.claude/settings.json` (user-level):
```json
{
"permissions": {
"allow": [
"Bash(browse:*)",
"Bash(BROWSE_SESSION=*)"
]
}
}
```
The first pattern covers plain `browse` commands. The second covers parallel sessions (`BROWSE_SESSION=signup browse open ...`). Both are needed to avoid approval prompts.
## Mode Selection
| Target | Mode | Command | Auth |
|--------|------|---------|------|
| `localhost` / `127.0.0.1` | Local | `browse open <url> --local` | None needed (clean isolated local browser by default) |
| Deployed/staging site | Remote | `browse open <url> --remote` | Browserbase credentials; use contexts where supported |
**Rule: If the target URL contains `localhost` or `127.0.0.1`, pass `--local` on the first `browse open`.**
### Local Mode (default for localhost)
```bash
browse open http://localhost:3000 --local
```
`browse open ... --local` uses a clean isolated local browser by default, which is best for reproducible localhost QA runs.
Use local-mode variants only when needed:
- `browse open <url> --auto-connect` — auto-discover an existing debuggable local Chrome. Use this only when the test explicitly needs existing local login/cookies/state.
- `browse open <url> --cdp <port|url>` — attach to a specific CDP target (explicit local browser attach).
### Remote Mode (deployed sites via cookie-sync)
```bash
# Step 1: Sync cookies from local Chrome to Browserbase
node .claude/skills/cookie-sync/scripts/cookie-sync.mjs --domains your-app.com
# Output: Context ID: ctx_abc123
# Step 2: Open in remote mode with the synced context
SESSION_JSON="$(browse cloud sessions create --context-id ctx_abc123 --persist --keep-alive)"
SESSION_ID="$(echo "$SESSION_JSON" | jq -r .id)"
CONNECT_URL="$(echo "$SESSION_JSON" | jq -r .connectUrl)"
browse open https://staging.your-app.com --cdp "$CONNECT_URL"
browse snapshot
# ... run tests ...
browse stop
browse cloud sessions update "$SESSION_ID" --status REQUEST_RELEASE
```
Cookie-sync flags: `--domains`, `--context`, `--verified`, `--proxy "City,ST,US"`
## Workflow A: Diff-Driven Testing
### Phase 1: Analyze the diff
```bash
git diff --name-only HEAD~1 # or: git diff --name-only / git diff --name-only main...HEAD
git diff HEAD~1 -- <file> # read actual changes
```
Categorize changed files:
| File pattern | UI impact | What to test |
|-------------|-----------|--------------|
| `*.tsx`, `*.jsx`, `*.vue`, `*.svelte` | Component | Render, interaction, state, edge cases |
| `pages/**`, `app/**`, `src/routes/**` | Route/page | Navigation, page load, content, 404 handling |
| `*.css`, `*.scss`, `*.module.css` | Style | Visual appearance (screenshot), responsive |
| `*form*`, `*input*`, `*field*` | Form | Validation, submission, empty input, long input, special chars |
| `*modal*`, `*dialog*`, `*dropdown*` | Interactive | Open/close, escape, focus trap, cancel vs confirm |
| `*nav*`, `*menu*`, `*header*` | Navigation | Links, active states, routing, keyboard nav |
| Non-UI files only | None | Skip — report "no UI tests needed" |
### Phase 2: Map files to URLs
Detect framework: `cat package.json | grep -E '"(next|react|vue|nuxt|svelte|@sveltejs|angular|vite)"'`
| Framework | Default port | File → URL pattern |
|-----------|-------------|-----|
| Next.js App Router | 3000 | `app/dashboard/page.tsx` → `/dashboard` |
| Next.js Pages Router | 3000 | `pages/about.tsx` → `/about` |
| Vite | 5173 | Check router config |
| Nuxt | 3000 | `pages/index.vue` → `/` |
| SvelteKit | 5173 | `src/routes/+page.svelte` → `/` |
| Angular | 4200 | Check routing module |
### Phase 3: Ensure the right code is running
Before testing, verify the dev server is serving the code from the diff — not a stale branch.
**If testing a PR or specific branch:**
```bash
# Check what branch is currently checked out
git branch --show-current
# If it's not the PR branch, switch to it
git fetch origin <branch> && git checkout <branch>
# Install deps — the lockfile may differ between branches
yarn install # or npm install / pnpm install
```
If the dev server was already running on a different branch, restart it after checkout.
**Find a running dev server:**
```bash
for port in 3000 3001 5173 4200 8080 8000 5000; do
s=$(curl -s -o /dev/null -w "%{http_code}" "http://localhost:$port" 2>/dev/null)
if [ "$s" != "000" ]; then echo "Dev server on port $port (HTTP $s)"; fi
done
```
If nothing found: tell the user to start their dev server.
**Verify it actually renders:**
After `browse open` + `browse snapshot`, check that the accessibility tree contains real page content (navigation, headings, interactive elements) — not just an error overlay or empty body. Next.js dev servers can return HTTP 200 while showing a full-screen build error dialog. If the snapshot is empty or dominated by an error dialog, the server is broken — fix the build before testing.
### Phase 4: Generate test plan
For each changed area, plan **both happy path AND adversarial tests**:
```
Test Plan (based on git diff)
=============================
Changed: src/components/SignupForm.tsx (added email validation)
1. [happy] Valid email submits successfully
URL: http://localhost:3000/signup
Steps: fill valid email → submit → verify success message appears
2. [adversarial] Invalid email shows error
Steps: fill "not-an-email" → submit → verify error message appears
3. [adversarial] Empty form submission
Steps: click submit without filling anything → verify error, no crash
4. [adversarial] XSS in email field
Steps: fill "<script>alert(1)</script>" → submit → verify sanitized/rejected
5. [adversarial] Rapid double-submit
Steps: click submit twice quickly → verify no duplicate submission
6. [adversarial] Keyboard-only flow
Steps: Tab to email → type → Tab to submit → Enter → verify success
```
### Phase 5: Execute tests
```bash
browse stop 2>/dev/null
mkdir -p .context/ui-test-screenshots
# localhost/default QA → clean, reproducible local run
browse open http://localhost:3000 --local
```
For each test, follow the **before/after pattern**:
```bash
# Navigate
browse open http://localhost:3000/path --local
browse wait load
# BEFORE snapshot
browse snapshot
# Note the current state: elements, refs, text
# ACT
browse click @0-ref
# or: browse fill "selector" "value"
# or: browse type "text"
# or: browse press Enter
# AFTER snapshot
browse snapshot
# Compare against BEFORE: what changed?
# ASSERT with marker
# STEP_PASS|step-id|evidence OR STEP_FAIL|step-id|expected → actual
```
### Phase 6: Report results
```
## UI Test Results
### STEP_PASS|valid-email-submit|status "Thanks!" appeared at @0-42 after submit
- URL: http://localhost:3000/signup
- Before: form with email input @0-3, submit button @0-7
- Action: filled "[email protected]", clicked @0-7
- After: form replaced by status element with "Thanks! We'll be in touch."
### STEP_FAIL|double-submit|expected single submission → form submitted twice|.context/ui-test-screenshots/double-submit.png
- URL: http://localhost:3000/signup
- Before: form with submit button @0-7
- Action: clicked @0-7 twice rapidly
- After: two success toasts appeared, suggesting duplicate submission
- Screenshot: .context/ui-test-screenshots/double-submit.png
- Suggestion: disable submit button after first click, or debounce the handler
---
**Summary: 4/6 passed, 2 failed**
Failed: double-submit, xss-sanitization
Screenshots saved to `.context/ui-test-screenshots/` — open any failed step's screenshot to see the broken state.
```
Always `browse stop` when done.
### Phase 7: Generate HTML report
After producing the text report, generate a standalone HTML report that a reviewer can open in a browser. The report embeds screenshots inline (base64) so it works as a single file — no external dependencies.
**Why:** Text reports are good for the agent conversation, but reviewers (PMs, designers, other engineers) want a visual artifact they can open, scan, and share. Screenshots inline make failures immediately obvious.
#### How to generate
1. Read the HTML template at [references/report-template.html](references/report-template.html)
2. Build the report by replacing the template placeholders with actual test data:
| Placeholder | Value |
|-------------|-------|
| `{{TITLE}}` | Report title for `<title>` tag (e.g., "UI Test: PR #1234 — OAuth Settings") |
| `{{TITLE_HTML}}` | Report title for the visible `<h1>`. If a PR URL is available, wrap the PR reference in an `<a>` tag so it's clickable (e.g., `UI Test: <a href="https://github.com/org/repo/pull/1234">PR #1234</a> — OAuth Settings`). If no URL, use plain text same as `{{TITLE}}`. |
| `{{META}}` | One-line context: date, app URL, user, branch |
| `{{TOTAL_TESTS}}` | Total STEP_PASS + STEP_FAIL count |
| `{{AGENT_COUNT}}` | Number of sub-agents that ran |
| `{{PASS_COUNT}}` | Number of STEP_PASS |
| `{{FAIL_COUNT}}` | Number of STEP_FAIL |
| `{{PASS_RATE}}` | Integer percentage (e.g., "92") |
| `{{RATE_CLASS}}` | `good` (≥90%), `warn` (70–89%), `bad` (<70%) |
| `{{FAILURES_SECTION}}` | HTML for failed test cards (see below) |
| `{{PASSES_SECTION}}` | HTML for passed test cards (see below) |
3. For each test result, generate a `<details>` card. Failed tests should be **open by default** so reviewers see them immediately:
```html
<!-- Failed test card (open by default) -->
<div class="section">
<h2>Failures <span class="count">{{FAIL_COUNT}}</span></h2>
<details class="test-card fail" open>
<summary>
<span class="badge fail">FAIL</span>
<span class="step-id">step-id-here</span>
<span class="evidence">expected → actual</span>
</summary>
<div class="body">
<dl>
<dt>URL</dt><dd>http://localhost:3000/path</dd>
<dt>Action</dt><dd>What was done</dd>
<dt>Expected</dt><dd>What should have happened</dd>
<dt>Actual</dt><dd>What happened instead</dd>
</dl>
<div class="suggestion">Fix: description of suggested fix</div>
<div class="screenshot">
<img src="data:image/png;base64,..." alt="Screenshot of failure">
<div class="caption">step-id.png — captured at moment of failure</div>
</div>
</div>
</details>
</div>
<!-- Passed test card (collapsed by default) -->
<div class="section">
<h2>Passed <span class="count">{{PASS_COUNT}}</span></h2>
<details class="test-card pass">
<summary>
<span class="badge pass">PASS</span>
<span class="step-id">step-id-here</span>
<span class="evidence">evidence summary</span>
</summary>
<div class="body">
<dl>
<dt>URL</dt><dd>http://localhost:3000/path</dd>
<dt>Evidence</dt><dd>What was observed</dd>
</dl>
</div>
</details>
</div>
```
4. **Embed screenshots as base64** so the HTML is fully self-contained:
```bash
# Convert screenshot to base64 data URI
base64 -i .context/ui-test-screenshots/step-id.png | tr -d '\n'
# Use as: src="data:image/png;base64,<output>"
```
Read each screenshot file referenced in STEP_FAIL markers, base64-encode it, and embed it as an `<img src="data:image/png;base64,...">` in the corresponding test card. For STEP_PASS, only embed a screenshot if one was explicitly taken (e.g., baseline screenshots).
5. Write the final HTML to `.context/ui-test-report.html`:
```bash
# Write the generated HTML
cat > .context/ui-test-report.html << 'REPORT_EOF'
<!DOCTYPE html>
...generated report...
REPORT_EOF
# Open it for the reviewer
open .context/ui-test-report.html # macOS
# xdg-open .context/ui-test-report.html # Linux
```
6. Tell the user: `Report saved to .context/ui-test-report.html` and offer to open it.
**Rules:**
- Failures section comes before passes — reviewers care about what's broken first
- Failed cards are `open` by default; passed cards are collapsed
- Every STEP_FAIL card MUST have an embedded screenshot — if the screenshot file is missing, note it in the card
- Include the suggestion/fix in each failure card if one was provided
- The report must work offline — no CDN links, no external assets
- Keep the HTML under 5MB — if screenshots push it over, reduce image quality or skip baseline screenshots for passes
## Adversarial Test Patterns
Apply these to every interactive element you test. Read [references/adversarial-patterns.md](references/adversarial-patterns.md) for the full pattern library (forms, modals, navigation, error states, keyboard accessibility).
## Deterministic Checks
These produce structured data, not judgment calls. Use them as the strongest form of assertion.
| Check | What it catches | Assertion |
|-------|----------------|-----------|
| axe-core | WCAG violations | `violations.length === 0` |
| Console errors | Runtime exceptions, failed requests | empty error array |
| Broken images | Missing/failed image loads | no images with `naturalWidth === 0` |
| Form labels | Inputs without accessible labels | every input has `hasLabel: true` |
For the exact `browse eval` recipes, read [references/browser-recipes.md](references/browser-recipes.md).
## Workflow B: Exploratory Testing
No diff, no plan — just open the app and try to break it. Use this when the user says "test my app", "find bugs", or "QA this site."
### Approach
1. **Discover the app** — read `package.json` to detect the framework, then open the root URL and snapshot to see what's there
2. **Navigate everything** — click through nav links, visit every reachable page, note what exists
3. **Test what you find** — for each page, apply the adversarial patterns below (forms, modals, navigation, keyboard, error states)
4. **Run deterministic checks** — axe-core, console errors, broken images, form labels on every page
5. **Report findings** — use STEP_PASS/STEP_FAIL markers, include reproduction steps for failures
Don't try to be systematic about coverage. Just explore like a user would, but with the intent to break things. The agent is good at this — let it roam.
### Tips for exploratory runs
- Start with the homepage, then follow the navigation naturally
- Try the 404 page (`/does-not-exist`) — is it custom or default?
- Look for empty states (pages with no data)
- Test forms with garbage input before valid input
- Check mobile viewport (375px) on every page — does it overflow?
- If the app has auth, use cookie-sync first
## Workflow C: Parallel Testing
Run independent test groups concurrently using named `browse` sessions (`BROWSE_SESSION=<name>`). Each session gets its own browser. Works with both local and remote mode.
Use when testing multiple pages or categories and you want faster wall clock time.
Read [references/parallel-testing.md](references/parallel-testing.md) for the full workflow: session setup, agent fan-out, cookie-sync for auth, and result merging.
## Design Consistency
Check whether changed UI matches the rest of the app visually. Read [references/design-consistency.md](references/design-consistency.md) when doing visual or design checks.
## Test Categories
| Category | How | Assertion type |
|----------|-----|---------------|
| Accessibility | axe-core + keyboard nav | Deterministic (violation count) |
| Visual Quality | Screenshot + heuristic evaluation | Visual judgment (weakest — note specifics) |
| Responsive | Viewport sweep + screenshots | Visual + deterministic (overflow check) |
| Console Health | Console capture eval | Deterministic (error count) |
| UX Heuristics | Snapshot + Laws of UX + Nielsen's | Structured judgment (cite specific heuristic) |
| Error States | Navigate to empty/error states | Before/after comparison |
| Data Display | Snapshot on tables/dashboards | Element match (column count, formatting) |
| Design Consistency | Screenshot baseline + changed page comparison | Visual judgment (cite specific property) |
| Exploratory | Free navigation + adversarial testing | Before/after + judgment |
Reference guides (load on demand):
- **Adversarial patterns** — [references/adversarial-patterns.md](references/adversarial-patterns.md) — load when testing forms, modals, navigation, or keyboard a11y
- **Browser recipes** — [references/browser-recipes.md](references/browser-recipes.md) — load when running deterministic checks (axe-core, console, images, form labels)
- **Exploratory testing** — [references/exploratory-testing.md](references/exploratory-testing.md) — load for Workflow B (no diff, open exploration)
- **UX heuristics** — [references/ux-heuristics.md](references/ux-heuristics.md) — load when evaluating UX quality or citing specific heuristics
- **Design system** — [references/design-system.example.md](references/design-system.example.md) — template for users to customize
- **Design consistency** — [references/design-consistency.md](references/design-consistency.md) — load when doing visual consistency checks
- **Parallel testing** — [references/parallel-testing.md](references/parallel-testing.md) — load for Workflow C (concurrent sessions)
- **Report template** — [references/report-template.html](references/report-template.html) — HTML template for Phase 7 report generation
For worked examples with exact commands, read [EXAMPLES.md](EXAMPLES.md) if you need to see the assertion protocol in action.
## Best Practices
1. **Be adversarial** — try to break things, don't just confirm they work
2. **Every assertion needs evidence** — snapshot ref, eval result, or before/after diff
3. **Before/after for every interaction** — snapshot, act, snapshot, compare
4. **Screenshot every failure** — `browse screenshot` immediately on STEP_FAIL, save to `.context/ui-test-screenshots/<step-id>.png`
5. **Deterministic checks first** — axe-core, console errors, form labels before visual judgment
6. **For localhost, start with clean local mode** — pass `--local` on the first `browse open` for reproducible runs; use `--auto-connect` only when existing local state is required
7. **Always `browse stop` when done** — for parallel runs, stop every named session
8. **Report failures with reproduction steps** — action, expected, actual, screenshot path, suggestion
9. **Parallelize independent tests** — use Workflow C with named sessions when testing multiple pages or categories on a deployed site
## Troubleshooting
- **"No active page"**: `browse stop`, retry. For zombies: `pkill -f "browse.*daemon"`
- **Dev server not responding**: `curl http://localhost:<port>` — ask user to start it
- **`browse eval` with `await` fails**: Use `.then()` instead — `browse eval` doesn't support top-level await
- **Element ref not found**: `browse snapshot` again — refs change on page update
- **Blank snapshot**: `browse wait load` or `browse wait selector ".expected"` before snapshotting
- **SPA deep links 404**: Navigate to `/` first, then click through
- **Remote auth fails**: Re-run cookie-sync with `--context <id>`, try `--verified`
- **Parallel session conflicts**: Ensure every `browse` command uses `BROWSE_SESSION=<name>` — without it, commands go to the default session
- **Session not stopping**: `BROWSE_SESSION=<name> browse stop`. For zombies: `pkill -f "browse.*<name>.*daemon"`





집
