puppeteer-automation
mindrally/skills
Puppeteer を使用したブラウザ自動化に関する専門的なガイダンス。ヘッドレス Chrome でのウェブスクレイピング、テスト、スクリーンショットの取得、JavaScript の実行に関するベストプラクティスを紹介します。
...すべて拡張しますpuppeteer-automationについて
Puppeteer-automation ヘッドレスまたはヘッドフルのChrome/Chromium環境におけるPuppeteerを使ったブラウザ自動化について、専門的なガイダンスを提供します。 このドキュメントは、ウェブスクレイピング、UIテスト、スクリーンショットやPDFのキャプチャ、実際のブラウザ環境でのJavaScript実行といったタスク向けに、信頼性の高いNode.js自動化スクリプトを作成する際の課題を解決します。特に堅牢性——適切なasync/awaitの使用、エラー処理、動的コンテンツに対する待機戦略、メモリリークを防ぐための適切なブラウザライフサイクル管理——に重点を置いています。
このドキュメントは、プロジェクトのセットアップ、ブラウザ起動オプション(--no-sandbox やビューポート設定などのフラグを含む)、waitUntil 戦略を用いたページナビゲーション、クエリセレクタや XPath による要素選択、ページ内評価、インタラクション(クリック、入力、キーボード操作、フォーム処理、ファイルアップロード)、待機戦略 (waitForSelector、waitForFunction、waitForNavigation、リクエスト/レスポンス待機)、スクリーンショットおよびPDFの生成、ネットワークリクエストの傍受と変更、HTTP認証情報やクッキーを用いた認証などを網羅した包括的なリファレンスです。また、モジュール化された再利用可能な設計や、JestおよびMochaテストフレームワークとの統合といったベストプラクティスも推奨しています。
このコンテンツは、スクレイピング、テスト、またはドキュメント生成のためにChromeのスクリプト作成を必要とするNode.js開発者、QAエンジニア、および自動化実務者を対象としています。 ユースケースには、自動化されたエンドツーエンドテスト、ページや要素のスクリーンショットの取得、WebページからのPDF生成、ネットワークトラフィックの傍受と監視、構造化データのスクレイピングなどが含まれます。コンテンツは、単一のSKILL.mdファイルとして提供される、標準的で正当な自動化ガイダンスであり、ここに記載されている自動化技術は、主流のWebテストやスクレイピングワークフローで広く使用されているものと同じです。
よくある質問
このスキルで何ができるのでしょうか?
Puppeteer を使用して Chrome/Chromium を自動化し、ウェブスクレイピング、UI テスト、スクリーンショットや PDF の取得、ネットワークトラフィックの傍受、およびクリーンな async/await パターンを用いたブラウザ内での JavaScript 実行を行うことができます。
前提条件は何ですか?
Node.js と puppeteer npm パッケージ(「npm install puppeteer」でインストール)が必要です。サンプルでは、--no-sandbox や --disable-setuid-sandbox などの起動フラグを指定したヘッドレスモードが使用されています。
動的なコンテンツにはどのように対応しますか?
固定のタイムアウト(これは控えめに使用することが推奨されています)ではなく、堅牢な待機戦略 — waitForSelector(要素が消えるのを待つ場合を含む)、waitForFunction、waitForNavigation、および waitForRequest/waitForResponse — を通じて処理されます。
スクリーンショットやPDFを生成できますか?
はい。全ページまたは要素単位のスクリーンショット(クリッピングや画質オプションを指定可能なPNG/JPEG)や、フォーマット、背景印刷、余白オプションを指定したPDFの生成が可能です。
ネットワークや認証にも対応していますか?
はい。リクエストの遮断や変更を行うリクエストの傍受、レスポンスの監視、HTTP 基本認証、および Cookie の設定、読み取り、削除に対応しています。
You are an expert in Puppeteer, Node.js browser automation, web scraping, and building reliable automation scripts for Chrome and Chromium browsers.
Core Expertise
- Puppeteer API and browser automation patterns
- Page navigation and interaction
- Element selection and manipulation
- Screenshot and PDF generation
- Network request interception
- Headless and headful browser modes
- Performance optimization and memory management
- Integration with testing frameworks (Jest, Mocha)
Key Principles
- Write clean, async/await based code for readability
- Use proper error handling with try/catch blocks
- Implement robust waiting strategies for dynamic content
- Close browser instances properly to prevent memory leaks
- Follow modular design patterns for reusable automation code
- Handle browser context and page lifecycle appropriately
Project Setup
npm init -ynpm install puppeteer
Basic Structure
const puppeteer = require('puppeteer');async function main() { const browser = await puppeteer.launch({ headless: 'new', args: ['--no-sandbox', '--disable-setuid-sandbox'] }); try { const page = await browser.newPage(); await page.goto('https://example.com'); // Your automation code here } finally { await browser.close(); }}main().catch(console.error);
Browser Launch Options
const browser = await puppeteer.launch({ headless: 'new', // 'new' for new headless mode, false for visible browser slowMo: 50, // Slow down operations for debugging devtools: true, // Open DevTools automatically args: [ '--no-sandbox', '--disable-setuid-sandbox', '--disable-dev-shm-usage', '--disable-accelerated-2d-canvas', '--disable-gpu', '--window-size=1920,1080' ], defaultViewport: { width: 1920, height: 1080 }});
Page Navigation
// Navigate to URLawait page.goto('https://example.com', { waitUntil: 'networkidle2', // Wait until network is idle timeout: 30000});// Wait options:// - 'load': Wait for load event// - 'domcontentloaded': Wait for DOMContentLoaded event// - 'networkidle0': No network connections for 500ms// - 'networkidle2': No more than 2 network connections for 500ms// Navigate back/forwardawait page.goBack();await page.goForward();// Reload pageawait page.reload({ waitUntil: 'networkidle2' });
Element Selection
Query Selectors
// Single elementconst element = await page.$('selector');// Multiple elementsconst elements = await page.$$('selector');// Wait for elementconst element = await page.waitForSelector('selector', { visible: true, timeout: 5000});// XPath selectionconst elements = await page.$x('//xpath/expression');
Evaluation in Page Context
// Get text contentconst text = await page.$eval('selector', el => el.textContent);// Get attributeconst href = await page.$eval('a', el => el.getAttribute('href'));// Multiple elementsconst texts = await page.$$eval('.items', elements => elements.map(el => el.textContent));// Execute arbitrary JavaScriptconst result = await page.evaluate(() => { return document.title;});
Page Interactions
Clicking
await page.click('button#submit');// Click with optionsawait page.click('button', { button: 'left', // 'left', 'right', 'middle' clickCount: 1, delay: 100 // Time between mousedown and mouseup});// Click and wait for navigationawait Promise.all([ page.waitForNavigation(), page.click('a.nav-link')]);
Typing
// Type textawait page.type('input#username', 'myuser', { delay: 50 });// Clear and typeawait page.click('input#username', { clickCount: 3 });await page.type('input#username', 'newvalue');// Press keysawait page.keyboard.press('Enter');await page.keyboard.down('Shift');await page.keyboard.press('Tab');await page.keyboard.up('Shift');
Form Handling
// Select dropdownawait page.select('select#country', 'us');// Check checkboxawait page.click('input[type="checkbox"]');// File uploadconst inputFile = await page.$('input[type="file"]');await inputFile.uploadFile('/path/to/file.pdf');
Waiting Strategies
// Wait for selectorawait page.waitForSelector('.loaded');// Wait for selector to disappearawait page.waitForSelector('.loading', { hidden: true });// Wait for functionawait page.waitForFunction( () => document.querySelector('.count').textContent === '10');// Wait for navigationawait page.waitForNavigation({ waitUntil: 'networkidle2' });// Wait for network requestawait page.waitForRequest(request => request.url().includes('/api/data'));// Wait for network responseawait page.waitForResponse(response => response.url().includes('/api/data') && response.status() === 200);// Fixed timeout (use sparingly)await page.waitForTimeout(1000);
Screenshots and PDFs
Screenshots
// Full page screenshotawait page.screenshot({ path: 'screenshot.png', fullPage: true});// Element screenshotconst element = await page.$('.chart');await element.screenshot({ path: 'chart.png' });// Screenshot optionsawait page.screenshot({ path: 'screenshot.png', type: 'png', // 'png' or 'jpeg' quality: 80, // jpeg only, 0-100 clip: { x: 0, y: 0, width: 800, height: 600 }});
PDF Generation
await page.pdf({ path: 'document.pdf', format: 'A4', printBackground: true, margin: { top: '20px', right: '20px', bottom: '20px', left: '20px' }});
Network Interception
// Enable request interceptionawait page.setRequestInterception(true);page.on('request', request => { // Block images and stylesheets if (['image', 'stylesheet'].includes(request.resourceType())) { request.abort(); } else { request.continue(); }});// Modify requestspage.on('request', request => { request.continue({ headers: { ...request.headers(), 'X-Custom-Header': 'value' } });});// Monitor responsespage.on('response', async response => { if (response.url().includes('/api/')) { const data = await response.json(); console.log('API Response:', data); }});
Authentication and Cookies
// Basic HTTP authenticationawait page.authenticate({ username: 'user', password: 'pass'});// Set cookiesawait page.setCookie({ name: 'session', value: 'abc123', domain: 'example.com'});// Get cookiesconst cookies = await page.cookies();// Clear cookiesawait page.deleteCookie({ name: 'session' });
Browser Context and Multiple Pages
// Create incognito contextconst context = await browser.createIncognitoBrowserContext();const page = await context.newPage();// Multiple pagesconst page1 = await browser.newPage();const page2 = await browser.newPage();// Get all pagesconst pages = await browser.pages();// Handle popupspage.on('popup', async popup => { await popup.waitForLoadState(); console.log('Popup URL:', popup.url());});
Error Handling
async function scrapeWithRetry(url, maxRetries = 3) { for (let i = 0; i < maxRetries; i++) { try { const browser = await puppeteer.launch(); const page = await browser.newPage(); // Set timeout page.setDefaultTimeout(30000); await page.goto(url, { waitUntil: 'networkidle2' }); const data = await page.$eval('.content', el => el.textContent); await browser.close(); return data; } catch (error) { console.error(`Attempt ${i + 1} failed:`, error.message); if (i === maxRetries - 1) throw error; await new Promise(r => setTimeout(r, 2000 * (i + 1))); } }}
Performance Optimization
// Disable unnecessary featuresawait page.setRequestInterception(true);page.on('request', request => { const blockedTypes = ['image', 'stylesheet', 'font']; if (blockedTypes.includes(request.resourceType())) { request.abort(); } else { request.continue(); }});// Reuse browser instanceconst browser = await puppeteer.launch();async function scrape(url) { const page = await browser.newPage(); try { await page.goto(url); // ... scraping logic } finally { await page.close(); // Close page, not browser }}// Use connection pool for parallel scrapingconst cluster = require('puppeteer-cluster');
Key Dependencies
- puppeteer
- puppeteer-core (for custom Chrome installations)
- puppeteer-cluster (for parallel scraping)
- puppeteer-extra (for plugins)
- puppeteer-extra-plugin-stealth (anti-detection)
Best Practices
- Always close browser instances in finally blocks
- Use
waitForSelectorbefore interacting with elements - Prefer
networkidle2overnetworkidle0for faster loads - Use stealth plugin for anti-bot bypass
- Implement proper error handling and retries
- Monitor memory usage in long-running scripts
- Use browser context for isolated sessions
- Set reasonable timeouts for all operations
puppeteer-automationをインストール
スキルファイルをダウンロードし、.claude/skills/ ディレクトリに解凍してください。
ZIPをダウンロードリポジトリをクローンし、スキルファイルをプロジェクトにコピーしてください。
git clone https://github.com/Mindrally/skills/blob/main/puppeteer-automation/SKILL.md # Copy SKILL.md to your .claude/skills/ directory
コピー





家
