選項
首頁首頁 Skill 資料庫管理 dummy-dataset

dummy-dataset

phuryn/pm-skills phuryn/pm-skills

產生逼真的模擬資料集供測試使用,並可自訂欄位、限制條件及輸出格式(CSV、JSON、SQL、Python 腳本)。

...展開全部
0
更新時間 2026-09-29

模擬資料集生成

生成逼真的模擬資料集以供測試,可自訂欄位、限制條件及輸出格式(CSV、JSON、SQL、Python 腳本)。可產生可執行的腳本或直接可用的資料檔案,供立即使用。

適用情境:建立測試資料、產生樣本資料集、為開發建立逼真的模擬資料,或填充測試環境。

參數:

  • $PRODUCT:產品或系統名稱
  • $DATASET_TYPE: 資料類型(例如:客戶回饋、交易紀錄、使用者檔案)
  • $ROWS: 要生成的列數(預設:100)
  • $COLUMNS: 需包含的特定欄位或字段
  • $FORMAT: 輸出格式(CSV、JSON、SQL、Python 腳本)
  • $CONSTRAINTS: 其他限制條件或業務規則

逐步流程

  1. 識別資料集類型 — 了解資料領域
  2. 定義欄位規格 — 名稱、資料類型及數值範圍
  3. 確定列數 — 需要多少筆樣本記錄
  4. 選擇輸出格式 — CSV、JSON、SQL INSERT 或 Python 腳本
  5. 套用符合實際情況的模式 — 確保資料看起來真實且有效
  6. 加入業務限制條件 — 遵循業務邏輯與關聯性
  7. 生成或編寫資料腳本 — 建立可執行的輸出
  8. 驗證輸出結果 — 確保資料品質與完整性

範本:Python 腳本輸出

import csv
import json
from datetime import datetime, timedelta
import random

# Configuration
ROWS = $ROWS
FILENAME = "$DATASET_TYPE.csv"

# Column definitions with realistic value generators
columns = {
    "id": "auto-increment",
    "name": "first_last_name",
    "email": "email",
    "created_at": "timestamp",
    # Add more columns...
}

def generate_dataset():
    """Generate realistic dummy dataset"""
    data = []
    for i in range(1, ROWS + 1):
        record = {
            "id": f"U{i:06d}",
            # Generate values based on column definitions
        }
        data.append(record)
    return data

def save_as_csv(data, filename):
    """Save dataset as CSV"""
    with open(filename, 'w', newline='') as f:
        writer = csv.DictWriter(f, fieldnames=data[0].keys())
        writer.writeheader()
        writer.writerows(data)

if __name__ == "__main__":
    dataset = generate_dataset()
    save_as_csv(dataset, FILENAME)
    print(f"Generated {len(dataset)} records in {FILENAME}")

範例資料集規格

資料集類型:客戶回饋

欄位:

  • feedback_id(自動遞增,U001、U002...)
  • customer_name(真實姓名)
  • email(有效電子郵件格式)
  • feedback_date(過去 90 天內的日期)
  • 評分 (1-5 顆星)
  • 類別 (錯誤、功能請求、投訴、讚揚)
  • text(真實的回饋內容)
  • 產品類別(電子產品、服飾、家居)

限制條件:

  • 評分分布:40% 為 5 星、30% 為 4 星、20% 為 3 星、10% 為 1 至 2 星
  • 僅限「錯誤」類別且評分介於 1 至 3 星
  • 功能請求僅限 3 至 5 星評分
  • 電子郵件網域須符合現實情境(gmail、yahoo、company.com)

輸出成果

  • 可直接執行的 Python 腳本 或 直接資料檔案
  • 具備正確標題與格式設定的 CSV 檔案
  • 結構與資料類型均正確的 JSON 檔案
  • 用於資料庫資料填入的 SQL INSERT 語句
  • 資料驗證與符合限制條件
  • 符合實際情況且適用于業務的數值
  • 資料生成邏輯的文件說明
  • 資料集的使用快速入門指南

輸出格式

CSV:平面表格格式,便於匯入試算表和資料庫

JSON:嵌套結構,非常適合用於 API 和 NoSQL 資料庫

SQL:INSERT 語句,可直接在關係型資料庫上執行

Python 腳本:適用於自訂或大型資料集的可執行產生器

在 GitHub 上查看
---
name: dummy-dataset
description: Generate realistic dummy datasets for testing with customizable columns, constraints, and output formats (CSV, JSON, SQL, Python script).
---
# Dummy Dataset Generation

Generate realistic dummy datasets for testing with customizable columns, constraints, and output formats (CSV, JSON, SQL, Python script). Creates executable scripts or direct data files for immediate use.

**Use when:** Creating test data, generating sample datasets, building realistic mock data for development, or populating test environments.

**Arguments:**
- `$PRODUCT`: The product or system name
- `$DATASET_TYPE`: Type of data (e.g., customer feedback, transactions, user profiles)
- `$ROWS`: Number of rows to generate (default: 100)
- `$COLUMNS`: Specific columns or fields to include
- `$FORMAT`: Output format (CSV, JSON, SQL, Python script)
- `$CONSTRAINTS`: Additional constraints or business rules

## Step-by-Step Process

1. **Identify dataset type** - Understand the data domain
2. **Define column specifications** - Names, data types, and value ranges
3. **Determine row count** - How many sample records needed
4. **Select output format** - CSV, JSON, SQL INSERT, or Python script
5. **Apply realistic patterns** - Ensure data looks authentic and valid
6. **Add business constraints** - Respect business logic and relationships
7. **Generate or script data** - Create executable output
8. **Validate output** - Ensure data quality and completeness

## Template: Python Script Output

```python
import csv
import json
from datetime import datetime, timedelta
import random

# Configuration
ROWS = $ROWS
FILENAME = "$DATASET_TYPE.csv"

# Column definitions with realistic value generators
columns = {
    "id": "auto-increment",
    "name": "first_last_name",
    "email": "email",
    "created_at": "timestamp",
    # Add more columns...
}

def generate_dataset():
    """Generate realistic dummy dataset"""
    data = []
    for i in range(1, ROWS + 1):
        record = {
            "id": f"U{i:06d}",
            # Generate values based on column definitions
        }
        data.append(record)
    return data

def save_as_csv(data, filename):
    """Save dataset as CSV"""
    with open(filename, 'w', newline='') as f:
        writer = csv.DictWriter(f, fieldnames=data[0].keys())
        writer.writeheader()
        writer.writerows(data)

if __name__ == "__main__":
    dataset = generate_dataset()
    save_as_csv(dataset, FILENAME)
    print(f"Generated {len(dataset)} records in {FILENAME}")
```

## Example Dataset Specification

**Dataset Type:** Customer Feedback

**Columns:**
- feedback_id (auto-increment, U001, U002...)
- customer_name (realistic names)
- email (valid email format)
- feedback_date (dates last 90 days)
- rating (1-5 stars)
- category (Bug, Feature Request, Complaint, Praise)
- text (realistic feedback)
- product (electronics, clothing, home)

**Constraints:**
- Ratings skewed: 40% 5-star, 30% 4-star, 20% 3-star, 10% 1-2 star
- Bug category only with ratings 1-3
- Feature requests only with ratings 3-5
- Email domains realistic (gmail, yahoo, company.com)

## Output Deliverables

- Ready-to-execute Python script OR direct data file
- CSV file with proper headers and formatting
- JSON file with valid structure and types
- SQL INSERT statements for database population
- Data validation and constraint compliance
- Realistic, business-appropriate values
- Documentation of data generation logic
- Quick-start instructions for using the dataset

## Output Formats

**CSV:** Flat tabular format, easy to import into spreadsheets and databases

**JSON:** Nested structure, ideal for APIs and NoSQL databases

**SQL:** INSERT statements, directly executable on relational databases

**Python Script:** Executable generator for custom or large datasets

所有檔案

1 個檔案

安裝 dummy-dataset

請下載並將技能檔案解壓縮至您的 .claude/skills/ 目錄中。

下載 ZIP

複製儲存庫並將技能檔案複製到您的專案中。

git clone https://github.com/phuryn/pm-skills/tree/main/pm-execution/skills/dummy-dataset # Copy SKILL.md to your .claude/skills/ directory

複製 複製
快速設定: 將技能資料夾複製到 .claude/skills/ Claude 會自動偵測並使用該技能
儲存庫 phuryn/pm-skills

相關技能

microservices-patterns
更新時間 2026-06-29
jpa-patterns
更新時間 2026-06-30
fabric-lakehouse
更新時間 2026-06-30
prisma-expert
更新時間 2026-06-29
OR