dummy-dataset
phuryn/pm-skills
사용자 정의 가능한 열, 제약 조건 및 출력 형식(CSV, JSON, SQL, Python 스크립트)을 사용하여 테스트용 사실적인 모의 데이터 세트를 생성합니다.
...모든 것을 확장하십시오모의 데이터셋 생성
사용자 정의 가능한 열, 제약 조건 및 출력 형식(CSV, JSON, SQL, Python 스크립트)을 사용하여 테스트용 사실적인 더미 데이터셋을 생성합니다. 즉시 사용할 수 있는 실행 가능한 스크립트나 데이터 파일을 생성합니다.
사용 시나리오: 테스트 데이터 생성, 샘플 데이터셋 생성, 개발을 위한 사실적인 모의 데이터 구축, 또는 테스트 환경 채우기.
인수:
$PRODUCT: 제품 또는 시스템 이름$DATASET_TYPE: 데이터 유형(예: 고객 피드백, 거래 내역, 사용자 프로필)$ROWS: 생성할 행 수 (기본값: 100)$COLUMNS: 포함할 특정 열 또는 필드$FORMAT: 출력 형식 (CSV, JSON, SQL, Python 스크립트)$CONSTRAINTS: 추가 제약 조건 또는 비즈니스 규칙
단계별 절차
- 데이터셋 유형 파악 - 데이터 도메인 이해
- 열 사양 정의 - 이름, 데이터 유형 및 값 범위
- 행 수 결정 - 필요한 샘플 레코드 수
- 출력 형식 선택 - CSV, JSON, SQL INSERT 또는 Python 스크립트
- 현실적인 패턴 적용 - 데이터가 사실적이고 유효해 보이도록 보장
- 비즈니스 제약 조건 추가 - 비즈니스 로직 및 관계 준수
- 데이터 생성 또는 스크립트 작성 - 실행 가능한 출력 생성
- 출력 검증 - 데이터 품질 및 완전성 확인
템플릿: Python 스크립트 출력
import csv
import json
from datetime import datetime, timedelta
import random
# Configuration
ROWS = $ROWS
FILENAME = "$DATASET_TYPE.csv"
# Column definitions with realistic value generators
columns = {
"id": "auto-increment",
"name": "first_last_name",
"email": "email",
"created_at": "timestamp",
# Add more columns...
}
def generate_dataset():
"""Generate realistic dummy dataset"""
data = []
for i in range(1, ROWS + 1):
record = {
"id": f"U{i:06d}",
# Generate values based on column definitions
}
data.append(record)
return data
def save_as_csv(data, filename):
"""Save dataset as CSV"""
with open(filename, 'w', newline='') as f:
writer = csv.DictWriter(f, fieldnames=data[0].keys())
writer.writeheader()
writer.writerows(data)
if __name__ == "__main__":
dataset = generate_dataset()
save_as_csv(dataset, FILENAME)
print(f"Generated {len(dataset)} records in {FILENAME}")
데이터셋 사양 예시
데이터셋 유형: 고객 피드백
열:
- feedback_id (자동 증가, U001, U002...)
- customer_name (실제 이름)
- email (유효한 이메일 형식)
- feedback_date (최근 90일 이내의 날짜)
- rating (1~5개 별)
- category (버그, 기능 요청, 불만, 칭찬)
- text (현실적인 피드백)
- 제품 (전자제품, 의류, 가정용품)
제약 조건:
- 평점 분포: 5성 40%, 4성 30%, 3성 20%, 1~2성 10%
- 1~3성 평점만 있는 오류 카테고리
- 기능 요청은 3~5점 평가만 포함
- 이메일 도메인은 현실적이어야 함(gmail, yahoo, company.com)
출력 결과물
- 바로 실행 가능한 Python 스크립트 또는 직접 데이터 파일
- 적절한 헤더와 서식이 적용된 CSV 파일
- 유효한 구조와 유형을 갖춘 JSON 파일
- 데이터베이스 채우기를 위한 SQL INSERT 문
- 데이터 유효성 검사 및 제약 조건 준수
- 현실적이고 비즈니스에 적합한 값
- 데이터 생성 로직에 대한 문서화
- 데이터셋 사용을 위한 빠른 시작 안내
출력 형식
CSV: 평면 표 형식, 스프레드시트 및 데이터베이스로 쉽게 가져올 수 있음
JSON: 중첩 구조로, API 및 NoSQL 데이터베이스에 이상적
SQL: INSERT 문으로, 관계형 데이터베이스에서 직접 실행 가능
Python 스크립트: 사용자 정의 또는 대용량 데이터셋을 위한 실행 가능한 생성기
---
name: dummy-dataset
description: Generate realistic dummy datasets for testing with customizable columns, constraints, and output formats (CSV, JSON, SQL, Python script).
---
# Dummy Dataset Generation
Generate realistic dummy datasets for testing with customizable columns, constraints, and output formats (CSV, JSON, SQL, Python script). Creates executable scripts or direct data files for immediate use.
**Use when:** Creating test data, generating sample datasets, building realistic mock data for development, or populating test environments.
**Arguments:**
- `$PRODUCT`: The product or system name
- `$DATASET_TYPE`: Type of data (e.g., customer feedback, transactions, user profiles)
- `$ROWS`: Number of rows to generate (default: 100)
- `$COLUMNS`: Specific columns or fields to include
- `$FORMAT`: Output format (CSV, JSON, SQL, Python script)
- `$CONSTRAINTS`: Additional constraints or business rules
## Step-by-Step Process
1. **Identify dataset type** - Understand the data domain
2. **Define column specifications** - Names, data types, and value ranges
3. **Determine row count** - How many sample records needed
4. **Select output format** - CSV, JSON, SQL INSERT, or Python script
5. **Apply realistic patterns** - Ensure data looks authentic and valid
6. **Add business constraints** - Respect business logic and relationships
7. **Generate or script data** - Create executable output
8. **Validate output** - Ensure data quality and completeness
## Template: Python Script Output
```python
import csv
import json
from datetime import datetime, timedelta
import random
# Configuration
ROWS = $ROWS
FILENAME = "$DATASET_TYPE.csv"
# Column definitions with realistic value generators
columns = {
"id": "auto-increment",
"name": "first_last_name",
"email": "email",
"created_at": "timestamp",
# Add more columns...
}
def generate_dataset():
"""Generate realistic dummy dataset"""
data = []
for i in range(1, ROWS + 1):
record = {
"id": f"U{i:06d}",
# Generate values based on column definitions
}
data.append(record)
return data
def save_as_csv(data, filename):
"""Save dataset as CSV"""
with open(filename, 'w', newline='') as f:
writer = csv.DictWriter(f, fieldnames=data[0].keys())
writer.writeheader()
writer.writerows(data)
if __name__ == "__main__":
dataset = generate_dataset()
save_as_csv(dataset, FILENAME)
print(f"Generated {len(dataset)} records in {FILENAME}")
```
## Example Dataset Specification
**Dataset Type:** Customer Feedback
**Columns:**
- feedback_id (auto-increment, U001, U002...)
- customer_name (realistic names)
- email (valid email format)
- feedback_date (dates last 90 days)
- rating (1-5 stars)
- category (Bug, Feature Request, Complaint, Praise)
- text (realistic feedback)
- product (electronics, clothing, home)
**Constraints:**
- Ratings skewed: 40% 5-star, 30% 4-star, 20% 3-star, 10% 1-2 star
- Bug category only with ratings 1-3
- Feature requests only with ratings 3-5
- Email domains realistic (gmail, yahoo, company.com)
## Output Deliverables
- Ready-to-execute Python script OR direct data file
- CSV file with proper headers and formatting
- JSON file with valid structure and types
- SQL INSERT statements for database population
- Data validation and constraint compliance
- Realistic, business-appropriate values
- Documentation of data generation logic
- Quick-start instructions for using the dataset
## Output Formats
**CSV:** Flat tabular format, easy to import into spreadsheets and databases
**JSON:** Nested structure, ideal for APIs and NoSQL databases
**SQL:** INSERT statements, directly executable on relational databases
**Python Script:** Executable generator for custom or large datasets
모든 파일
1개 파일dummy-dataset 설치
스킬 파일을 다운로드하여 .claude/skills/ 디렉터리에 압축을 풀어주세요.
ZIP 다운로드저장소를 클론하고 스킬 파일을 프로젝트에 복사하세요.
git clone https://github.com/phuryn/pm-skills/tree/main/pm-execution/skills/dummy-dataset # Copy SKILL.md to your .claude/skills/ directory
복사





집
