옵션
집집 Skill 데이터베이스 관리 nemo-data-designer-plugin

nemo-data-designer-plugin

NVIDIA/skills NVIDIA/skills

Data Designer 라이브러리를 사용하여 합성 데이터셋과 데이터 생성 파이프라인을 구축합니다.

...모든 것을 확장하십시오
1
업데이트 된 시간 2026년 9월 27일

시작하기 전에

먼저 작업 공간을 둘러보지 마세요. 워크플로우의 ‘배움’ 단계에서 필요한 모든 정보를 확인할 수 있습니다.

목표

Data Designer 라이브러리를 사용하여 다음 설명과 일치하는 합성 데이터셋을 구축하십시오:

$ARGUMENTS

워크플로우

사용자가 질문에 답하고 싶지 않다는 뜻을 내비치는 경우(예: “주관적으로 처리해 주세요”, “당신이 결정하세요”, “합리적인 가정을 해 주세요”, “그냥 만들어 주세요”, “놀라게 해 주세요” 등)에는 오토파일럿 모드를 사용하십시오. 그렇지 않은 경우 대화형 모드(기본값)를 사용하십시오.

선택한 모드에 해당하는 워크플로 파일만 읽은 다음, 그 지침을 따르십시오:

  • 대화형 → 읽기 workflows/interactive.md
  • 자동 모드 → 읽기 workflows/autopilot.md

규칙

  • 기본적으로 출력에 모든 열을 포함하십시오. 열을 제외할 수 있는 유일한 예외는 (1) 사용자가 명시적으로 요청한 경우, 또는 (2) 다른 열을 도출하기 위해서만 존재하는 보조 열인 경우(예: 이름, 도시 등을 추출하는 데 사용되는 샘플링된 사람 객체)입니다. 확실하지 않은 경우에는 해당 열을 유지하십시오.
  • 시드 데이터셋에 대해 제안하거나 묻지 마십시오. 사용자가 명시적으로 시드 데이터를 제공하거나 기존 레코드를 기반으로 구축해 달라고 요청할 때만 사용하십시오. 시드를 사용할 때는 다음을 참조하십시오. references/seed-datasets.md.
  • 데이터셋에 개인 정보(이름, 인구통계, 주소)가 필요한 경우, 다음을 참조하십시오. references/person-sampling.md.
  • 데이터셋 설명과 일치하는 데이터셋 스크립트가 이미 존재하는 경우, 사용자에게 기존 스크립트를 수정할지 아니면 새 스크립트를 생성할지 물어보십시오.
  • 이 NeMo 플랫폼 플러그인에 특화된 명령어 및 컨텍스트(예: IGW 공급자로부터 모델 구성을 가져오거나 스크립트 내 ModelConfig, Nemotron Personas 로케일 설치 또는 게시, 플랫폼 측 리소스 포인터 등)에 대해서는 references/nemo-platform-plugin-additions.md.

사용 요령 및 일반적인 실수

  • 샘플러 및 유효성 검사 열에는 유형과 매개변수가 모두 필요합니다. 예: sampler_type="category" Jinja2 템플릿을 사용하는 경우 params=dd.CategorySamplerParams(...).
  • Jinja2 템플릿을 prompt, system_prompt, 그리고 expr 필드: {{ column_name }}, 중첩된 필드인 {{ column_name.field }}.
  • **SamplerColumnConfig:** 다음을 params, sampler_params.
  • LLM 심사위원 점수 참조: LLMJudgeColumnConfig 각 점수 이름이 {reasoning: str, score: int}. 숫자 점수를 얻으려면 .score 속성을 사용합니다. 예를 들어, quality 인 채점자 열의 점수 이름이 correctness인 경우, {{ quality.correctness.score }}를 사용합니다. {{ quality.correctness }} 를 사용하면 숫자 점수가 아닌 전체 딕셔너리가 반환됩니다.

문제 해결

  • **nemo data-designer CLI를 찾을 수 없음:** 사용자에게 nemo data-designer 이 환경에 설치되어 있지 않다고 알립니다(Python 3.11 이상 필요). 가상 환경을 생성하여 설치해 드릴지, 아니면 사용자가 직접 설치할지 물어보세요. 사용자의 허락 없이 아무것도 설치하지 마십시오.
  • 미리보기 중 네트워크 오류: 샌드박스 환경이 외부로 나가는 요청을 차단하고 있을 수 있습니다. 샌드박스를 비활성화한 상태에서 명령을 다시 시도할 수 있도록 사용자에게 허락을 구하십시오. 샌드박스 외부에서 재시도해도 실패하는 경우, 최후의 수단으로만 사용자에게 직접 명령을 실행하도록 안내하십시오.

출력 템플릿

현재 디렉터리에 다음 함수를 반환하는 Python 파일을 작성하십시오. load_config_builder() 를 반환하는 함수를 포함하도록 하십시오 DataDesignerConfigBuilder를 반환하는 함수를 포함하십시오. 파일 이름은 내용을 명확히 나타내도록 지정하십시오(예: customer_reviews.py). 종속성에 대해서는 PEP 723의 인라인 메타데이터를 사용하십시오.

# /// script
# dependencies = [
#   "data-designer", # always required
#   "pydantic", # only if this script imports from pydantic
#   # add additional dependencies here
# ]
# ///
import data_designer.config as dd
from pydantic import BaseModel, Field


# Use Pydantic models when the output needs to conform to a specific schema
class MyStructuredOutput(BaseModel):
    field_one: str = Field(description="...")
    field_two: int = Field(description="...")


# Use custom generators when built-in column types aren't enough
@dd.custom_column_generator(
    required_columns=["col_a"],
    side_effect_columns=["extra_col"],
)
def generator_function(row: dict) -> dict:
    # add custom logic here that depends on "col_a" and update row in place
    row["name_in_custom_column_config"] = "custom value"
    row["extra_col"] = "extra value"
    return row


def load_config_builder() -> dd.DataDesignerConfigBuilder:
    config_builder = dd.DataDesignerConfigBuilder(
        # Declaring model configs programmatically here is the portable path:
        # it works for both local `run` and cluster `submit`, while the local
        # YAML registry alternative only works for `run`. The provider below
        # is a common default created during `nemo setup` — confirm it (or
        # discover others) with `nemo inference providers list`. See
        # references/nemo-platform-plugin-additions.md for the local-YAML alternative.
        model_configs=[
            dd.ModelConfig(
                alias="text",
                model="...",
                provider="default/nvidia-build",
                inference_parameters=dd.ChatCompletionInferenceParams(),
            ),
        ],
    )

    # Seed dataset (only if the user explicitly mentions a seed dataset path)
    # config_builder.with_seed_dataset(dd.LocalFileSeedSource(path="path/to/seed.parquet"))

    # config_builder.add_column(...)
    # config_builder.add_processor(...)

    return config_builder

Pydantic 모델, 사용자 정의 생성기, 시드 데이터셋 및 추가 종속성은 작업에 필요한 경우에만 포함하십시오. model_configs 를 포함하는 것을 선호하십시오 — 스크립트에서 이를 선언하면 로컬 run 클러스터 submit간 이식성을 유지해 주지만, 로컬 YAML 레지스트리 방식은 run.

GitHub에서 보기
---
name: nemo-data-designer-plugin
description: Build synthetic datasets and data generation pipelines using the Data Designer library.
license: Apache-2.0
---

# Before You Start

Do not explore the workspace first. The workflow's Learn step gives you everything you need.

# Goal

Build a synthetic dataset using the Data Designer library that matches this description:

$ARGUMENTS

# Workflow

Use **Autopilot** mode if the user implies they don't want to answer questions — e.g., they say something like "be opinionated", "you decide", "make reasonable assumptions", "just build it", "surprise me", etc. Otherwise, use **Interactive** mode (default).

Read **only** the workflow file that matches the selected mode, then follow it:

- **Interactive** → read `workflows/interactive.md`
- **Autopilot** → read `workflows/autopilot.md`

# Rules

- Keep all columns in the output by default. The only exceptions for dropping a column are: (1) the user explicitly asks, or (2) it is a helper column that exists solely to derive other columns (e.g., a sampled person object used to extract name, city, etc.). When in doubt, keep the column.
- Do not suggest or ask about seed datasets. Only use one when the user explicitly provides seed data or asks to build from existing records. When using a seed, read `references/seed-datasets.md`.
- When the dataset requires person data (names, demographics, addresses), read `references/person-sampling.md`.
- If a dataset script that matches the dataset description already exists, ask the user whether to edit it or create a new one.
- For commands and context specific to this NeMo Platform plugin (e.g., sourcing model configs from IGW providers or in-script `ModelConfig`s, installing or publishing Nemotron Personas locales, platform-side resource pointers), read `references/nemo-platform-plugin-additions.md`.

# Usage Tips and Common Pitfalls

- **Sampler and validation columns need both a type and params.** E.g., `sampler_type="category"` with `params=dd.CategorySamplerParams(...)`.
- **Jinja2 templates** in `prompt`, `system_prompt`, and `expr` fields: reference columns with `{{ column_name }}`, nested fields with `{{ column_name.field }}`.
- `**SamplerColumnConfig`:** Takes `params`, not `sampler_params`.
- **LLM judge score access:** `LLMJudgeColumnConfig` produces a nested dict where each score name maps to `{reasoning: str, score: int}`. To get the numeric score, use the `.score` attribute. For example, for a judge column named `quality` with a score named `correctness`, use `{{ quality.correctness.score }}`. Using `{{ quality.correctness }}` returns the full dict, not the numeric score.

# Troubleshooting

- `**nemo data-designer` CLI not found:** Tell the user that `nemo data-designer` is not installed in this environment (requires Python >= 3.11). Ask if they would like you to create a virtual environment and install it, or if they prefer to do it themselves. Do not install anything without the user's permission.
- **Network errors during preview:** A sandbox environment may be blocking outbound requests. Ask the user for permission to retry the command with the sandbox disabled. Only as a last resort, if retrying outside the sandbox also fails, tell the user to run the command themselves.

# Output Template

Write a Python file to the current directory with a `load_config_builder()` function returning a `DataDesignerConfigBuilder`. Name the file descriptively (e.g., `customer_reviews.py`). Use PEP 723 inline metadata for dependencies.

```python
# /// script
# dependencies = [
#   "data-designer", # always required
#   "pydantic", # only if this script imports from pydantic
#   # add additional dependencies here
# ]
# ///
import data_designer.config as dd
from pydantic import BaseModel, Field


# Use Pydantic models when the output needs to conform to a specific schema
class MyStructuredOutput(BaseModel):
    field_one: str = Field(description="...")
    field_two: int = Field(description="...")


# Use custom generators when built-in column types aren't enough
@dd.custom_column_generator(
    required_columns=["col_a"],
    side_effect_columns=["extra_col"],
)
def generator_function(row: dict) -> dict:
    # add custom logic here that depends on "col_a" and update row in place
    row["name_in_custom_column_config"] = "custom value"
    row["extra_col"] = "extra value"
    return row


def load_config_builder() -> dd.DataDesignerConfigBuilder:
    config_builder = dd.DataDesignerConfigBuilder(
        # Declaring model configs programmatically here is the portable path:
        # it works for both local `run` and cluster `submit`, while the local
        # YAML registry alternative only works for `run`. The provider below
        # is a common default created during `nemo setup` — confirm it (or
        # discover others) with `nemo inference providers list`. See
        # references/nemo-platform-plugin-additions.md for the local-YAML alternative.
        model_configs=[
            dd.ModelConfig(
                alias="text",
                model="...",
                provider="default/nvidia-build",
                inference_parameters=dd.ChatCompletionInferenceParams(),
            ),
        ],
    )

    # Seed dataset (only if the user explicitly mentions a seed dataset path)
    # config_builder.with_seed_dataset(dd.LocalFileSeedSource(path="path/to/seed.parquet"))

    # config_builder.add_column(...)
    # config_builder.add_processor(...)

    return config_builder
```

Only include Pydantic models, custom generators, seed datasets, and extra dependencies when the task requires them. Prefer including `model_configs` when the dataset uses LLM columns — declaring it in the script keeps the config portable between local `run` and cluster `submit`, while the local YAML registry alternative only works for `run`.

nemo-data-designer-plugin 설치

스킬 파일을 다운로드하여 .claude/skills/ 디렉터리에 압축을 풀어주세요.

ZIP 다운로드

저장소를 클론하고 스킬 파일을 프로젝트에 복사하세요.

git clone https://github.com/NVIDIA/skills/tree/main/skills/nemo-data-designer-plugin # Copy SKILL.md to your .claude/skills/ directory

복사 복사
빠른 설정: 스킬 폴더를 .claude/skills/로 복사하세요. Claude가 해당 스킬을 자동으로 감지하여 사용합니다.
저장소 NVIDIA/skills

관련 스킬

microservices-patterns
업데이트 된 시간 2026년 6월 29일
jpa-patterns
업데이트 된 시간 2026년 6월 30일
fabric-lakehouse
업데이트 된 시간 2026년 6월 30일
prisma-expert
업데이트 된 시간 2026년 6월 29일
OR