nemo-data-designer-plugin
NVIDIA/skills
使用 Data Designer 函式庫建立合成資料集與資料生成管線。
...展開全部開始之前
請勿先自行探索工作區。此工作流程中的「學習」步驟已提供您所需的一切資訊。
目標
使用 Data Designer 函式庫建立一個符合以下描述的合成資料集:
$ARGUMENTS
工作流程
若使用者暗示不希望回答問題(例如說出「請展現主見」、「由你決定」、「做出合理假設」、「直接建立就好」、「給我驚喜」等),請使用「自動駕駛」模式;否則,請使用「互動」模式(預設)。
僅讀取與所選模式相符的工作流程檔案,並依照該檔案執行:
- 互動式 → 讀取
workflows/interactive.md - 自動駕駛 → 讀取
workflows/autopilot.md
規則
- 預設情況下,請保留輸出中的所有欄位。唯一可刪除欄位的例外情況為:(1) 使用者明確要求,或 (2) 該欄位純粹是為了推導其他欄位而存在的輔助欄位(例如:用於擷取姓名、城市等資訊的抽樣人員物件)。如有疑慮,請保留該欄位。
- 請勿建議或詢問關於種子資料集的事宜。僅在使用者明確提供種子資料,或要求基於現有記錄建構時,才使用種子資料集。使用種子資料集時,請參閱
references/seed-datasets.md. - 當資料集需要個人資料(姓名、人口統計資料、地址)時,請參閱
references/person-sampling.md. - 若已存在與資料集描述相符的資料集腳本,請詢問使用者是要編輯現有腳本還是建立新的腳本。
- 關於此 NeMo Platform 外掛程式專屬的指令與情境(例如:從 IGW 提供者獲取模型配置,或腳本中的
ModelConfig、安裝或發佈 Nemotron Personas 地區設定、平台端資源指標),請參閱references/nemo-platform-plugin-additions.md.
使用提示與常見陷阱
- 取樣器和驗證欄位均需指定資料型別與參數。例如:
sampler_type="category"搭配params=dd.CategorySamplerParams(...). - Jinja2 模板位於
prompt,system_prompt,以及expr欄位:使用{{ column_name }},以及嵌套欄位{{ column_name.field }}. **SamplerColumnConfig:** 接受params,而非sampler_params.- LLM 評審分數存取:
LLMJudgeColumnConfig會產生一個嵌套字典,其中每個分數名稱對應至{reasoning: str, score: int}。若要取得數值分數,請使用.score屬性。例如,對於名為quality,且評分名為correctness,請使用{{ quality.correctness.score }}。若使用{{ quality.correctness }}會傳回完整的字典,而非數值分數。
疑難排解
**nemo data-designer找不到 CLI:** 告知使用者nemo data-designer未安裝於此環境中(需 Python >= 3.11)。詢問使用者是否希望您建立虛擬環境並進行安裝,或是他們寧願自行處理。未經使用者許可,請勿安裝任何軟體。- 預覽期間發生網路錯誤:沙盒環境可能正在封鎖外發請求。請徵求使用者同意,在停用沙盒的情況下重新嘗試執行該指令。僅在萬不得已的情況下(即在沙盒外重新嘗試仍失敗時),才告知使用者自行執行該指令。
輸出範本
在當前目錄中建立一個 Python 檔案,其中包含一個 load_config_builder() 函式,該函式會傳回一個 DataDesignerConfigBuilder。為檔案命名時請具描述性(例如: customer_reviews.py)。請使用 PEP 723 規範的內嵌元資料來標示依賴項。
# /// script
# dependencies = [
# "data-designer", # always required
# "pydantic", # only if this script imports from pydantic
# # add additional dependencies here
# ]
# ///
import data_designer.config as dd
from pydantic import BaseModel, Field
# Use Pydantic models when the output needs to conform to a specific schema
class MyStructuredOutput(BaseModel):
field_one: str = Field(description="...")
field_two: int = Field(description="...")
# Use custom generators when built-in column types aren't enough
@dd.custom_column_generator(
required_columns=["col_a"],
side_effect_columns=["extra_col"],
)
def generator_function(row: dict) -> dict:
# add custom logic here that depends on "col_a" and update row in place
row["name_in_custom_column_config"] = "custom value"
row["extra_col"] = "extra value"
return row
def load_config_builder() -> dd.DataDesignerConfigBuilder:
config_builder = dd.DataDesignerConfigBuilder(
# Declaring model configs programmatically here is the portable path:
# it works for both local `run` and cluster `submit`, while the local
# YAML registry alternative only works for `run`. The provider below
# is a common default created during `nemo setup` — confirm it (or
# discover others) with `nemo inference providers list`. See
# references/nemo-platform-plugin-additions.md for the local-YAML alternative.
model_configs=[
dd.ModelConfig(
alias="text",
model="...",
provider="default/nvidia-build",
inference_parameters=dd.ChatCompletionInferenceParams(),
),
],
)
# Seed dataset (only if the user explicitly mentions a seed dataset path)
# config_builder.with_seed_dataset(dd.LocalFileSeedSource(path="path/to/seed.parquet"))
# config_builder.add_column(...)
# config_builder.add_processor(...)
return config_builder
僅在任務有需求時,才包含 Pydantic 模型、自訂產生器、種子資料集及額外依賴項。若資料集使用 LLM 欄位,請優先在腳本中宣告 model_configs ——若資料集使用 LLM 欄位,請在腳本中宣告此依賴項,以確保配置在本地端 run 與叢集 submit,而本地的 YAML 註冊表替代方案僅適用於 run.
---
name: nemo-data-designer-plugin
description: Build synthetic datasets and data generation pipelines using the Data Designer library.
license: Apache-2.0
---
# Before You Start
Do not explore the workspace first. The workflow's Learn step gives you everything you need.
# Goal
Build a synthetic dataset using the Data Designer library that matches this description:
$ARGUMENTS
# Workflow
Use **Autopilot** mode if the user implies they don't want to answer questions — e.g., they say something like "be opinionated", "you decide", "make reasonable assumptions", "just build it", "surprise me", etc. Otherwise, use **Interactive** mode (default).
Read **only** the workflow file that matches the selected mode, then follow it:
- **Interactive** → read `workflows/interactive.md`
- **Autopilot** → read `workflows/autopilot.md`
# Rules
- Keep all columns in the output by default. The only exceptions for dropping a column are: (1) the user explicitly asks, or (2) it is a helper column that exists solely to derive other columns (e.g., a sampled person object used to extract name, city, etc.). When in doubt, keep the column.
- Do not suggest or ask about seed datasets. Only use one when the user explicitly provides seed data or asks to build from existing records. When using a seed, read `references/seed-datasets.md`.
- When the dataset requires person data (names, demographics, addresses), read `references/person-sampling.md`.
- If a dataset script that matches the dataset description already exists, ask the user whether to edit it or create a new one.
- For commands and context specific to this NeMo Platform plugin (e.g., sourcing model configs from IGW providers or in-script `ModelConfig`s, installing or publishing Nemotron Personas locales, platform-side resource pointers), read `references/nemo-platform-plugin-additions.md`.
# Usage Tips and Common Pitfalls
- **Sampler and validation columns need both a type and params.** E.g., `sampler_type="category"` with `params=dd.CategorySamplerParams(...)`.
- **Jinja2 templates** in `prompt`, `system_prompt`, and `expr` fields: reference columns with `{{ column_name }}`, nested fields with `{{ column_name.field }}`.
- `**SamplerColumnConfig`:** Takes `params`, not `sampler_params`.
- **LLM judge score access:** `LLMJudgeColumnConfig` produces a nested dict where each score name maps to `{reasoning: str, score: int}`. To get the numeric score, use the `.score` attribute. For example, for a judge column named `quality` with a score named `correctness`, use `{{ quality.correctness.score }}`. Using `{{ quality.correctness }}` returns the full dict, not the numeric score.
# Troubleshooting
- `**nemo data-designer` CLI not found:** Tell the user that `nemo data-designer` is not installed in this environment (requires Python >= 3.11). Ask if they would like you to create a virtual environment and install it, or if they prefer to do it themselves. Do not install anything without the user's permission.
- **Network errors during preview:** A sandbox environment may be blocking outbound requests. Ask the user for permission to retry the command with the sandbox disabled. Only as a last resort, if retrying outside the sandbox also fails, tell the user to run the command themselves.
# Output Template
Write a Python file to the current directory with a `load_config_builder()` function returning a `DataDesignerConfigBuilder`. Name the file descriptively (e.g., `customer_reviews.py`). Use PEP 723 inline metadata for dependencies.
```python
# /// script
# dependencies = [
# "data-designer", # always required
# "pydantic", # only if this script imports from pydantic
# # add additional dependencies here
# ]
# ///
import data_designer.config as dd
from pydantic import BaseModel, Field
# Use Pydantic models when the output needs to conform to a specific schema
class MyStructuredOutput(BaseModel):
field_one: str = Field(description="...")
field_two: int = Field(description="...")
# Use custom generators when built-in column types aren't enough
@dd.custom_column_generator(
required_columns=["col_a"],
side_effect_columns=["extra_col"],
)
def generator_function(row: dict) -> dict:
# add custom logic here that depends on "col_a" and update row in place
row["name_in_custom_column_config"] = "custom value"
row["extra_col"] = "extra value"
return row
def load_config_builder() -> dd.DataDesignerConfigBuilder:
config_builder = dd.DataDesignerConfigBuilder(
# Declaring model configs programmatically here is the portable path:
# it works for both local `run` and cluster `submit`, while the local
# YAML registry alternative only works for `run`. The provider below
# is a common default created during `nemo setup` — confirm it (or
# discover others) with `nemo inference providers list`. See
# references/nemo-platform-plugin-additions.md for the local-YAML alternative.
model_configs=[
dd.ModelConfig(
alias="text",
model="...",
provider="default/nvidia-build",
inference_parameters=dd.ChatCompletionInferenceParams(),
),
],
)
# Seed dataset (only if the user explicitly mentions a seed dataset path)
# config_builder.with_seed_dataset(dd.LocalFileSeedSource(path="path/to/seed.parquet"))
# config_builder.add_column(...)
# config_builder.add_processor(...)
return config_builder
```
Only include Pydantic models, custom generators, seed datasets, and extra dependencies when the task requires them. Prefer including `model_configs` when the dataset uses LLM columns — declaring it in the script keeps the config portable between local `run` and cluster `submit`, while the local YAML registry alternative only works for `run`.
所有檔案
12 個檔案安裝 nemo-data-designer-plugin
請下載並將技能檔案解壓縮至您的 .claude/skills/ 目錄中。
下載 ZIP複製儲存庫並將技能檔案複製到您的專案中。
git clone https://github.com/NVIDIA/skills/tree/main/skills/nemo-data-designer-plugin # Copy SKILL.md to your .claude/skills/ directory
複製





首頁
