nemo-data-designer-plugin
NVIDIA/skills
Data Designer ライブラリを使用して、合成データセットとデータ生成パイプラインを構築します。
...すべて拡張します始める前に
最初にワークスペースを探索しないでください。ワークフローの「学習」ステップには、必要な情報がすべて含まれています。
目標
Data Designer ライブラリを使用して、以下の説明に合致する合成データセットを構築してください:
$ARGUMENTS
ワークフロー
ユーザーが質問に答えたくない意向を示している場合(例:「独自の判断で」「あなたが決めて」「妥当な仮定を立てて」「とにかく構築して」「驚かせて」などと言った場合)、オートパイロットモードを使用してください。それ以外の場合は、対話モード(デフォルト)を使用してください。
選択されたモードに対応するワークフローファイルのみを読み込み、それに従って処理を行う:
- インタラクティブ → 読み込み
workflows/interactive.md - オートパイロット → 読み込み
workflows/autopilot.md
ルール
- デフォルトでは、出力にすべての列を含めます。列を削除する唯一の例外は、(1) ユーザーが明示的に要求した場合、または (2) 他の列を導出するためだけに存在する補助列(例:名前や都市などを抽出するために使用されるサンプリングされた人物オブジェクト)である場合です。判断に迷った場合は、その列を残してください。
- シードデータセットの提案や確認は行わないでください。ユーザーが明示的にシードデータを提供した場合、または既存のレコードから構築するよう要求した場合にのみ使用してください。シードを使用する場合は、以下を参照してください
references/seed-datasets.md. - データセットに個人データ(氏名、人口統計情報、住所)が必要な場合は、以下を参照してください
references/person-sampling.md. - データセットの説明と一致するデータセットスクリプトがすでに存在する場合は、それを編集するか、新しいスクリプトを作成するかをユーザーに確認してください。
- このNeMo Platformプラグインに固有のコマンドやコンテキスト(例:IGWプロバイダーからのモデル設定の取得、スクリプト内の
ModelConfig、Nemotron Personasのロケールのインストールや公開、プラットフォーム側のリソースポインタなど)については、references/nemo-platform-plugin-additions.md.
「使用上のヒントとよくある落とし穴」を参照してください
- サンプラーおよび検証カラムには、型とパラメータの両方が必要です。例:
sampler_type="category"Jinja2 テンプレートを使用する場合、params=dd.CategorySamplerParams(...). - Jinja2テンプレートが
prompt,system_prompt、およびexprフィールド:参照カラムには{{ column_name }}、ネストされたフィールドには{{ column_name.field }}. **SamplerColumnConfig:** 受け取るparamsを使用し、sampler_params.- LLMによる採点スコアへのアクセス:
LLMJudgeColumnConfig各スコア名が{reasoning: str, score: int}にマッピングされます。数値のスコアを取得するには、.score属性を使用します。例えば、qualityというスコアがcorrectnessというスコア名を持つ場合、{{ quality.correctness.score }}を使用します。{{ quality.correctness }}を使用すると、数値スコアではなく完全な辞書が返されます。
トラブルシューティング
**nemo data-designerCLIが見つかりません:** ユーザーに、nemo data-designerがインストールされていないことを伝えてください(Python 3.11 以上が必要です)。仮想環境を作成してインストールするか、ユーザー自身でインストールするかを確認してください。ユーザーの許可なしに何もインストールしないでください。- プレビュー中のネットワークエラー:サンドボックス環境が外部へのリクエストをブロックしている可能性があります。サンドボックスを無効にしてコマンドを再実行する許可をユーザーに求めます。サンドボックス外での再実行も失敗した場合に限り、最後の手段として、ユーザー自身にコマンドを実行するよう伝えてください。
出力テンプレート
現在のディレクトリに、 load_config_builder() を返す関数を含むPythonファイルを現在のディレクトリに書き出します DataDesignerConfigBuilderを返す関数を含むPythonファイルを現在のディレクトリに作成してください。ファイル名は内容を明確に表すものにしてください(例: customer_reviews.py)。依存関係については、PEP 723に準拠したインラインメタデータを使用してください。
# /// script
# dependencies = [
# "data-designer", # always required
# "pydantic", # only if this script imports from pydantic
# # add additional dependencies here
# ]
# ///
import data_designer.config as dd
from pydantic import BaseModel, Field
# Use Pydantic models when the output needs to conform to a specific schema
class MyStructuredOutput(BaseModel):
field_one: str = Field(description="...")
field_two: int = Field(description="...")
# Use custom generators when built-in column types aren't enough
@dd.custom_column_generator(
required_columns=["col_a"],
side_effect_columns=["extra_col"],
)
def generator_function(row: dict) -> dict:
# add custom logic here that depends on "col_a" and update row in place
row["name_in_custom_column_config"] = "custom value"
row["extra_col"] = "extra value"
return row
def load_config_builder() -> dd.DataDesignerConfigBuilder:
config_builder = dd.DataDesignerConfigBuilder(
# Declaring model configs programmatically here is the portable path:
# it works for both local `run` and cluster `submit`, while the local
# YAML registry alternative only works for `run`. The provider below
# is a common default created during `nemo setup` — confirm it (or
# discover others) with `nemo inference providers list`. See
# references/nemo-platform-plugin-additions.md for the local-YAML alternative.
model_configs=[
dd.ModelConfig(
alias="text",
model="...",
provider="default/nvidia-build",
inference_parameters=dd.ChatCompletionInferenceParams(),
),
],
)
# Seed dataset (only if the user explicitly mentions a seed dataset path)
# config_builder.with_seed_dataset(dd.LocalFileSeedSource(path="path/to/seed.parquet"))
# config_builder.add_column(...)
# config_builder.add_processor(...)
return config_builder
Pydanticモデル、カスタムジェネレータ、シードデータセット、および追加の依存関係は、タスクで必要とされる場合にのみ含めてください。 model_configs を含めることを推奨します。スクリプト内で宣言することで、ローカル run とクラスタ submit間で設定の移植性を維持できますが、ローカルの YAML レジストリによる代替手段は、 run.
---
name: nemo-data-designer-plugin
description: Build synthetic datasets and data generation pipelines using the Data Designer library.
license: Apache-2.0
---
# Before You Start
Do not explore the workspace first. The workflow's Learn step gives you everything you need.
# Goal
Build a synthetic dataset using the Data Designer library that matches this description:
$ARGUMENTS
# Workflow
Use **Autopilot** mode if the user implies they don't want to answer questions — e.g., they say something like "be opinionated", "you decide", "make reasonable assumptions", "just build it", "surprise me", etc. Otherwise, use **Interactive** mode (default).
Read **only** the workflow file that matches the selected mode, then follow it:
- **Interactive** → read `workflows/interactive.md`
- **Autopilot** → read `workflows/autopilot.md`
# Rules
- Keep all columns in the output by default. The only exceptions for dropping a column are: (1) the user explicitly asks, or (2) it is a helper column that exists solely to derive other columns (e.g., a sampled person object used to extract name, city, etc.). When in doubt, keep the column.
- Do not suggest or ask about seed datasets. Only use one when the user explicitly provides seed data or asks to build from existing records. When using a seed, read `references/seed-datasets.md`.
- When the dataset requires person data (names, demographics, addresses), read `references/person-sampling.md`.
- If a dataset script that matches the dataset description already exists, ask the user whether to edit it or create a new one.
- For commands and context specific to this NeMo Platform plugin (e.g., sourcing model configs from IGW providers or in-script `ModelConfig`s, installing or publishing Nemotron Personas locales, platform-side resource pointers), read `references/nemo-platform-plugin-additions.md`.
# Usage Tips and Common Pitfalls
- **Sampler and validation columns need both a type and params.** E.g., `sampler_type="category"` with `params=dd.CategorySamplerParams(...)`.
- **Jinja2 templates** in `prompt`, `system_prompt`, and `expr` fields: reference columns with `{{ column_name }}`, nested fields with `{{ column_name.field }}`.
- `**SamplerColumnConfig`:** Takes `params`, not `sampler_params`.
- **LLM judge score access:** `LLMJudgeColumnConfig` produces a nested dict where each score name maps to `{reasoning: str, score: int}`. To get the numeric score, use the `.score` attribute. For example, for a judge column named `quality` with a score named `correctness`, use `{{ quality.correctness.score }}`. Using `{{ quality.correctness }}` returns the full dict, not the numeric score.
# Troubleshooting
- `**nemo data-designer` CLI not found:** Tell the user that `nemo data-designer` is not installed in this environment (requires Python >= 3.11). Ask if they would like you to create a virtual environment and install it, or if they prefer to do it themselves. Do not install anything without the user's permission.
- **Network errors during preview:** A sandbox environment may be blocking outbound requests. Ask the user for permission to retry the command with the sandbox disabled. Only as a last resort, if retrying outside the sandbox also fails, tell the user to run the command themselves.
# Output Template
Write a Python file to the current directory with a `load_config_builder()` function returning a `DataDesignerConfigBuilder`. Name the file descriptively (e.g., `customer_reviews.py`). Use PEP 723 inline metadata for dependencies.
```python
# /// script
# dependencies = [
# "data-designer", # always required
# "pydantic", # only if this script imports from pydantic
# # add additional dependencies here
# ]
# ///
import data_designer.config as dd
from pydantic import BaseModel, Field
# Use Pydantic models when the output needs to conform to a specific schema
class MyStructuredOutput(BaseModel):
field_one: str = Field(description="...")
field_two: int = Field(description="...")
# Use custom generators when built-in column types aren't enough
@dd.custom_column_generator(
required_columns=["col_a"],
side_effect_columns=["extra_col"],
)
def generator_function(row: dict) -> dict:
# add custom logic here that depends on "col_a" and update row in place
row["name_in_custom_column_config"] = "custom value"
row["extra_col"] = "extra value"
return row
def load_config_builder() -> dd.DataDesignerConfigBuilder:
config_builder = dd.DataDesignerConfigBuilder(
# Declaring model configs programmatically here is the portable path:
# it works for both local `run` and cluster `submit`, while the local
# YAML registry alternative only works for `run`. The provider below
# is a common default created during `nemo setup` — confirm it (or
# discover others) with `nemo inference providers list`. See
# references/nemo-platform-plugin-additions.md for the local-YAML alternative.
model_configs=[
dd.ModelConfig(
alias="text",
model="...",
provider="default/nvidia-build",
inference_parameters=dd.ChatCompletionInferenceParams(),
),
],
)
# Seed dataset (only if the user explicitly mentions a seed dataset path)
# config_builder.with_seed_dataset(dd.LocalFileSeedSource(path="path/to/seed.parquet"))
# config_builder.add_column(...)
# config_builder.add_processor(...)
return config_builder
```
Only include Pydantic models, custom generators, seed datasets, and extra dependencies when the task requires them. Prefer including `model_configs` when the dataset uses LLM columns — declaring it in the script keeps the config portable between local `run` and cluster `submit`, while the local YAML registry alternative only works for `run`.
すべてのファイル
12件のファイルnemo-data-designer-pluginをインストール
スキルファイルをダウンロードし、.claude/skills/ ディレクトリに解凍してください。
ZIPをダウンロードリポジトリをクローンし、スキルファイルをプロジェクトにコピーしてください。
git clone https://github.com/NVIDIA/skills/tree/main/skills/nemo-data-designer-plugin # Copy SKILL.md to your .claude/skills/ directory
コピー





家
