nemo-automodel-launcher-config
NVIDIA/skills
対話型実行、Slurmクラスタ、およびSkyPilotクラウド実行向けに、NeMo AutoModelのジョブ起動を設定します。
...すべて拡張しますランチャーの設定
NeMo AutoModel では、対話型(torchrun)、Slurm(HPC クラスタ)、SkyPilot(クラウド非依存)の 3 つの起動方法をサポートしています。
手順
ランチャーに関する質問については、ユーザーからファイルの編集を求められない限り、 リポジトリを確認せずに、このスキルから直接回答してください。回答は、 関連する起動用 YAML、必須フィールド、および期待される実行時の動作に焦点を当ててください。
よくある質問には、以下の簡潔な回答パターンを使用してください:
- Slurm マルチノード:
slurm:YAMLブロックを表示し、job_name、nodes、ntasks_per_node、time、accountまたはpartition、container_image、hf_home、オプションのextra_mounts、env_vars、およびmaster_portを含めてください。また、 ランチャーがWORLD_SIZE = nodes * ntasks_per_nodeを算出し、MASTER_ADDRおよびMASTER_PORTを設定することを説明してください。 - SkyPilotスポット:
cloud、accelerators、num_nodes、use_spot: true、disk_size、region、setup、およびenv_varsを含むskypilot:YAMLブロックを表示する。スポットインスタンスはプリエンプトされる可能性があることを警告し、短いstep_scheduler.checkpoint_intervalを設定し、restore_from.pathで再開する。 - Slurm 上の Nsight Systems:通常の
Slurm フィールドとともに
`slurm.nsys_enabled: true` を表示し、ランチャーがトレーニングコマンドを`nsys` プロファイルでラップすることを説明し、それが `.nsys-rep`レポートファイルを生成することを明記する。 プロファイリングは診断目的のみに扱う:プロファイリングの実行時間を短くし、 通常の本番トレーニングではオーバーヘッドや大きなアーティファクトが発生するため無効にする。
Slurm に関する回答については、この最小限のテンプレートから始め、ユーザーが質問した フィールドのみを調整してください:
slurm:
job_name: llm_finetune
nodes: 2
ntasks_per_node: 8
time: "04:00:00"
account: my_account
partition: batch
container_image: nvcr.io/nvidia/nemo:dev
hf_home: ~/.cache/huggingface
master_port: 13742
env_vars:
HF_TOKEN: "${HF_TOKEN}"
Slurm に関する質問のみの場合、ユーザーから
求められない限り、SkyPilot やプロファイリングについては言及しないでください。プロファイリングに関する質問については、.nsys-repレポートが
Slurm ジョブの作業ディレクトリまたは出力ディレクトリに書き込まれていること、また、設定されている場合はランチャーの Nsys 出力設定が
使用されていることを伝えてください。
対応範囲
このスキルは、起動の仕組み(対話型実行、Slurm、SkyPilot、コンテナ、マウント、環境変数、ランデブー設定、プロファイリング)にのみ使用してください。
新しいモデルアーキテクチャ、Hugging Faceの状態指定アダプタ、モデルファイル、または機能フラグの実装や登録には、このスキルを使用しないでください。これらはモデルのオンボーディングタスクであり、ランチャーの設定タスクではありません。
起動方法
- 対話型(デフォルト):現在のノード上で torchrun を実行します。シングルノードでの開発やデバッグに適しています。
- Slurm:HPCクラスタースケジューラにバッチジョブをサブミットします。マルチノード設定、コンテナ管理、環境設定を処理します。
- SkyPilot:AWS、GCP、Azure、Lambda、またはKubernetesへの、クラウドに依存しないジョブ送信。スポットインスタンスをサポートしています。
対話型起動
# シングルGPU
automodel finetune llm -c config.yaml
# マルチGPU(現在のノード上のすべてのGPU)
torchrun --nproc_per_node=8 -m nemo_automodel._cli.app finetune llm -c config.yaml
対話モードでは、追加の YAML セクションは必要ありません。設定にslurm:またはskypilot:セクションが含まれていない場合、CLI は自動的に torchrun にリダイレクトします。
Slurmの設定
SlurmConfigデータクラスは、テンプレートから SBATCH スクリプトを生成します。
YAMLの例
slurm:
job_name: llm_finetune
nodes: 2
ntasks_per_node: 8
time: "04:00:00"
account: my_account
partition: batch
container_image: nvcr.io/nvidia/nemo:dev
hf_home: ~/.cache/huggingface
extra_mounts:
- source: /data
dest: /data
env_vars:
WANDB_API_KEY: "${WANDB_API_KEY}"
HF_TOKEN: "${HF_TOKEN}"
主要なフィールド
job_name: Slurm ジョブ識別子nodes: 要求するノード数ntasks_per_node: ノードあたりのタスク数(GPU数)time: ウォールタイム制限(HH:MM:SS形式)account,partition: Slurm スケジューリングパラメータcontainer_image: Enroot/Pyxis コンテナイメージのパスnemo_mount: コンテナ内の NeMo AutoModel ソースのマウントポイントhf_home: HuggingFaceキャッシュディレクトリのパスextra_mounts: 追加のコンテナバインドマウント用のVolumeMapping(ソース, 宛先)のリストmaster_port: 分散通信用のポート(デフォルト 13742)env_vars: ジョブに渡される環境変数nsys_enabled: true の場合、Nsight Systems のプロファイリング用にトレーニングコマンドをnsys プロファイルでラップする
SkyPilot の設定
SkyPilotConfigデータクラスは、クラウドジョブのパラメータを定義します。
YAML の例
skypilot:
cloud: aws
accelerators: "H100:8"
num_nodes: 2
use_spot: true
disk_size: 200
region: us-east-1
setup: "pip install nemo-automodel"
env_vars:
HF_TOKEN: "${HF_TOKEN}"
主要なフィールド
cloud: 対象のクラウドプロバイダー (aws,gcp,azure,lambda,kubernetes)accelerators: GPUの種類と台数(例:"H100:8","A100-80GB:4")num_nodes: クラウドインスタンスの数use_spot: コスト削減のためにプリエンプティブル/スポットインスタンスを使用するdisk_size: ノードあたりのディスク容量(GB単位)region: インスタンスを配置するクラウドリージョンsetup: トレーニングジョブの実行前に実行するシェルコマンド(例:依存関係のインストール)env_vars: ジョブ用の環境変数
SkyPilot スポットチェックリスト
スポットインスタンスまたはプリエンプティブルインスタンスを使用する場合:
skypilot:セクションでuse_spot: trueを設定してください。accelerators、num_nodes、disk_size、region、setup、および必要なenv_vars を指定してください。- スポットインスタンスはプリエンプトされる可能性があるため、レシピ(例:
step_scheduler.checkpoint_interval)ではチェックポイントの間隔を短く設定してください。 - プリエンプト発生後は、レシピの
`restore_from` 設定を使用して、最新のチェックポイントから再開してください。
Spot 再開用レシピの最小限のキー:
step_scheduler:
checkpoint_interval: 100
restore_from:
path: /checkpoints/latest
マルチノード環境
マルチノードでのトレーニング(Slurm および SkyPilot の両方)の場合、ランチャーは自動的に以下を設定します:
MASTER_ADDR: 最初のノードのホスト名MASTER_PORT: ランデブー用のポート(デフォルト 13742)WORLD_SIZE: プロセスの総数 (ノード数 × ntasks_per_node)- 集合通信を最適化するためのNCCL環境変数
Nsys プロファイリング
Slurm ジョブで Nsight Systems プロファイリングを有効にする:
slurm:
job_name: llm_profile
nodes: 1
ntasks_per_node: 8
time: "00:30:00"
account: my_account
partition: batch
container_image: nvcr.io/nvidia/nemo:dev
nsys_enabled: true
これは Slurm ランチャーの設定です。job_name、
nodes、ntasks_per_node、time、account、partition、および
container_imageといった通常の Slurm フィールドは引き続き適用されます。
nsys_enabled: true の場合、ランチャーはトレーニングコマンドを
nsys プロファイルでラップし、パフォーマンス分析用の.nsys-repレポートファイルを
Slurm ジョブの作業ディレクトリまたは出力ディレクトリに書き込みます。
プロファイリングは診断目的のみです。短時間の調査のために実行し、オーバーヘッドや
大規模なアーティファクトの発生を想定した上で、通常の本番トレーニングでは無効にしてください。
コードの参照先
components/launcher/slurm/config.py- SlurmConfig データクラス、VolumeMappingcomponents/launcher/slurm/template.py- SBATCHスクリプトテンプレートの生成components/launcher/slurm/utils.py- Slurm サブミッションユーティリティcomponents/launcher/skypilot/config.py- SkyPilotConfig データクラス_cli/app.py- CLIのエントリポイントおよびランチャーのルーティングロジック
注意点
- ポートの競合:デフォルトの
master_port(13742) が同じノード上の別のジョブで使用されている場合は、接続エラーを回避するために変更してください。 - コンテナのマウント:
extra_mounts内のソースパスは、割り当てられたすべてのノード上に存在している必要があります。パスが存在しない場合、コンテナの起動に失敗します。 - Slurm のフォールトトレランス:フォールトトレランスプラグインは Slurm 専用であり、SkyPilot や対話モードでは動作しません。
- SkyPilot スポットインスタンスのプリエンプション:スポットインスタンス (
use_spot: true) は、クラウドプロバイダーによってプリエンプションされる可能性があります。作業の損失を最小限に抑えるため、短い間隔でチェックポイントを作成するように設定してください。 - 環境変数の構文:YAML でのシェル変数の展開には
${VAR}構文を使用してください。変数名のみを記述した場合は展開されません。 - 時間制限と非同期チェックポイント: Slurm
の時間制限が短すぎると、進行中の非同期チェックポイントの書き込みが完了前に強制終了され、チェックポイントが破損する可能性があります。少なくとも 5~10 分程度の余裕を持たせてください。
---
name: nemo-automodel-launcher-config
description: Configure NeMo AutoModel job launches for interactive runs, Slurm clusters, and SkyPilot cloud execution.
license: Apache-2.0
---
# Launcher Configuration
NeMo AutoModel supports three launch methods: interactive (torchrun), Slurm (HPC clusters), and SkyPilot (cloud-agnostic).
## Instructions
For launcher questions, answer directly from this skill without inspecting the
repository unless the user asks you to edit files. Keep the answer focused on
the relevant launch YAML, required fields, and the expected runtime behavior.
Use these compact answer patterns for common questions:
- Slurm multi-node: show a `slurm:` YAML block with `job_name`, `nodes`,
`ntasks_per_node`, `time`, `account` or `partition`, `container_image`,
`hf_home`, optional `extra_mounts`, `env_vars`, and `master_port`; explain
that the launcher derives `WORLD_SIZE = nodes * ntasks_per_node` and sets
`MASTER_ADDR` and `MASTER_PORT`.
- SkyPilot spot: show a `skypilot:` YAML block with `cloud`, `accelerators`,
`num_nodes`, `use_spot: true`, `disk_size`, `region`, `setup`, and
`env_vars`; warn that spot instances can be preempted, set a short
`step_scheduler.checkpoint_interval`, and resume with `restore_from.path`.
- Nsight Systems on Slurm: show `slurm.nsys_enabled: true` alongside normal
Slurm fields, say the launcher wraps the training command with
`nsys profile`, and state that it produces a `.nsys-rep` report file.
Treat profiling as diagnostic-only: use short profiling runs and disable it
for normal production training because it adds overhead and large artifacts.
For Slurm answers, start with this minimal template and then adjust only the
fields the user asked about:
```yaml
slurm:
job_name: llm_finetune
nodes: 2
ntasks_per_node: 8
time: "04:00:00"
account: my_account
partition: batch
container_image: nvcr.io/nvidia/nemo:dev
hf_home: ~/.cache/huggingface
master_port: 13742
env_vars:
HF_TOKEN: "${HF_TOKEN}"
```
For Slurm-only questions, do not discuss SkyPilot or profiling unless the user
asks. For profiling questions, say the `.nsys-rep` report is written in the
Slurm job working or output directory, using the launcher's Nsys output setting
when one is configured.
## Routing Boundary
Use this skill only for launch mechanics: interactive execution, Slurm, SkyPilot, containers, mounts, environment variables, rendezvous settings, and profiling.
Do not use this skill for implementing or registering new model architectures, Hugging Face state-dict adapters, model files, or capability flags. Those are model onboarding tasks, not launcher configuration tasks.
## Launch Methods
1. **Interactive** (default): runs torchrun on the current node. Suitable for single-node development and debugging.
2. **Slurm**: submits a batch job to an HPC cluster scheduler. Handles multi-node setup, container management, and environment configuration.
3. **SkyPilot**: cloud-agnostic job submission to AWS, GCP, Azure, Lambda, or Kubernetes. Supports spot instances.
## Interactive Launch
```bash
# Single GPU
automodel finetune llm -c config.yaml
# Multi-GPU (all GPUs on current node)
torchrun --nproc_per_node=8 -m nemo_automodel._cli.app finetune llm -c config.yaml
```
No additional YAML section is needed for interactive mode. The CLI routes to torchrun automatically when no `slurm:` or `skypilot:` section is present in the config.
## Slurm Configuration
The `SlurmConfig` dataclass generates an SBATCH script from a template.
### YAML Example
```yaml
slurm:
job_name: llm_finetune
nodes: 2
ntasks_per_node: 8
time: "04:00:00"
account: my_account
partition: batch
container_image: nvcr.io/nvidia/nemo:dev
hf_home: ~/.cache/huggingface
extra_mounts:
- source: /data
dest: /data
env_vars:
WANDB_API_KEY: "${WANDB_API_KEY}"
HF_TOKEN: "${HF_TOKEN}"
```
### Key Fields
- `job_name`: Slurm job identifier
- `nodes`: number of nodes to request
- `ntasks_per_node`: number of tasks (GPUs) per node
- `time`: wall-time limit in HH:MM:SS format
- `account`, `partition`: Slurm scheduling parameters
- `container_image`: Enroot/Pyxis container image path
- `nemo_mount`: mount point for NeMo AutoModel source inside the container
- `hf_home`: HuggingFace cache directory path
- `extra_mounts`: list of `VolumeMapping(source, dest)` for additional container bind mounts
- `master_port`: port for distributed communication (default 13742)
- `env_vars`: environment variables passed into the job
- `nsys_enabled`: when true, wraps the training command with `nsys profile` for Nsight Systems profiling
## SkyPilot Configuration
The `SkyPilotConfig` dataclass defines cloud job parameters.
### YAML Example
```yaml
skypilot:
cloud: aws
accelerators: "H100:8"
num_nodes: 2
use_spot: true
disk_size: 200
region: us-east-1
setup: "pip install nemo-automodel"
env_vars:
HF_TOKEN: "${HF_TOKEN}"
```
### Key Fields
- `cloud`: target cloud provider (`aws`, `gcp`, `azure`, `lambda`, `kubernetes`)
- `accelerators`: GPU type and count (e.g., `"H100:8"`, `"A100-80GB:4"`)
- `num_nodes`: number of cloud instances
- `use_spot`: use preemptible/spot instances for cost savings
- `disk_size`: disk size in GB per node
- `region`: cloud region for instance placement
- `setup`: shell commands to run before the training job (e.g., install dependencies)
- `env_vars`: environment variables for the job
### SkyPilot spot checklist
When using spot or preemptible instances:
- Set `use_spot: true` in the `skypilot:` section.
- Include `accelerators`, `num_nodes`, `disk_size`, `region`, `setup`, and required `env_vars`.
- Use short checkpoint intervals in the recipe, for example `step_scheduler.checkpoint_interval`, because spot instances can be preempted.
- Resume from the most recent checkpoint after preemption with the recipe's `restore_from` setting.
Minimal spot-resume recipe keys:
```yaml
step_scheduler:
checkpoint_interval: 100
restore_from:
path: /checkpoints/latest
```
## Multi-Node Environment
For multi-node training (both Slurm and SkyPilot), the launcher automatically configures:
- `MASTER_ADDR`: hostname of the first node
- `MASTER_PORT`: port for rendezvous (default 13742)
- `WORLD_SIZE`: total number of processes (`nodes * ntasks_per_node`)
- NCCL environment variables for optimized collective communication
## Nsys Profiling
Enable Nsight Systems profiling in Slurm jobs:
```yaml
slurm:
job_name: llm_profile
nodes: 1
ntasks_per_node: 8
time: "00:30:00"
account: my_account
partition: batch
container_image: nvcr.io/nvidia/nemo:dev
nsys_enabled: true
```
This is a Slurm launcher setting. Normal Slurm fields such as `job_name`,
`nodes`, `ntasks_per_node`, `time`, `account` or `partition`, and
`container_image` still apply.
When `nsys_enabled: true`, the launcher wraps the training command with
`nsys profile` and writes a `.nsys-rep` report file for performance analysis
in the Slurm job working or output directory.
Profiling is diagnostic-only: run it for a short investigation, expect overhead
and large artifacts, and turn it off for normal production training.
## Code Anchors
- `components/launcher/slurm/config.py` - SlurmConfig dataclass, VolumeMapping
- `components/launcher/slurm/template.py` - SBATCH script template generation
- `components/launcher/slurm/utils.py` - Slurm submission utilities
- `components/launcher/skypilot/config.py` - SkyPilotConfig dataclass
- `_cli/app.py` - CLI entry point and launcher routing logic
## Pitfalls
- **Port collisions**: if the default `master_port` (13742) is in use by another job on the same node, change it to avoid connection failures.
- **Container mounts**: the `source` path in `extra_mounts` must exist on all nodes in the allocation. Missing paths cause container startup failures.
- **Slurm fault tolerance**: the fault tolerance plugin is Slurm-specific and does not work with SkyPilot or interactive mode.
- **SkyPilot spot preemption**: spot instances (`use_spot: true`) may be preempted by the cloud provider. Enable checkpointing with short intervals to minimize lost work.
- **Environment variable syntax**: use `${VAR}` syntax in YAML for shell variable expansion. Bare variable names will not be expanded.
- **Time limit vs async checkpoint**: if the Slurm `time` limit is too short, an in-progress async checkpoint write may be killed before completion, resulting in a corrupted checkpoint. Leave at least 5-10 minutes of margin.
すべてのファイル
5件のファイルnemo-automodel-launcher-configをインストール
スキルファイルをダウンロードし、.claude/skills/ ディレクトリに解凍してください。
ZIPをダウンロードリポジトリをクローンし、スキルファイルをプロジェクトにコピーしてください。
git clone https://github.com/NVIDIA/skills/tree/main/skills/nemo-automodel-launcher-config # Copy SKILL.md to your .claude/skills/ directory
コピー





家
