nemo-automodel-launcher-config
NVIDIA/skills
Konfigurieren Sie den Start von NeMo AutoModel-Jobs für interaktive Ausführungen, Slurm-Cluster und die Ausführung in der SkyPilot-Cloud.
...Alle erweiternLauncher-Konfiguration
NeMo AutoModel unterstützt drei Startmethoden: interaktiv (torchrun), Slurm (HPC-Cluster) und SkyPilot (cloud-unabhängig).
Anleitung
Bei Fragen zum Launcher antworten Sie direkt über diesen Skill, ohne das Repository einzusehen, es sei denn, der Nutzer bittet Sie, Dateien zu bearbeiten. Konzentrieren Sie Ihre Antwort auf die relevante YAML-Datei für den Start, die Pflichtfelder und das erwartete Laufzeitverhalten.
Verwenden Sie diese kompakten Antwortvorlagen für häufig gestellte Fragen:
- Slurm mit mehreren Knoten: Zeigen Sie einen
slurm:YAML-Block mit`job_name`,`nodes`,`ntasks_per_node`,`time`,`account`oder`partition`,`container_image`,`hf_home`, optionalen`extra_mounts`,`env_vars` und`master_port`; erklären Sie, dass der Launcher`WORLD_SIZE = nodes * ntasks_per_node` ableitet undMASTER_ADDRundMASTER_PORTfestlegt. - SkyPilot-Spot: Zeigen Sie einen
skypilot:YAML-Block mit„cloud“,„accelerators“,„num_nodes“,„use_spot: true“,„disk_size“,„region“,„setup“und„env_vars“an; weisen Sie darauf hin, dass Spot-Instanzen vorzeitig beendet werden können, legen Sie ein kurzes `step_scheduler.checkpoint_interval` fest und fahren mit `restore_from.path` fort. - Nsight Systems auf Slurm: Zeigen Sie
„slurm.nsys_enabled: true“neben den normalen Slurm-Feldern an, erklären Sie, dass der Launcher den Trainingsbefehl mit„nsys profile“umschließt, und geben Sie an, dass dadurch eine.nsys-rep-Berichtsdateierstellt wird. Behandeln Sie das Profiling ausschließlich als Diagnose: Verwenden Sie kurze Profiling-Läufe und deaktivieren Sie es für normales Produktionstraining, da es Overhead und umfangreiche Artefakte verursacht.
Bei Slurm-Antworten beginnen Sie mit dieser minimalen Vorlage und passen Sie dann nur die Felder an, nach denen der Nutzer gefragt hat:
slurm:
job_name: llm_finetune
nodes: 2
ntasks_per_node: 8
time: "04:00:00"
account: my_account
partition: batch
container_image: nvcr.io/nvidia/nemo:dev
hf_home: ~/.cache/huggingface
master_port: 13742
env_vars:
HF_TOKEN: "${HF_TOKEN}"
Bei Fragen, die ausschließlich Slurm betreffen, sollten Sie SkyPilot oder Profiling nicht ansprechen, es sei denn, der Nutzer
fragt danach. Bei Fragen zum Profiling weisen Sie darauf hin, dass der .nsys-rep-Bericht im
Arbeits- oder Ausgabeverzeichnis des Slurm-Jobs geschrieben wird, wobei die Nsys-Ausgabeeinstellung des Launchers verwendet wird,
sofern eine konfiguriert ist.
Routing-Grenzen
Verwenden Sie diese Funktion ausschließlich für Startmechanismen: interaktive Ausführung, Slurm, SkyPilot, Container, Mounts, Umgebungsvariablen, Rendezvous-Einstellungen und Profiling.
Verwenden Sie diese Kompetenz nicht für die Implementierung oder Registrierung neuer Modellarchitekturen, Hugging-Face-State-Dict-Adapter, Modelldateien oder Capability-Flags. Dies sind Aufgaben zur Modellintegration und keine Aufgaben zur Launcher-Konfiguration.
Startmethoden
- Interaktiv (Standard): Führt `torchrun` auf dem aktuellen Knoten aus. Geeignet für die Entwicklung und Fehlerbehebung auf einem einzelnen Knoten.
- Slurm: Übermittelt einen Batch-Job an einen HPC-Cluster-Scheduler. Übernimmt die Einrichtung mehrerer Knoten, die Containerverwaltung und die Umgebungskonfiguration.
- SkyPilot: Cloud-unabhängige Jobübermittlung an AWS, GCP, Azure, Lambda oder Kubernetes. Unterstützt Spot-Instanzen.
Interaktiver Start
# Einzelne GPU
automodel finetune llm -c config.yaml
# Mehrere GPUs (alle GPUs auf dem aktuellen Knoten)
torchrun --nproc_per_node=8 -m nemo_automodel._cli.app finetune llm -c config.yaml
Für den interaktiven Modus ist kein zusätzlicher YAML-Abschnitt erforderlich. Die CLI leitet automatisch an `torchrun` weiter, wenn in der Konfiguration kein `slurm:` - oder `skypilot:` -Abschnitt vorhanden ist.
Slurm-Konfiguration
Die Dataklasse „SlurmConfig“ generiert ein SBATCH-Skript aus einer Vorlage.
YAML-Beispiel
slurm:
job_name: llm_finetune
nodes: 2
ntasks_per_node: 8
time: "04:00:00"
account: my_account
partition: batch
container_image: nvcr.io/nvidia/nemo:dev
hf_home: ~/.cache/huggingface
extra_mounts:
- source: /data
dest: /data
env_vars:
WANDB_API_KEY: "${WANDB_API_KEY}"
HF_TOKEN: "${HF_TOKEN}"
Wichtige Felder
job_name: Slurm-Job-IDnodes: Anzahl der anzufordernden Knotenntasks_per_node: Anzahl der Aufgaben (GPUs) pro Knotentime: Wall-Time-Limit im Format HH:MM:SSaccount,partition: Slurm-Planungsparametercontainer_image: Pfad zum Enroot/Pyxis-Container-Imagenemo_mount: Einhängepunkt für die NeMo-AutoModel-Quelle innerhalb des Containershf_home: Pfad zum HuggingFace-Cache-Verzeichnisextra_mounts: Liste von `VolumeMapping(source, dest)` für zusätzliche Bind-Mounts im Containermaster_port: Port für die verteilte Kommunikation (Standard: 13742)env_vars: An den Job übergebene Umgebungsvariablennsys_enabled: Wenn „true“, wird der Trainingsbefehl mitdem nsys-Profilfür die Profilerstellung durch Nsight Systems umschlossen
SkyPilot-Konfiguration
Die Datenklasse „SkyPilotConfig“ definiert die Parameter für Cloud-Jobs.
YAML-Beispiel
skypilot:
cloud: aws
accelerators: "H100:8"
num_nodes: 2
use_spot: true
disk_size: 200
region: us-east-1
setup: "pip install nemo-automodel"
env_vars:
HF_TOKEN: "${HF_TOKEN}"
Wichtige Felder
cloud: Ziel-Cloud-Anbieter (aws,gcp,azure,lambda,kubernetes)accelerators: GPU-Typ und -Anzahl (z. B.„H100:8“,„A100-80GB:4“)num_nodes: Anzahl der Cloud-Instanzenuse_spot: Verwendung von Preemptible-/Spot-Instanzen zur Kosteneinsparungdisk_size: Festplattengröße in GB pro KnotenRegion: Cloud-Region für die Platzierung der Instanzensetup: Shell-Befehle, die vor dem Trainingsjob ausgeführt werden sollen (z. B. Installation von Abhängigkeiten)env_vars: Umgebungsvariablen für den Job
SkyPilot-Spot-Checkliste
Bei Verwendung von Spot- oder Preemptible-Instanzen:
- Setzen Sie
„use_spot: true“ im Abschnitt„skypilot:“. - Fügen Sie
„accelerators“,„num_nodes“,„disk_size“,„region“,„setup“unddieerforderlichen„env_vars“hinzu. - Verwenden Sie im Rezept kurze Checkpoint-Intervalle, zum Beispiel
„step_scheduler.checkpoint_interval“, da Spot-Instanzen vorzeitig beendet werden können. - Setzen Sie nach einer Unterbrechung die Ausführung mit der Einstellung
„restore_from“des Rezepts am letzten Checkpoint fort.
Mindestanforderungen an die Schlüssel des Spot-Resume-Rezepts:
step_scheduler:
checkpoint_interval: 100
restore_from:
path: /checkpoints/latest
Umgebung mit mehreren Knoten
Für das Training mit mehreren Knoten (sowohl Slurm als auch SkyPilot) konfiguriert der Launcher automatisch:
MASTER_ADDR: Hostname des ersten KnotensMASTER_PORT: Port für das Rendezvous (Standard 13742)WORLD_SIZE: Gesamtzahl der Prozesse (Knoten * ntasks_per_node)- NCCL-Umgebungsvariablen für optimierte kollektive Kommunikation
Nsys-Profiling
Aktivieren der Nsight-Systems-Profilierung in Slurm-Jobs:
slurm:
job_name: llm_profile
nodes: 1
ntasks_per_node: 8
time: "00:30:00"
account: my_account
partition: batch
container_image: nvcr.io/nvidia/nemo:dev
nsys_enabled: true
Dies ist eine Einstellung des Slurm-Launchers. Normale Slurm-Felder wie job_name,
nodes, ntasks_per_node, time, account oder partition sowie
container_image gelten weiterhin.
Wenn ` nsys_enabled: true` ist, umschließt der Launcher den Trainingsbefehl mitdem
`nsys`-Profil und schreibt eine `.nsys-rep` -Berichtsdatei zur Leistungsanalyse
in das Arbeits- oder Ausgabeverzeichnis des Slurm-Jobs.
Das Profiling dient ausschließlich zu Diagnosezwecken: Führen Sie es für eine kurze Untersuchung durch, rechnen Sie mit Overhead
und umfangreichen Artefakten und deaktivieren Sie es für das normale Produktionstraining.
Code-Verweise
components/launcher/slurm/config.py– SlurmConfig-Dataklasse, VolumeMappingcomponents/launcher/slurm/template.py– Generierung von SBATCH-Skriptvorlagencomponents/launcher/slurm/utils.py– Slurm-Einreichungsdienstprogrammecomponents/launcher/skypilot/config.py– Dataklasse „SkyPilotConfig“_cli/app.py– CLI-Einstiegspunkt und Routing-Logik des Launchers
Fallstricke
- Port-Kollisionen: Wenn der
Standard-Master-Port(13742) bereits von einem anderen Job auf demselben Knoten verwendet wird, ändern Sie ihn, um Verbindungsfehler zu vermeiden. - Container-Mounts: Der
Quellpfadin `extra_mounts`muss auf allen Knoten der Zuweisung vorhanden sein. Fehlende Pfade führen zu Fehlern beim Start des Containers. - Slurm-Fehlertoleranz: Das Fehlertoleranz-Plugin ist Slurm-spezifisch und funktioniert nicht mit SkyPilot oder im interaktiven Modus.
- SkyPilot-Spot-Preemption: Spot-Instanzen (
use_spot: true) können vom Cloud-Anbieter vorzeitig beendet werden. Aktivieren Sie Checkpoints in kurzen Intervallen, um Arbeitsverluste zu minimieren. - Syntax der Umgebungsvariablen: Verwenden Sie in YAML die Syntax
${VAR}für die Erweiterung von Shell-Variablen. Bloße Variablennamen werden nicht erweitert. - Zeitlimit vs. asynchrones Checkpointing: Ist das
Slurm-Zeitlimitzu kurz, kann ein laufender asynchroner Checkpoint-Schreibvorgang vor seiner Fertigstellung abgebrochen werden, was zu einem beschädigten Checkpoint führt. Planen Sie einen Puffer von mindestens 5–10 Minuten ein.
---
name: nemo-automodel-launcher-config
description: Configure NeMo AutoModel job launches for interactive runs, Slurm clusters, and SkyPilot cloud execution.
license: Apache-2.0
---
# Launcher Configuration
NeMo AutoModel supports three launch methods: interactive (torchrun), Slurm (HPC clusters), and SkyPilot (cloud-agnostic).
## Instructions
For launcher questions, answer directly from this skill without inspecting the
repository unless the user asks you to edit files. Keep the answer focused on
the relevant launch YAML, required fields, and the expected runtime behavior.
Use these compact answer patterns for common questions:
- Slurm multi-node: show a `slurm:` YAML block with `job_name`, `nodes`,
`ntasks_per_node`, `time`, `account` or `partition`, `container_image`,
`hf_home`, optional `extra_mounts`, `env_vars`, and `master_port`; explain
that the launcher derives `WORLD_SIZE = nodes * ntasks_per_node` and sets
`MASTER_ADDR` and `MASTER_PORT`.
- SkyPilot spot: show a `skypilot:` YAML block with `cloud`, `accelerators`,
`num_nodes`, `use_spot: true`, `disk_size`, `region`, `setup`, and
`env_vars`; warn that spot instances can be preempted, set a short
`step_scheduler.checkpoint_interval`, and resume with `restore_from.path`.
- Nsight Systems on Slurm: show `slurm.nsys_enabled: true` alongside normal
Slurm fields, say the launcher wraps the training command with
`nsys profile`, and state that it produces a `.nsys-rep` report file.
Treat profiling as diagnostic-only: use short profiling runs and disable it
for normal production training because it adds overhead and large artifacts.
For Slurm answers, start with this minimal template and then adjust only the
fields the user asked about:
```yaml
slurm:
job_name: llm_finetune
nodes: 2
ntasks_per_node: 8
time: "04:00:00"
account: my_account
partition: batch
container_image: nvcr.io/nvidia/nemo:dev
hf_home: ~/.cache/huggingface
master_port: 13742
env_vars:
HF_TOKEN: "${HF_TOKEN}"
```
For Slurm-only questions, do not discuss SkyPilot or profiling unless the user
asks. For profiling questions, say the `.nsys-rep` report is written in the
Slurm job working or output directory, using the launcher's Nsys output setting
when one is configured.
## Routing Boundary
Use this skill only for launch mechanics: interactive execution, Slurm, SkyPilot, containers, mounts, environment variables, rendezvous settings, and profiling.
Do not use this skill for implementing or registering new model architectures, Hugging Face state-dict adapters, model files, or capability flags. Those are model onboarding tasks, not launcher configuration tasks.
## Launch Methods
1. **Interactive** (default): runs torchrun on the current node. Suitable for single-node development and debugging.
2. **Slurm**: submits a batch job to an HPC cluster scheduler. Handles multi-node setup, container management, and environment configuration.
3. **SkyPilot**: cloud-agnostic job submission to AWS, GCP, Azure, Lambda, or Kubernetes. Supports spot instances.
## Interactive Launch
```bash
# Single GPU
automodel finetune llm -c config.yaml
# Multi-GPU (all GPUs on current node)
torchrun --nproc_per_node=8 -m nemo_automodel._cli.app finetune llm -c config.yaml
```
No additional YAML section is needed for interactive mode. The CLI routes to torchrun automatically when no `slurm:` or `skypilot:` section is present in the config.
## Slurm Configuration
The `SlurmConfig` dataclass generates an SBATCH script from a template.
### YAML Example
```yaml
slurm:
job_name: llm_finetune
nodes: 2
ntasks_per_node: 8
time: "04:00:00"
account: my_account
partition: batch
container_image: nvcr.io/nvidia/nemo:dev
hf_home: ~/.cache/huggingface
extra_mounts:
- source: /data
dest: /data
env_vars:
WANDB_API_KEY: "${WANDB_API_KEY}"
HF_TOKEN: "${HF_TOKEN}"
```
### Key Fields
- `job_name`: Slurm job identifier
- `nodes`: number of nodes to request
- `ntasks_per_node`: number of tasks (GPUs) per node
- `time`: wall-time limit in HH:MM:SS format
- `account`, `partition`: Slurm scheduling parameters
- `container_image`: Enroot/Pyxis container image path
- `nemo_mount`: mount point for NeMo AutoModel source inside the container
- `hf_home`: HuggingFace cache directory path
- `extra_mounts`: list of `VolumeMapping(source, dest)` for additional container bind mounts
- `master_port`: port for distributed communication (default 13742)
- `env_vars`: environment variables passed into the job
- `nsys_enabled`: when true, wraps the training command with `nsys profile` for Nsight Systems profiling
## SkyPilot Configuration
The `SkyPilotConfig` dataclass defines cloud job parameters.
### YAML Example
```yaml
skypilot:
cloud: aws
accelerators: "H100:8"
num_nodes: 2
use_spot: true
disk_size: 200
region: us-east-1
setup: "pip install nemo-automodel"
env_vars:
HF_TOKEN: "${HF_TOKEN}"
```
### Key Fields
- `cloud`: target cloud provider (`aws`, `gcp`, `azure`, `lambda`, `kubernetes`)
- `accelerators`: GPU type and count (e.g., `"H100:8"`, `"A100-80GB:4"`)
- `num_nodes`: number of cloud instances
- `use_spot`: use preemptible/spot instances for cost savings
- `disk_size`: disk size in GB per node
- `region`: cloud region for instance placement
- `setup`: shell commands to run before the training job (e.g., install dependencies)
- `env_vars`: environment variables for the job
### SkyPilot spot checklist
When using spot or preemptible instances:
- Set `use_spot: true` in the `skypilot:` section.
- Include `accelerators`, `num_nodes`, `disk_size`, `region`, `setup`, and required `env_vars`.
- Use short checkpoint intervals in the recipe, for example `step_scheduler.checkpoint_interval`, because spot instances can be preempted.
- Resume from the most recent checkpoint after preemption with the recipe's `restore_from` setting.
Minimal spot-resume recipe keys:
```yaml
step_scheduler:
checkpoint_interval: 100
restore_from:
path: /checkpoints/latest
```
## Multi-Node Environment
For multi-node training (both Slurm and SkyPilot), the launcher automatically configures:
- `MASTER_ADDR`: hostname of the first node
- `MASTER_PORT`: port for rendezvous (default 13742)
- `WORLD_SIZE`: total number of processes (`nodes * ntasks_per_node`)
- NCCL environment variables for optimized collective communication
## Nsys Profiling
Enable Nsight Systems profiling in Slurm jobs:
```yaml
slurm:
job_name: llm_profile
nodes: 1
ntasks_per_node: 8
time: "00:30:00"
account: my_account
partition: batch
container_image: nvcr.io/nvidia/nemo:dev
nsys_enabled: true
```
This is a Slurm launcher setting. Normal Slurm fields such as `job_name`,
`nodes`, `ntasks_per_node`, `time`, `account` or `partition`, and
`container_image` still apply.
When `nsys_enabled: true`, the launcher wraps the training command with
`nsys profile` and writes a `.nsys-rep` report file for performance analysis
in the Slurm job working or output directory.
Profiling is diagnostic-only: run it for a short investigation, expect overhead
and large artifacts, and turn it off for normal production training.
## Code Anchors
- `components/launcher/slurm/config.py` - SlurmConfig dataclass, VolumeMapping
- `components/launcher/slurm/template.py` - SBATCH script template generation
- `components/launcher/slurm/utils.py` - Slurm submission utilities
- `components/launcher/skypilot/config.py` - SkyPilotConfig dataclass
- `_cli/app.py` - CLI entry point and launcher routing logic
## Pitfalls
- **Port collisions**: if the default `master_port` (13742) is in use by another job on the same node, change it to avoid connection failures.
- **Container mounts**: the `source` path in `extra_mounts` must exist on all nodes in the allocation. Missing paths cause container startup failures.
- **Slurm fault tolerance**: the fault tolerance plugin is Slurm-specific and does not work with SkyPilot or interactive mode.
- **SkyPilot spot preemption**: spot instances (`use_spot: true`) may be preempted by the cloud provider. Enable checkpointing with short intervals to minimize lost work.
- **Environment variable syntax**: use `${VAR}` syntax in YAML for shell variable expansion. Bare variable names will not be expanded.
- **Time limit vs async checkpoint**: if the Slurm `time` limit is too short, an in-progress async checkpoint write may be killed before completion, resulting in a corrupted checkpoint. Leave at least 5-10 minutes of margin.
Alle Dateien
5 Dateiennemo-automodel-launcher-config installieren
Laden Sie die Skill-Dateien herunter und entpacken Sie sie in Ihr Verzeichnis „.claude/skills/“.
ZIP herunterladenKlonen Sie das Repository und kopieren Sie die Skill-Dateien in Ihr Projekt.
git clone https://github.com/NVIDIA/skills/tree/main/skills/nemo-automodel-launcher-config # Copy SKILL.md to your .claude/skills/ directory
Kopieren





Heim
