nemo-automodel-launcher-config
NVIDIA/skills
대화형 실행, Slurm 클러스터 및 SkyPilot 클라우드 실행을 위해 NeMo AutoModel 작업 실행을 구성합니다.
...모든 것을 확장하십시오런처 구성
NeMo AutoModel은 대화형(torchrun), Slurm(HPC 클러스터), SkyPilot(클라우드 독립형)의 세 가지 실행 방식을 지원합니다.
지침
런처 관련 질문의 경우, 사용자가 파일 수정을 요청하지 않는 한 리포지토리를 확인하지 말고 이 스킬에서 직접 답변해 주세요. 답변은 관련 런치 YAML, 필수 입력 항목, 예상되는 실행 시 동작에 초점을 맞추어 주세요.
일반적인 질문에는 다음과 같은 간결한 답변 패턴을 사용하십시오:
- Slurm 멀티노드:
slurm:YAML 블록을 표시하고,job_name,nodes,ntasks_per_node,time,account또는partition,container_image,hf_home, 선택 사항인extra_mounts,env_vars및master_port를포함시키세요. 또한 런처가WORLD_SIZE = nodes * ntasks_per_node를도출하고MASTER_ADDR및MASTER_PORT를설정한다고 설명하십시오. - SkyPilot 스팟:
cloud,accelerators,num_nodes,use_spot: true,disk_size,region,setup및env_vars가포함된skypilot:YAML 블록을 표시하고, 스팟 인스턴스는 선점될 수 있음을 경고하며, 짧은step_scheduler.checkpoint_interval을설정하고,restore_from.path로 재개합니다. - Slurm상의 Nsight Systems: 일반
Slurm 필드와 함께
slurm.nsys_enabled: true를표시하고, 런처가 훈련 명령을nsys 프로파일로감싸며,.nsys-rep보고서 파일을 생성한다고 설명합니다. 프로파일링은 진단 목적으로만 취급하십시오: 짧은 프로파일링 실행을 사용하고, 오버헤드와 큰 아티팩트를 유발하므로 일반적인 운영 환경의 훈련에서는 비활성화하십시오.
Slurm 관련 답변의 경우, 이 최소 템플릿을 바탕으로 시작하고 사용자가 문의한 필드만 조정하십시오:
slurm:
job_name: llm_finetune
nodes: 2
ntasks_per_node: 8
time: "04:00:00"
account: my_account
partition: batch
container_image: nvcr.io/nvidia/nemo:dev
hf_home: ~/.cache/huggingface
master_port: 13742
env_vars:
HF_TOKEN: "${HF_TOKEN}"
Slurm 관련 질문의 경우, 사용자가 직접 묻지 않는 한 SkyPilot이나 프로파일링에 대해
논의하지 마십시오. 프로파일링 관련 질문이 있을 경우, .nsys-rep 보고서가
Slurm 작업의 작업 디렉터리 또는 출력 디렉터리에 작성된다고 설명하고,
설정되어 있는 경우 런처의 Nsys 출력 설정을 사용한다고 덧붙이십시오.
라우팅 경계
이 스킬은 실행 메커니즘(대화형 실행, Slurm, SkyPilot, 컨테이너, 마운트, 환경 변수, 랑데뷰 설정, 프로파일링)에만 사용하십시오.
새로운 모델 아키텍처, Hugging Face 상태 기반 어댑터, 모델 파일 또는 기능 플래그의 구현이나 등록에는 이 스킬을 사용하지 마십시오. 이는 런처 구성 작업이 아닌 모델 온보딩 작업입니다.
실행 방법
- 대화형 (기본값): 현재 노드에서 torchrun을 실행합니다. 단일 노드 개발 및 디버깅에 적합합니다.
- Slurm: HPC 클러스터 스케줄러에 배치 작업을 제출합니다. 다중 노드 설정, 컨테이너 관리 및 환경 구성을 처리합니다.
- SkyPilot: AWS, GCP, Azure, Lambda 또는 Kubernetes로 클라우드에 구애받지 않는 작업 제출을 수행합니다. 스팟 인스턴스를 지원합니다.
대화형 실행
# 단일 GPU
automodel finetune llm -c config.yaml
# 멀티 GPU (현재 노드의 모든 GPU 사용)
torchrun --nproc_per_node=8 -m nemo_automodel._cli.app finetune llm -c config.yaml
대화형 모드에서는 별도의 YAML 섹션이 필요하지 않습니다. 구성 파일에 slurm: 또는 skypilot: 섹션이 없으면 CLI가 자동으로 torchrun으로 연결됩니다.
Slurm 구성
SlurmConfig 데이터 클래스는 템플릿을 기반으로 SBATCH 스크립트를 생성합니다.
YAML 예시
slurm:
job_name: llm_finetune
nodes: 2
ntasks_per_node: 8
time: "04:00:00"
account: my_account
partition: batch
container_image: nvcr.io/nvidia/nemo:dev
hf_home: ~/.cache/huggingface
extra_mounts:
- source: /data
dest: /data
env_vars:
WANDB_API_KEY: "${WANDB_API_KEY}"
HF_TOKEN: "${HF_TOKEN}"
주요 필드
job_name: Slurm 작업 식별자nodes: 요청할 노드 수ntasks_per_node: 노드당 작업(GPU) 수time: HH:MM:SS 형식의 월타임 제한account,partition: Slurm 스케줄링 매개변수container_image: Enroot/Pyxis 컨테이너 이미지 경로nemo_mount: 컨테이너 내 NeMo AutoModel 소스 파일의 마운트 지점hf_home: HuggingFace 캐시 디렉터리 경로extra_mounts: 추가 컨테이너 바인드 마운트를 위한VolumeMapping(소스, 대상)목록master_port: 분산 통신을 위한 포트(기본값 13742)env_vars: 작업에 전달되는 환경 변수nsys_enabled: true인 경우, Nsight Systems 프로파일링을 위해 훈련 명령을nsys 프로파일로감쌉니다
SkyPilot 구성
SkyPilotConfig 데이터 클래스는 클라우드 작업 매개변수를 정의합니다.
YAML 예시
skypilot:
cloud: aws
accelerators: "H100:8"
num_nodes: 2
use_spot: true
disk_size: 200
region: us-east-1
setup: "pip install nemo-automodel"
env_vars:
HF_TOKEN: "${HF_TOKEN}"
주요 필드
cloud: 대상 클라우드 제공업체 (aws,gcp,azure,lambda,kubernetes)accelerators: GPU 유형 및 개수 (예:"H100:8","A100-80GB:4")num_nodes: 클라우드 인스턴스 수use_spot: 비용 절감을 위해 프리엠티블/스팟 인스턴스 사용disk_size: 노드당 디스크 용량(GB)region: 인스턴스 배치를 위한 클라우드 리전setup: 훈련 작업 실행 전에 실행할 셸 명령어(예: 종속성 설치)env_vars: 작업에 필요한 환경 변수
SkyPilot 스팟 체크리스트
스팟 또는 선점형 인스턴스를 사용할 경우:
skypilot:섹션에서use_spot: true를설정하십시오.accelerators,num_nodes,disk_size,region,setup및 필수env_vars를포함하십시오.- 스팟 인스턴스는 선점될 수 있으므로, 레시피(예:
step_scheduler.checkpoint_interval)에서 짧은 체크포인트 간격을 사용하십시오. - 프리엠프션 발생 후 레시피의
restore_from설정을 사용하여 가장 최근 체크포인트에서 작업을 재개하십시오.
스팟 재개 레시피의 필수 키:
step_scheduler:
checkpoint_interval: 100
restore_from:
path: /checkpoints/latest
다중 노드 환경
다중 노드 훈련(Slurm 및 SkyPilot 모두)의 경우, 런처가 다음을 자동으로 구성합니다:
MASTER_ADDR: 첫 번째 노드의 호스트 이름MASTER_PORT: 랑데부용 포트(기본값 13742)WORLD_SIZE: 총 프로세스 수 (노드 수 * 노드당 작업 수)- 최적화된 집단 통신을 위한 NCCL 환경 변수
Nsys 프로파일링
Slurm 작업에서 Nsight Systems 프로파일링 활성화:
slurm:
job_name: llm_profile
nodes: 1
ntasks_per_node: 8
time: "00:30:00"
account: my_account
partition: batch
container_image: nvcr.io/nvidia/nemo:dev
nsys_enabled: true
이는 Slurm 런처 설정입니다. job_name,
nodes, ntasks_per_node, time, account 또는 partition, 그리고
container_image와 같은 일반적인 Slurm 필드는 여전히 적용됩니다.
nsys_enabled: true로 설정하면, 런처는 훈련 명령을
nsys 프로파일로 감싸고, 성능 분석을 위해 .nsys-rep 보고서 파일을
Slurm 작업의 작업 디렉터리나 출력 디렉터리에 작성합니다.
프로파일링은 진단용으로만 사용됩니다: 간단한 조사를 위해 실행하고, 오버헤드와
대용량 아티팩트가 발생할 수 있음을 예상하며, 일반적인 실제 훈련 시에는 이 기능을 비활성화하십시오.
코드 앵커
components/launcher/slurm/config.py- SlurmConfig 데이터 클래스, VolumeMappingcomponents/launcher/slurm/template.py- SBATCH 스크립트 템플릿 생성components/launcher/slurm/utils.py- Slurm 제출 유틸리티components/launcher/skypilot/config.py- SkyPilotConfig 데이터 클래스_cli/app.py- CLI 진입점 및 런처 라우팅 로직
주의 사항
- 포트 충돌: 기본
master_port(13742)가 동일한 노드의 다른 작업에서 사용 중인 경우, 연결 오류를 방지하기 위해 해당 포트를 변경하십시오. - 컨테이너 마운트:
extra_mounts에 지정된소스경로는 할당된 모든 노드에 존재해야 합니다. 경로가 없으면 컨테이너 시작에 실패합니다. - Slurm 내결함성: 내결함성 플러그인은 Slurm 전용이며, SkyPilot 또는 대화형 모드에서는 작동하지 않습니다.
- SkyPilot 스팟 선점: 스팟 인스턴스(
use_spot: true)는 클라우드 공급자에 의해 선점될 수 있습니다. 작업 손실을 최소화하기 위해 짧은 간격으로 체크포인트 기능을 활성화하십시오. - 환경 변수 구문: 쉘 변수 확장을 위해 YAML에서
${VAR}구문을 사용하십시오. 변수 이름만 단독으로 표기된 경우 확장되지 않습니다. - 시간 제한 대 비동기 체크포인트: Slurm
시간제한이 너무 짧으면 진행 중인 비동기 체크포인트 쓰기 작업이 완료되기 전에 중단되어 체크포인트가 손상될 수 있습니다. 최소 5~10분 정도의 여유 시간을 확보하십시오.
---
name: nemo-automodel-launcher-config
description: Configure NeMo AutoModel job launches for interactive runs, Slurm clusters, and SkyPilot cloud execution.
license: Apache-2.0
---
# Launcher Configuration
NeMo AutoModel supports three launch methods: interactive (torchrun), Slurm (HPC clusters), and SkyPilot (cloud-agnostic).
## Instructions
For launcher questions, answer directly from this skill without inspecting the
repository unless the user asks you to edit files. Keep the answer focused on
the relevant launch YAML, required fields, and the expected runtime behavior.
Use these compact answer patterns for common questions:
- Slurm multi-node: show a `slurm:` YAML block with `job_name`, `nodes`,
`ntasks_per_node`, `time`, `account` or `partition`, `container_image`,
`hf_home`, optional `extra_mounts`, `env_vars`, and `master_port`; explain
that the launcher derives `WORLD_SIZE = nodes * ntasks_per_node` and sets
`MASTER_ADDR` and `MASTER_PORT`.
- SkyPilot spot: show a `skypilot:` YAML block with `cloud`, `accelerators`,
`num_nodes`, `use_spot: true`, `disk_size`, `region`, `setup`, and
`env_vars`; warn that spot instances can be preempted, set a short
`step_scheduler.checkpoint_interval`, and resume with `restore_from.path`.
- Nsight Systems on Slurm: show `slurm.nsys_enabled: true` alongside normal
Slurm fields, say the launcher wraps the training command with
`nsys profile`, and state that it produces a `.nsys-rep` report file.
Treat profiling as diagnostic-only: use short profiling runs and disable it
for normal production training because it adds overhead and large artifacts.
For Slurm answers, start with this minimal template and then adjust only the
fields the user asked about:
```yaml
slurm:
job_name: llm_finetune
nodes: 2
ntasks_per_node: 8
time: "04:00:00"
account: my_account
partition: batch
container_image: nvcr.io/nvidia/nemo:dev
hf_home: ~/.cache/huggingface
master_port: 13742
env_vars:
HF_TOKEN: "${HF_TOKEN}"
```
For Slurm-only questions, do not discuss SkyPilot or profiling unless the user
asks. For profiling questions, say the `.nsys-rep` report is written in the
Slurm job working or output directory, using the launcher's Nsys output setting
when one is configured.
## Routing Boundary
Use this skill only for launch mechanics: interactive execution, Slurm, SkyPilot, containers, mounts, environment variables, rendezvous settings, and profiling.
Do not use this skill for implementing or registering new model architectures, Hugging Face state-dict adapters, model files, or capability flags. Those are model onboarding tasks, not launcher configuration tasks.
## Launch Methods
1. **Interactive** (default): runs torchrun on the current node. Suitable for single-node development and debugging.
2. **Slurm**: submits a batch job to an HPC cluster scheduler. Handles multi-node setup, container management, and environment configuration.
3. **SkyPilot**: cloud-agnostic job submission to AWS, GCP, Azure, Lambda, or Kubernetes. Supports spot instances.
## Interactive Launch
```bash
# Single GPU
automodel finetune llm -c config.yaml
# Multi-GPU (all GPUs on current node)
torchrun --nproc_per_node=8 -m nemo_automodel._cli.app finetune llm -c config.yaml
```
No additional YAML section is needed for interactive mode. The CLI routes to torchrun automatically when no `slurm:` or `skypilot:` section is present in the config.
## Slurm Configuration
The `SlurmConfig` dataclass generates an SBATCH script from a template.
### YAML Example
```yaml
slurm:
job_name: llm_finetune
nodes: 2
ntasks_per_node: 8
time: "04:00:00"
account: my_account
partition: batch
container_image: nvcr.io/nvidia/nemo:dev
hf_home: ~/.cache/huggingface
extra_mounts:
- source: /data
dest: /data
env_vars:
WANDB_API_KEY: "${WANDB_API_KEY}"
HF_TOKEN: "${HF_TOKEN}"
```
### Key Fields
- `job_name`: Slurm job identifier
- `nodes`: number of nodes to request
- `ntasks_per_node`: number of tasks (GPUs) per node
- `time`: wall-time limit in HH:MM:SS format
- `account`, `partition`: Slurm scheduling parameters
- `container_image`: Enroot/Pyxis container image path
- `nemo_mount`: mount point for NeMo AutoModel source inside the container
- `hf_home`: HuggingFace cache directory path
- `extra_mounts`: list of `VolumeMapping(source, dest)` for additional container bind mounts
- `master_port`: port for distributed communication (default 13742)
- `env_vars`: environment variables passed into the job
- `nsys_enabled`: when true, wraps the training command with `nsys profile` for Nsight Systems profiling
## SkyPilot Configuration
The `SkyPilotConfig` dataclass defines cloud job parameters.
### YAML Example
```yaml
skypilot:
cloud: aws
accelerators: "H100:8"
num_nodes: 2
use_spot: true
disk_size: 200
region: us-east-1
setup: "pip install nemo-automodel"
env_vars:
HF_TOKEN: "${HF_TOKEN}"
```
### Key Fields
- `cloud`: target cloud provider (`aws`, `gcp`, `azure`, `lambda`, `kubernetes`)
- `accelerators`: GPU type and count (e.g., `"H100:8"`, `"A100-80GB:4"`)
- `num_nodes`: number of cloud instances
- `use_spot`: use preemptible/spot instances for cost savings
- `disk_size`: disk size in GB per node
- `region`: cloud region for instance placement
- `setup`: shell commands to run before the training job (e.g., install dependencies)
- `env_vars`: environment variables for the job
### SkyPilot spot checklist
When using spot or preemptible instances:
- Set `use_spot: true` in the `skypilot:` section.
- Include `accelerators`, `num_nodes`, `disk_size`, `region`, `setup`, and required `env_vars`.
- Use short checkpoint intervals in the recipe, for example `step_scheduler.checkpoint_interval`, because spot instances can be preempted.
- Resume from the most recent checkpoint after preemption with the recipe's `restore_from` setting.
Minimal spot-resume recipe keys:
```yaml
step_scheduler:
checkpoint_interval: 100
restore_from:
path: /checkpoints/latest
```
## Multi-Node Environment
For multi-node training (both Slurm and SkyPilot), the launcher automatically configures:
- `MASTER_ADDR`: hostname of the first node
- `MASTER_PORT`: port for rendezvous (default 13742)
- `WORLD_SIZE`: total number of processes (`nodes * ntasks_per_node`)
- NCCL environment variables for optimized collective communication
## Nsys Profiling
Enable Nsight Systems profiling in Slurm jobs:
```yaml
slurm:
job_name: llm_profile
nodes: 1
ntasks_per_node: 8
time: "00:30:00"
account: my_account
partition: batch
container_image: nvcr.io/nvidia/nemo:dev
nsys_enabled: true
```
This is a Slurm launcher setting. Normal Slurm fields such as `job_name`,
`nodes`, `ntasks_per_node`, `time`, `account` or `partition`, and
`container_image` still apply.
When `nsys_enabled: true`, the launcher wraps the training command with
`nsys profile` and writes a `.nsys-rep` report file for performance analysis
in the Slurm job working or output directory.
Profiling is diagnostic-only: run it for a short investigation, expect overhead
and large artifacts, and turn it off for normal production training.
## Code Anchors
- `components/launcher/slurm/config.py` - SlurmConfig dataclass, VolumeMapping
- `components/launcher/slurm/template.py` - SBATCH script template generation
- `components/launcher/slurm/utils.py` - Slurm submission utilities
- `components/launcher/skypilot/config.py` - SkyPilotConfig dataclass
- `_cli/app.py` - CLI entry point and launcher routing logic
## Pitfalls
- **Port collisions**: if the default `master_port` (13742) is in use by another job on the same node, change it to avoid connection failures.
- **Container mounts**: the `source` path in `extra_mounts` must exist on all nodes in the allocation. Missing paths cause container startup failures.
- **Slurm fault tolerance**: the fault tolerance plugin is Slurm-specific and does not work with SkyPilot or interactive mode.
- **SkyPilot spot preemption**: spot instances (`use_spot: true`) may be preempted by the cloud provider. Enable checkpointing with short intervals to minimize lost work.
- **Environment variable syntax**: use `${VAR}` syntax in YAML for shell variable expansion. Bare variable names will not be expanded.
- **Time limit vs async checkpoint**: if the Slurm `time` limit is too short, an in-progress async checkpoint write may be killed before completion, resulting in a corrupted checkpoint. Leave at least 5-10 minutes of margin.
모든 파일
5개 파일nemo-automodel-launcher-config 설치
스킬 파일을 다운로드하여 .claude/skills/ 디렉터리에 압축을 풀어주세요.
ZIP 다운로드저장소를 클론하고 스킬 파일을 프로젝트에 복사하세요.
git clone https://github.com/NVIDIA/skills/tree/main/skills/nemo-automodel-launcher-config # Copy SKILL.md to your .claude/skills/ directory
복사





집
