选项
首页首页 Skill 开发运营和 CI/CD nemo-automodel-launcher-config

nemo-automodel-launcher-config

NVIDIA/skills NVIDIA/skills

配置 NeMo AutoModel 任务的启动,以支持交互式运行、Slurm 集群以及 SkyPilot 云端执行。

...展开全部
1
更新时间 2026-09-28

启动器配置

NeMo AutoModel 支持三种启动方式:交互式(torchrun)、Slurm(HPC 集群)和 SkyPilot(云平台无关)。

操作指南

对于启动器相关的问题,请直接通过本技能进行回答,无需查看 代码库,除非用户要求您编辑文件。请将回答重点放在 相关的启动 YAML 文件、必填字段以及预期的运行时行为上。

针对常见问题,请使用以下简洁的回答模板:

  • Slurm 多节点:展示包含job_name、nodes、YAML 代码块,其中包含job_name、nodes、 ntasks_per_node、time、account或partition、container_image、 hf_home、可选的extra_mounts、env_vars 以及master_port;说明 启动器会推导出WORLD_SIZE = nodes * ntasks_per_node,并设置 MASTER_ADDR和MASTER_PORT。
  • SkyPilot 按需实例:展示一个包含cloud、accelerators、 num_nodes、use_spot: true、disk_size、region、setup 以及 env_vars 的 skypilot:YAML 代码块;提醒用户按需实例可能被抢占,设置一个较短的 step_scheduler.checkpoint_interval,并使用restore_from.path 恢复运行。
  • Nsight Systems 在 Slurm 上:展示slurm.nsys_enabled: true以及常规 Slurm 字段,说明启动器会将训练命令包裹在 nsys 剖析中,并指出它会生成一个.nsys-rep报告文件。 将性能分析视为仅用于诊断:使用短时间的性能分析运行,并在 常规生产训练中禁用该功能,因为它会增加开销并产生大量伪影。

对于 Slurm 相关问题,请从以下简约模板开始,然后仅调整 用户询问的字段:

slurm:
  job_name: llm_finetune
  nodes: 2
  ntasks_per_node: 8
  time: "04:00:00"
  account: my_account
  partition: batch
  container_image: nvcr.io/nvidia/nemo:dev
  hf_home: ~/.cache/huggingface
  master_port: 13742
  env_vars:
    HF_TOKEN: "${HF_TOKEN}"

对于仅涉及 Slurm 的问题,除非用户 主动询问,否则不要讨论 SkyPilot 或性能分析。对于性能分析相关的问题,应说明.nsys-rep报告会写入 Slurm 作业的工作目录或输出目录中,如果已配置,则使用启动器的 Nsys 输出设置 。

问题分流边界

仅将此技能用于启动机制相关问题:交互式执行、Slurm、SkyPilot、容器、挂载、环境变量、会合设置以及性能分析。

请勿将此技能用于实现或注册新的模型架构、Hugging Face 状态字典适配器、模型文件或功能标志。这些属于模型接入任务,而非启动器配置任务。

启动方法

  1. 交互式(默认):在当前节点上运行 torchrun。适用于单节点开发和调试。
  2. Slurm:向 HPC 集群调度器提交批处理任务。负责多节点设置、容器管理和环境配置。
  3. SkyPilot:向 AWS、GCP、Azure、Lambda 或 Kubernetes 提交与云平台无关的任务。支持竞价实例。

交互式启动

# 单 GPU
automodel finetune llm -c config.yaml

# 多 GPU(当前节点上的所有 GPU)
torchrun --nproc_per_node=8 -m nemo_automodel._cli.app finetune llm -c config.yaml

交互模式下无需额外添加 YAML 部分。当配置中不存在slurm:或skypilot:部分时,CLI 会自动调用 torchrun。

Slurm 配置

SlurmConfig数据类会根据模板生成 SBATCH 脚本。

YAML 示例

slurm:
  job_name: llm_finetune
  nodes: 2
  ntasks_per_node: 8
  time: "04:00:00"
  account: my_account
  partition: batch
  container_image: nvcr.io/nvidia/nemo:dev
  hf_home: ~/.cache/huggingface
  extra_mounts:
    - source: /data
      dest: /data
  env_vars:
    WANDB_API_KEY: "${WANDB_API_KEY}"
    HF_TOKEN: "${HF_TOKEN}"

关键字段

  • job_name:Slurm 作业标识符
  • nodes: 请求的节点数量
  • ntasks_per_node: 每个节点上的任务(GPU)数量
  • time:墙时限制,格式为 HH:MM:SS
  • account,partition:Slurm 调度参数
  • container_image:Enroot/Pyxis 容器镜像路径
  • nemo_mount:容器内 NeMo AutoModel 源文件的挂载点
  • hf_home:HuggingFace 缓存目录路径
  • extra_mounts:用于额外容器绑定挂载的VolumeMapping(源, 目标)列表
  • master_port:分布式通信端口(默认 13742)
  • env_vars:传递给任务的环境变量
  • nsys_enabled:当值为 true 时,会使用nsys 配置文件对训练命令进行封装,以便进行 Nsight Systems 性能分析

SkyPilot 配置

SkyPilotConfig数据类用于定义云任务参数。

YAML 示例

skypilot:
  cloud: aws
  accelerators: "H100:8"
  num_nodes: 2
  use_spot: true
  disk_size: 200
  region: us-east-1
  setup: "pip install nemo-automodel"
  env_vars:
    HF_TOKEN: "${HF_TOKEN}"

关键字段

  • cloud:目标云服务提供商(aws、gcp、azure、lambda、kubernetes)
  • accelerators:GPU 类型和数量(例如,“H100:8”、“A100-80GB:4”)
  • num_nodes:云实例数量
  • use_spot:使用可抢占型/竞价型实例以节省成本
  • disk_size:每个节点的磁盘大小(单位:GB)
  • region:用于实例部署的云区域
  • setup:训练任务开始前需执行的 shell 命令(例如,安装依赖项)
  • env_vars:任务所需的环境变量

SkyPilot 竞价实例检查清单

使用竞价型或可抢占型实例时:

  • 在skypilot:部分中将use_spot设置为true。
  • 请包含accelerators、num_nodes、disk_size、region、setup 以及必需的env_vars。
  • 在配方中使用较短的检查点间隔(例如step_scheduler.checkpoint_interval),因为竞价实例可能会被抢占。
  • 在被抢占后,通过配方中的restore_from设置从最近的检查点恢复。

Spot 恢复配方所需的最小配置项:

step_scheduler:
  checkpoint_interval: 100

restore_from:
  path: /checkpoints/latest

多节点环境

对于多节点训练(包括 Slurm 和 SkyPilot),启动器会自动配置:

  • MASTER_ADDR:第一个节点的主机名
  • MASTER_PORT:会合端口(默认 13742)
  • WORLD_SIZE:进程总数(节点数 * 每个节点的任务数)
  • 用于优化集体通信的 NCCL 环境变量

Nsys 性能分析

在 Slurm 作业中启用 Nsight Systems 性能分析:

slurm:
  job_name: llm_profile
  nodes: 1
  ntasks_per_node: 8
  time: "00:30:00"
  account: my_account
  partition: batch
  container_image: nvcr.io/nvidia/nemo:dev
  nsys_enabled: true

这是 Slurm 启动器的设置。常规的 Slurm 字段(如job_name、 nodes、ntasks_per_node、time、account或partition,以及 container_image)仍然适用。

当nsys_enabled: true 时,启动器会使用 nsys 配置文件对训练命令进行封装,并在 Slurm 作业的工作目录或输出目录中 生成用于性能分析的.nsys-rep报告文件。 性能分析仅用于诊断:仅在短期调查时运行,需注意会产生开销 和大量生成文件,在常规生产训练中应将其关闭。

代码锚点

  • components/launcher/slurm/config.py- SlurmConfig 数据类、VolumeMapping
  • components/launcher/slurm/template.py- SBATCH 脚本模板生成
  • components/launcher/slurm/utils.py- Slurm 提交实用工具
  • components/launcher/skypilot/config.py- SkyPilotConfig 数据类
  • _cli/app.py- CLI 入口点和启动器路由逻辑

注意事项

  • 端口冲突:如果默认的master_port(13742)正在被同一节点上的另一个作业使用,请将其更改以避免连接失败。
  • 容器挂载:extra_mounts中的源路径必须在分配中的所有节点上都存在。路径缺失会导致容器启动失败。
  • Slurm 容错:该容错插件专为 Slurm 设计,不适用于 SkyPilot 或交互模式。
  • SkyPilot 按需实例抢占:按需实例(use_spot: true)可能会被云服务商抢占。请启用短间隔检查点功能,以最大限度地减少工作数据丢失。
  • 环境变量语法:在 YAML 中使用${VAR}语法进行 shell 变量展开。仅使用变量名的形式将不会被展开。
  • 时间限制与异步检查点:如果 Slurm时间限制过短,正在进行的异步检查点写入操作可能会在完成前被终止,导致检查点损坏。请预留至少 5-10 分钟的余量。
在 GitHub 上查看
---
name: nemo-automodel-launcher-config
description: Configure NeMo AutoModel job launches for interactive runs, Slurm clusters, and SkyPilot cloud execution.
license: Apache-2.0
---

# Launcher Configuration

NeMo AutoModel supports three launch methods: interactive (torchrun), Slurm (HPC clusters), and SkyPilot (cloud-agnostic).

## Instructions

For launcher questions, answer directly from this skill without inspecting the
repository unless the user asks you to edit files. Keep the answer focused on
the relevant launch YAML, required fields, and the expected runtime behavior.

Use these compact answer patterns for common questions:

- Slurm multi-node: show a `slurm:` YAML block with `job_name`, `nodes`,
  `ntasks_per_node`, `time`, `account` or `partition`, `container_image`,
  `hf_home`, optional `extra_mounts`, `env_vars`, and `master_port`; explain
  that the launcher derives `WORLD_SIZE = nodes * ntasks_per_node` and sets
  `MASTER_ADDR` and `MASTER_PORT`.
- SkyPilot spot: show a `skypilot:` YAML block with `cloud`, `accelerators`,
  `num_nodes`, `use_spot: true`, `disk_size`, `region`, `setup`, and
  `env_vars`; warn that spot instances can be preempted, set a short
  `step_scheduler.checkpoint_interval`, and resume with `restore_from.path`.
- Nsight Systems on Slurm: show `slurm.nsys_enabled: true` alongside normal
  Slurm fields, say the launcher wraps the training command with
  `nsys profile`, and state that it produces a `.nsys-rep` report file.
  Treat profiling as diagnostic-only: use short profiling runs and disable it
  for normal production training because it adds overhead and large artifacts.

For Slurm answers, start with this minimal template and then adjust only the
fields the user asked about:

```yaml
slurm:
  job_name: llm_finetune
  nodes: 2
  ntasks_per_node: 8
  time: "04:00:00"
  account: my_account
  partition: batch
  container_image: nvcr.io/nvidia/nemo:dev
  hf_home: ~/.cache/huggingface
  master_port: 13742
  env_vars:
    HF_TOKEN: "${HF_TOKEN}"
```

For Slurm-only questions, do not discuss SkyPilot or profiling unless the user
asks. For profiling questions, say the `.nsys-rep` report is written in the
Slurm job working or output directory, using the launcher's Nsys output setting
when one is configured.

## Routing Boundary

Use this skill only for launch mechanics: interactive execution, Slurm, SkyPilot, containers, mounts, environment variables, rendezvous settings, and profiling.

Do not use this skill for implementing or registering new model architectures, Hugging Face state-dict adapters, model files, or capability flags. Those are model onboarding tasks, not launcher configuration tasks.

## Launch Methods

1. **Interactive** (default): runs torchrun on the current node. Suitable for single-node development and debugging.
2. **Slurm**: submits a batch job to an HPC cluster scheduler. Handles multi-node setup, container management, and environment configuration.
3. **SkyPilot**: cloud-agnostic job submission to AWS, GCP, Azure, Lambda, or Kubernetes. Supports spot instances.

## Interactive Launch

```bash
# Single GPU
automodel finetune llm -c config.yaml

# Multi-GPU (all GPUs on current node)
torchrun --nproc_per_node=8 -m nemo_automodel._cli.app finetune llm -c config.yaml
```

No additional YAML section is needed for interactive mode. The CLI routes to torchrun automatically when no `slurm:` or `skypilot:` section is present in the config.

## Slurm Configuration

The `SlurmConfig` dataclass generates an SBATCH script from a template.

### YAML Example

```yaml
slurm:
  job_name: llm_finetune
  nodes: 2
  ntasks_per_node: 8
  time: "04:00:00"
  account: my_account
  partition: batch
  container_image: nvcr.io/nvidia/nemo:dev
  hf_home: ~/.cache/huggingface
  extra_mounts:
    - source: /data
      dest: /data
  env_vars:
    WANDB_API_KEY: "${WANDB_API_KEY}"
    HF_TOKEN: "${HF_TOKEN}"
```

### Key Fields

- `job_name`: Slurm job identifier
- `nodes`: number of nodes to request
- `ntasks_per_node`: number of tasks (GPUs) per node
- `time`: wall-time limit in HH:MM:SS format
- `account`, `partition`: Slurm scheduling parameters
- `container_image`: Enroot/Pyxis container image path
- `nemo_mount`: mount point for NeMo AutoModel source inside the container
- `hf_home`: HuggingFace cache directory path
- `extra_mounts`: list of `VolumeMapping(source, dest)` for additional container bind mounts
- `master_port`: port for distributed communication (default 13742)
- `env_vars`: environment variables passed into the job
- `nsys_enabled`: when true, wraps the training command with `nsys profile` for Nsight Systems profiling

## SkyPilot Configuration

The `SkyPilotConfig` dataclass defines cloud job parameters.

### YAML Example

```yaml
skypilot:
  cloud: aws
  accelerators: "H100:8"
  num_nodes: 2
  use_spot: true
  disk_size: 200
  region: us-east-1
  setup: "pip install nemo-automodel"
  env_vars:
    HF_TOKEN: "${HF_TOKEN}"
```

### Key Fields

- `cloud`: target cloud provider (`aws`, `gcp`, `azure`, `lambda`, `kubernetes`)
- `accelerators`: GPU type and count (e.g., `"H100:8"`, `"A100-80GB:4"`)
- `num_nodes`: number of cloud instances
- `use_spot`: use preemptible/spot instances for cost savings
- `disk_size`: disk size in GB per node
- `region`: cloud region for instance placement
- `setup`: shell commands to run before the training job (e.g., install dependencies)
- `env_vars`: environment variables for the job

### SkyPilot spot checklist

When using spot or preemptible instances:

- Set `use_spot: true` in the `skypilot:` section.
- Include `accelerators`, `num_nodes`, `disk_size`, `region`, `setup`, and required `env_vars`.
- Use short checkpoint intervals in the recipe, for example `step_scheduler.checkpoint_interval`, because spot instances can be preempted.
- Resume from the most recent checkpoint after preemption with the recipe's `restore_from` setting.

Minimal spot-resume recipe keys:

```yaml
step_scheduler:
  checkpoint_interval: 100

restore_from:
  path: /checkpoints/latest
```

## Multi-Node Environment

For multi-node training (both Slurm and SkyPilot), the launcher automatically configures:
- `MASTER_ADDR`: hostname of the first node
- `MASTER_PORT`: port for rendezvous (default 13742)
- `WORLD_SIZE`: total number of processes (`nodes * ntasks_per_node`)
- NCCL environment variables for optimized collective communication

## Nsys Profiling

Enable Nsight Systems profiling in Slurm jobs:

```yaml
slurm:
  job_name: llm_profile
  nodes: 1
  ntasks_per_node: 8
  time: "00:30:00"
  account: my_account
  partition: batch
  container_image: nvcr.io/nvidia/nemo:dev
  nsys_enabled: true
```

This is a Slurm launcher setting. Normal Slurm fields such as `job_name`,
`nodes`, `ntasks_per_node`, `time`, `account` or `partition`, and
`container_image` still apply.

When `nsys_enabled: true`, the launcher wraps the training command with
`nsys profile` and writes a `.nsys-rep` report file for performance analysis
in the Slurm job working or output directory.
Profiling is diagnostic-only: run it for a short investigation, expect overhead
and large artifacts, and turn it off for normal production training.

## Code Anchors

- `components/launcher/slurm/config.py` - SlurmConfig dataclass, VolumeMapping
- `components/launcher/slurm/template.py` - SBATCH script template generation
- `components/launcher/slurm/utils.py` - Slurm submission utilities
- `components/launcher/skypilot/config.py` - SkyPilotConfig dataclass
- `_cli/app.py` - CLI entry point and launcher routing logic

## Pitfalls

- **Port collisions**: if the default `master_port` (13742) is in use by another job on the same node, change it to avoid connection failures.
- **Container mounts**: the `source` path in `extra_mounts` must exist on all nodes in the allocation. Missing paths cause container startup failures.
- **Slurm fault tolerance**: the fault tolerance plugin is Slurm-specific and does not work with SkyPilot or interactive mode.
- **SkyPilot spot preemption**: spot instances (`use_spot: true`) may be preempted by the cloud provider. Enable checkpointing with short intervals to minimize lost work.
- **Environment variable syntax**: use `${VAR}` syntax in YAML for shell variable expansion. Bare variable names will not be expanded.
- **Time limit vs async checkpoint**: if the Slurm `time` limit is too short, an in-progress async checkpoint write may be killed before completion, resulting in a corrupted checkpoint. Leave at least 5-10 minutes of margin.

安装 nemo-automodel-launcher-config

下载技能文件并将其解压到 .claude/skills/ 目录中。

下载ZIP

克隆仓库并复制技能文件到您的项目中。

git clone https://github.com/NVIDIA/skills/tree/main/skills/nemo-automodel-launcher-config # Copy SKILL.md to your .claude/skills/ directory

复制 复制
快速设置: 将技能文件夹复制到 .claude/skills/ Claude 将自动检测并使用该技能
仓库 NVIDIA/skills

相关技能

klingai-upgrade-migration
更新时间 2026-07-03
Verification & Quality Assurance
更新时间 2026-06-29
base44-cli
更新时间 2026-06-29
Railway CLI Management
更新时间 2026-07-02
OR