選項
首頁首頁 Skill 開發營運和 CI/CD nemo-automodel-launcher-config

nemo-automodel-launcher-config

NVIDIA/skills NVIDIA/skills

設定 NeMo AutoModel 工作執行,以支援互動式執行、Slurm 叢集及 SkyPilot 雲端執行。

...展開全部
1
更新時間 2026-09-28

啟動器設定

NeMo AutoModel 支援三種啟動方式:互動式(torchrun)、Slurm(HPC 叢集)以及 SkyPilot(雲端中立)。

操作說明

若遇到關於啟動器的問題,請直接透過此技能進行回覆,無需檢視 儲存庫,除非使用者要求您編輯檔案。請將回覆重點放在 相關的啟動 YAML 檔案、必填欄位以及預期的執行時行為上。

針對常見問題,請使用以下簡潔的回答範本:

  • Slurm 多節點:展示包含job_name、nodes、YAML 區塊,內容包含job_name、nodes、 ntasks_per_node、time、account或partition、container_image、 hf_home、可選的extra_mounts、env_vars 以及master_port;並說明 啟動器會推導出WORLD_SIZE = nodes * ntasks_per_node,並設定 MASTER_ADDR和MASTER_PORT。
  • SkyPilot 按需實例:展示一個包含cloud、accelerators、 num_nodes、use_spot: true、disk_size、region、setup 以及 env_vars 的skypilot:YAML 區塊;提醒按需實例可能被預先終止,設定一個短暫的 step_scheduler.checkpoint_interval,並透過restore_from.path 恢復運作。
  • Nsight Systems 在 Slurm 上:展示slurm.nsys_enabled: true以及標準的 Slurm 欄位,說明啟動器會將訓練指令包裹在 nsys 剖析中,並指出其會產生一個.nsys-rep報告檔。 將效能分析視為僅供診斷之用:採用短暫的效能分析執行,並在 一般生產環境訓練中停用此功能,因為它會增加開銷並產生大量偽影。

針對 Slurm 的問題,請先使用此最小範本,然後僅調整 使用者詢問的欄位:

slurm:
  job_name: llm_finetune
  nodes: 2
  ntasks_per_node: 8
  time: "04:00:00"
  account: my_account
  partition: batch
  container_image: nvcr.io/nvidia/nemo:dev
  hf_home: ~/.cache/huggingface
  master_port: 13742
  env_vars:
    HF_TOKEN: "${HF_TOKEN}"

若問題僅涉及 Slurm,除非使用者主動詢問,否則請勿討論 SkyPilot 或效能分析。 針對效能分析相關問題,請說明.nsys-rep報告會寫入 Slurm 工作執行或輸出目錄中, 並在已設定的情況下,使用啟動器的 Nsys 輸出設定。

問題導向範圍

此技能僅適用於啟動機制相關事宜:互動式執行、Slurm、SkyPilot、容器、掛載、環境變數、會合點設定及效能分析。

請勿將此技能用於實作或註冊新的模型架構、Hugging Face 狀態字典轉接器、模型檔案或功能標誌。這些屬於模型導入任務,而非啟動器配置任務。

啟動方式

  1. 互動式(預設):在當前節點上執行 torchrun。適用於單節點開發與除錯。
  2. Slurm:將批次工作提交至 HPC 叢集排程器。負責處理多節點設定、容器管理及環境配置。
  3. SkyPilot:向 AWS、GCP、Azure、Lambda 或 Kubernetes 提交雲端平台中立的作業。支援按需實例。

互動式啟動

# 單一 GPU
automodel finetune llm -c config.yaml

# 多 GPU(當前節點上的所有 GPU)
torchrun --nproc_per_node=8 -m nemo_automodel._cli.app finetune llm -c config.yaml

互動模式下無需額外的 YAML 區段。當設定檔中不存在slurm:或skypilot:區段時,CLI 會自動轉由 torchrun 執行。

Slurm 配置

SlurmConfig資料類別會根據範本生成 SBATCH 腳本。

YAML 範例

slurm:
  job_name: llm_finetune
  nodes: 2
  ntasks_per_node: 8
  time: "04:00:00"
  account: my_account
  partition: batch
  container_image: nvcr.io/nvidia/nemo:dev
  hf_home: ~/.cache/huggingface
  extra_mounts:
    - 來源: /data
      目標: /data
  環境變數:
    WANDB_API_KEY: "${WANDB_API_KEY}"
    HF_TOKEN: "${HF_TOKEN}"

關鍵欄位

  • job_name:Slurm 工作識別碼
  • nodes:要請求的節點數
  • ntasks_per_node:每個節點的任務(GPU)數量
  • time:以 HH:MM:SS 格式表示的牆上時間限制
  • account,partition:Slurm 排程參數
  • container_image:Enroot/Pyxis 容器映像路徑
  • nemo_mount:容器內 NeMo AutoModel 原始碼的掛載點
  • hf_home:HuggingFace 快取目錄路徑
  • extra_mounts:用於額外容器綁定掛載的VolumeMapping(來源, 目標)清單
  • master_port:分散式通訊的埠號(預設為 13742)
  • env_vars:傳入工作(job)的環境變數
  • nsys_enabled:若為 true,則會使用nsys 配置檔包裹訓練指令,以便進行 Nsight Systems 效能分析

SkyPilot 配置

SkyPilotConfig資料類別用以定義雲端工作任務的參數。

YAML 範例

skypilot:
  cloud: aws
  accelerators: "H100:8"
  num_nodes: 2
  use_spot: true
  disk_size: 200
  region: us-east-1
  setup: "pip install nemo-automodel"
  env_vars:
    HF_TOKEN: "${HF_TOKEN}"

關鍵欄位

  • cloud:目標雲端服務供應商(aws、gcp、azure、lambda、kubernetes)
  • accelerators:GPU 類型與數量(例如:「H100:8」、「A100-80GB:4」)
  • num_nodes:雲端執行個體的數量
  • use_spot:使用可預先終止/競價型實例以節省成本
  • disk_size:每個節點的磁碟容量(以 GB 為單位)
  • region:用於放置實例的雲端區域
  • setup:訓練工作開始前需執行的 shell 指令(例如:安裝依賴項)
  • env_vars:工作所需的环境變數

SkyPilot 限時競標檢查清單

使用按需或可預先終止實例時:

  • 請在skypilot:區段中設定use_spot: true。
  • 請包含accelerators、num_nodes、disk_size、region、setup 以及所需的env_vars。
  • 由於隨機競標實例可能被預先終止,請在配方中設定較短的檢查點間隔,例如step_scheduler.checkpoint_interval。
  • 在遭預先終止後,透過配方中的 `restore_from` 設定,從最近的檢查點恢復執行。

Spot 恢復配方所需的最低鍵值:

step_scheduler:
  checkpoint_interval: 100

restore_from:
  path: /checkpoints/latest

多節點環境

對於多節點訓練(包括 Slurm 和 SkyPilot),啟動器會自動配置:

  • MASTER_ADDR:第一個節點的主機名稱
  • MASTER_PORT:會合埠(預設 13742)
  • WORLD_SIZE:總進程數(節點數 × 每個節點的任務數)
  • 用於優化集體通訊的 NCCL 環境變數

Nsys 效能分析

在 Slurm 工作任務中啟用 Nsight Systems 效能分析:

slurm:
  job_name: llm_profile
  nodes: 1
  ntasks_per_node: 8
  time: "00:30:00"
  account: my_account
  partition: batch
  container_image: nvcr.io/nvidia/nemo:dev
  nsys_enabled: true

這是 Slurm 啟動器的設定。一般的 Slurm 欄位,例如job_name、 nodes、ntasks_per_node、time、account或partition,以及 container_image仍適用。

當nsys_enabled: true 時,啟動器會將訓練指令以 nsys 剖析模式封裝,並在 Slurm 工作目錄或輸出目錄中 寫入.nsys-rep報告檔案以供效能分析。 剖析功能僅供診斷使用:僅在短期調查時執行,需預期會產生 額外開銷及大量分析產出,並應在正常生產訓練時關閉此功能。

程式碼錨點

  • components/launcher/slurm/config.py- SlurmConfig 資料類別、VolumeMapping
  • components/launcher/slurm/template.py- SBATCH 腳本範本產生
  • components/launcher/slurm/utils.py- Slurm 提交工具
  • components/launcher/skypilot/config.py- SkyPilotConfig 資料類別
  • _cli/app.py- 命令列介面 (CLI) 入口點與啟動器路由邏輯

注意事項

  • 埠號衝突:若預設的master_port(13742)正被同一節點上的其他工作佔用,請變更該埠號以避免連線失敗。
  • 容器掛載:extra_mounts中的來源路徑必須存在於分配中的所有節點上。若路徑不存在,將導致容器啟動失敗。
  • Slurm 容錯機制:此容錯外掛程式專為 Slurm 設計,不適用於 SkyPilot 或互動模式。
  • SkyPilot 按需實例預先終止:按需實例(use_spot: true)可能會被雲端供應商預先終止。請啟用短間隔的檢查點功能,以將工作損失降至最低。
  • 環境變數語法:請在 YAML 中使用${VAR}語法進行 shell 變數擴展。未帶前綴的變數名稱將不會被擴展。
  • 時間限制與非同步檢查點:若 Slurm的時間限制過短,進行中的非同步檢查點寫入作業可能會在完成前被終止,導致檢查點損毀。請預留至少 5 至 10 分鐘的緩衝時間。
在 GitHub 上查看
---
name: nemo-automodel-launcher-config
description: Configure NeMo AutoModel job launches for interactive runs, Slurm clusters, and SkyPilot cloud execution.
license: Apache-2.0
---

# Launcher Configuration

NeMo AutoModel supports three launch methods: interactive (torchrun), Slurm (HPC clusters), and SkyPilot (cloud-agnostic).

## Instructions

For launcher questions, answer directly from this skill without inspecting the
repository unless the user asks you to edit files. Keep the answer focused on
the relevant launch YAML, required fields, and the expected runtime behavior.

Use these compact answer patterns for common questions:

- Slurm multi-node: show a `slurm:` YAML block with `job_name`, `nodes`,
  `ntasks_per_node`, `time`, `account` or `partition`, `container_image`,
  `hf_home`, optional `extra_mounts`, `env_vars`, and `master_port`; explain
  that the launcher derives `WORLD_SIZE = nodes * ntasks_per_node` and sets
  `MASTER_ADDR` and `MASTER_PORT`.
- SkyPilot spot: show a `skypilot:` YAML block with `cloud`, `accelerators`,
  `num_nodes`, `use_spot: true`, `disk_size`, `region`, `setup`, and
  `env_vars`; warn that spot instances can be preempted, set a short
  `step_scheduler.checkpoint_interval`, and resume with `restore_from.path`.
- Nsight Systems on Slurm: show `slurm.nsys_enabled: true` alongside normal
  Slurm fields, say the launcher wraps the training command with
  `nsys profile`, and state that it produces a `.nsys-rep` report file.
  Treat profiling as diagnostic-only: use short profiling runs and disable it
  for normal production training because it adds overhead and large artifacts.

For Slurm answers, start with this minimal template and then adjust only the
fields the user asked about:

```yaml
slurm:
  job_name: llm_finetune
  nodes: 2
  ntasks_per_node: 8
  time: "04:00:00"
  account: my_account
  partition: batch
  container_image: nvcr.io/nvidia/nemo:dev
  hf_home: ~/.cache/huggingface
  master_port: 13742
  env_vars:
    HF_TOKEN: "${HF_TOKEN}"
```

For Slurm-only questions, do not discuss SkyPilot or profiling unless the user
asks. For profiling questions, say the `.nsys-rep` report is written in the
Slurm job working or output directory, using the launcher's Nsys output setting
when one is configured.

## Routing Boundary

Use this skill only for launch mechanics: interactive execution, Slurm, SkyPilot, containers, mounts, environment variables, rendezvous settings, and profiling.

Do not use this skill for implementing or registering new model architectures, Hugging Face state-dict adapters, model files, or capability flags. Those are model onboarding tasks, not launcher configuration tasks.

## Launch Methods

1. **Interactive** (default): runs torchrun on the current node. Suitable for single-node development and debugging.
2. **Slurm**: submits a batch job to an HPC cluster scheduler. Handles multi-node setup, container management, and environment configuration.
3. **SkyPilot**: cloud-agnostic job submission to AWS, GCP, Azure, Lambda, or Kubernetes. Supports spot instances.

## Interactive Launch

```bash
# Single GPU
automodel finetune llm -c config.yaml

# Multi-GPU (all GPUs on current node)
torchrun --nproc_per_node=8 -m nemo_automodel._cli.app finetune llm -c config.yaml
```

No additional YAML section is needed for interactive mode. The CLI routes to torchrun automatically when no `slurm:` or `skypilot:` section is present in the config.

## Slurm Configuration

The `SlurmConfig` dataclass generates an SBATCH script from a template.

### YAML Example

```yaml
slurm:
  job_name: llm_finetune
  nodes: 2
  ntasks_per_node: 8
  time: "04:00:00"
  account: my_account
  partition: batch
  container_image: nvcr.io/nvidia/nemo:dev
  hf_home: ~/.cache/huggingface
  extra_mounts:
    - source: /data
      dest: /data
  env_vars:
    WANDB_API_KEY: "${WANDB_API_KEY}"
    HF_TOKEN: "${HF_TOKEN}"
```

### Key Fields

- `job_name`: Slurm job identifier
- `nodes`: number of nodes to request
- `ntasks_per_node`: number of tasks (GPUs) per node
- `time`: wall-time limit in HH:MM:SS format
- `account`, `partition`: Slurm scheduling parameters
- `container_image`: Enroot/Pyxis container image path
- `nemo_mount`: mount point for NeMo AutoModel source inside the container
- `hf_home`: HuggingFace cache directory path
- `extra_mounts`: list of `VolumeMapping(source, dest)` for additional container bind mounts
- `master_port`: port for distributed communication (default 13742)
- `env_vars`: environment variables passed into the job
- `nsys_enabled`: when true, wraps the training command with `nsys profile` for Nsight Systems profiling

## SkyPilot Configuration

The `SkyPilotConfig` dataclass defines cloud job parameters.

### YAML Example

```yaml
skypilot:
  cloud: aws
  accelerators: "H100:8"
  num_nodes: 2
  use_spot: true
  disk_size: 200
  region: us-east-1
  setup: "pip install nemo-automodel"
  env_vars:
    HF_TOKEN: "${HF_TOKEN}"
```

### Key Fields

- `cloud`: target cloud provider (`aws`, `gcp`, `azure`, `lambda`, `kubernetes`)
- `accelerators`: GPU type and count (e.g., `"H100:8"`, `"A100-80GB:4"`)
- `num_nodes`: number of cloud instances
- `use_spot`: use preemptible/spot instances for cost savings
- `disk_size`: disk size in GB per node
- `region`: cloud region for instance placement
- `setup`: shell commands to run before the training job (e.g., install dependencies)
- `env_vars`: environment variables for the job

### SkyPilot spot checklist

When using spot or preemptible instances:

- Set `use_spot: true` in the `skypilot:` section.
- Include `accelerators`, `num_nodes`, `disk_size`, `region`, `setup`, and required `env_vars`.
- Use short checkpoint intervals in the recipe, for example `step_scheduler.checkpoint_interval`, because spot instances can be preempted.
- Resume from the most recent checkpoint after preemption with the recipe's `restore_from` setting.

Minimal spot-resume recipe keys:

```yaml
step_scheduler:
  checkpoint_interval: 100

restore_from:
  path: /checkpoints/latest
```

## Multi-Node Environment

For multi-node training (both Slurm and SkyPilot), the launcher automatically configures:
- `MASTER_ADDR`: hostname of the first node
- `MASTER_PORT`: port for rendezvous (default 13742)
- `WORLD_SIZE`: total number of processes (`nodes * ntasks_per_node`)
- NCCL environment variables for optimized collective communication

## Nsys Profiling

Enable Nsight Systems profiling in Slurm jobs:

```yaml
slurm:
  job_name: llm_profile
  nodes: 1
  ntasks_per_node: 8
  time: "00:30:00"
  account: my_account
  partition: batch
  container_image: nvcr.io/nvidia/nemo:dev
  nsys_enabled: true
```

This is a Slurm launcher setting. Normal Slurm fields such as `job_name`,
`nodes`, `ntasks_per_node`, `time`, `account` or `partition`, and
`container_image` still apply.

When `nsys_enabled: true`, the launcher wraps the training command with
`nsys profile` and writes a `.nsys-rep` report file for performance analysis
in the Slurm job working or output directory.
Profiling is diagnostic-only: run it for a short investigation, expect overhead
and large artifacts, and turn it off for normal production training.

## Code Anchors

- `components/launcher/slurm/config.py` - SlurmConfig dataclass, VolumeMapping
- `components/launcher/slurm/template.py` - SBATCH script template generation
- `components/launcher/slurm/utils.py` - Slurm submission utilities
- `components/launcher/skypilot/config.py` - SkyPilotConfig dataclass
- `_cli/app.py` - CLI entry point and launcher routing logic

## Pitfalls

- **Port collisions**: if the default `master_port` (13742) is in use by another job on the same node, change it to avoid connection failures.
- **Container mounts**: the `source` path in `extra_mounts` must exist on all nodes in the allocation. Missing paths cause container startup failures.
- **Slurm fault tolerance**: the fault tolerance plugin is Slurm-specific and does not work with SkyPilot or interactive mode.
- **SkyPilot spot preemption**: spot instances (`use_spot: true`) may be preempted by the cloud provider. Enable checkpointing with short intervals to minimize lost work.
- **Environment variable syntax**: use `${VAR}` syntax in YAML for shell variable expansion. Bare variable names will not be expanded.
- **Time limit vs async checkpoint**: if the Slurm `time` limit is too short, an in-progress async checkpoint write may be killed before completion, resulting in a corrupted checkpoint. Leave at least 5-10 minutes of margin.

安裝 nemo-automodel-launcher-config

請下載並將技能檔案解壓縮至您的 .claude/skills/ 目錄中。

下載 ZIP

複製儲存庫並將技能檔案複製到您的專案中。

git clone https://github.com/NVIDIA/skills/tree/main/skills/nemo-automodel-launcher-config # Copy SKILL.md to your .claude/skills/ directory

複製 複製
快速設定: 將技能資料夾複製到 .claude/skills/ Claude 會自動偵測並使用該技能
儲存庫 NVIDIA/skills

相關技能

klingai-upgrade-migration
更新時間 2026-07-03
Verification & Quality Assurance
更新時間 2026-06-29
base44-cli
更新時間 2026-06-29
Railway CLI Management
更新時間 2026-07-02
OR