选项
首页首页 Skill 数据科学与机器学习 nemo-mbridge-perf-cpu-offloading

nemo-mbridge-perf-cpu-offloading

NVIDIA/skills NVIDIA/skills

配置并验证 Megatron Bridge 训练的 CPU 卸载功能,包括激活函数卸载以及使用 HybridDeviceOptimizer 进行的优化器状态卸载。

...展开全部
2
更新时间 2026-09-28

CPU 卸载

参考资料

  • 稳定文档:@docs/training/cpu-offloading.md
  • 结构化元数据:@skills/nemo-mbridge-perf-cpu-offloading/card.yaml

什么是 CPU 卸载

两种将数据从 GPU 转移到 CPU 内存的独立机制:

机制 配置命名空间 卸载的内容 PP限制
激活卸载 model.cpu_offloading* 每个变换器层的激活值(以及可选的权重) PP 必须为 1
优化器卸载 optimizer.optimizer_cpu_offload 通过HybridDeviceOptimizer设置 Adam 优化器的状态(动量 + 方差) 无

快速决策

情况 建议
大型MoE模型(30B+),需要PP > 1 优化器卸载——由于 PP=1,激活层卸载被阻止
中小型模型,PP=1适用,激活内存是主要瓶颈 激活函数卸载
希望实现可调的内存-速度权衡 通过优化器卸载因子(optimizer_offload_fraction)实现优化器卸载
吞吐量是首要优先级 不要启用——卸载总是会增加开销
需要 CUDA 图 仅支持优化器卸载——激活卸载不兼容
内存压力适中 为获得最佳效率,优化器卸载比例应控制在25%–50%

启用

优化器 CPU 卸载(推荐用于大型模型)

cfg.optimizer.optimizer_cpu_offload = True
cfg.optimizer.optimizer_offload_fraction = 1.0
cfg.optimizer.overlap_cpu_optimizer_d2h_h2d = True

CLI 覆盖设置:

optimizer.optimizer_cpu_offload=True \
optimizer.optimizer_offload_fraction=0.5 \
optimizer.overlap_cpu_optimizer_d2h_h2d=True

激活函数 CPU 卸载(仅限小型/中型模型)

cfg.model.cpu_offloading = True
cfg.model.cpu_offloading_num_layers = 16
cfg.model.cpu_offloading_activations = True
cfg.model.cpu_offloading_weights = False

cfg.model.pipeline_model_parallel_size = 1
cfg.model.recompute_granularity = None
cfg.model.cuda_graph_impl = "none"

配置参数参考

优化器卸载

参数 默认值 描述
optimizer_cpu_offload False 主开关
optimizer_offload_fraction 0.0 CPU 上优化器状态所占的比例 (0.0–1.0)
overlap_cpu_optimizer_d2h_h2d false 将 GPU↔CPU 数据传输与计算操作重叠
use_torch_optimizer_for_cpu_offload False 对 CPU 部分使用torch.optim代替融合优化器

激活函数卸载

参数 默认值 描述
cpu_offloading False 主开关
cpu_offloading_num_layers 0 要卸载的变压器层数(0 到 num_layers-1)
cpu_offloading_activations True 卸载激活层
cpu_offloading_weights False 卸载权重
cpu_offloading_double_buffering False 在重新加载时跨层进行双缓冲

兼容性与限制

激活卸载

  • pipeline_model_parallel_size必须为 1
  • recompute_granularity必须为None
  • 不能与fine_grained_activation_offloading结合使用
  • 不能与 CUDA 图结合使用
  • cpu_offloading_num_layers必须在[0, num_layers-1)范围内

优化器卸载

  • 需要use_distributed_optimizer = True(在大多数示例中为默认值)
  • 无 PP、重新计算或 CUDA 图的限制
  • optimizer_offload_fraction必须在[0.0, 1.0]范围内

实际应用:大型 MoE 模型

对于 Qwen3-30B-A3B 及类似的大型 MoE 模型,激活函数卸载功能被禁用。PP=1 的限制意味着每块 GPU 需承载全部 48 个层;仅模型 权重 + 优化器状态(约 70 GB)就已超过 H100 的 80 GB 容量。

最简可运行命令

uv run python scripts/training/run_recipe.py \
  --recipe qwen3_30b_a3b_pretrain_config \
  optimizer.optimizer_cpu_offload=True \
  optimizer.optimizer_offload_fraction=0.5 \
  train.train_iters=20 \
  train.global_batch_size=8 \
  train.micro_batch_size=1

验证

单元测试

uv run python -m pytest \
  tests/unit_tests/models/test_gpt_full_te_layer_autocast_spec.py -k "cpu_offload" \
  tests/unit_tests/peft/test_utils.py -k "cpu_offload" -q

成功标准

  • 所选卸载模式的配置验证通过
  • 训练完成且未出现 OOM 或 NCCL 错误
  • 损失值与未卸载的基线一致(最大偏差 < 0.001)
  • 内存使用量随卸载比例呈正比下降

代码锚点

MCore 激活功能卸载约束

       if self.cpu_offloading 且 (
            self.cpu_offloading_num_layers< 0 or self.cpu_offloading_num_layers >=self.num_layers
        ):
            raise ValueError(...)

        if self.cpu_offloading 且 self.pipeline_model_parallel_size > 1:
            raise ValueError(
                "目前不支持在启用 CPU 卸载时进行管道并行处理"
            )

        if self.cpu_offloading 且 self.recompute_granularity 不为 None:
            raise ValueError(
                "启用激活函数重新计算时,CPU 卸载无法正常工作"
            )

MCore CUDA 图不兼容

           if self.cpu_offloading:
                raise ValueError("CPU 卸载不支持 CUDA 图。")

MCore 细粒度卸载互斥

       if self.fine_grained_activation_offloading:
            assert (
                not self.cpu_offloading
            ), "启用 cpu_offloading 时无法启用 fine_grained_activation_offloading。"

MCore HybridDeviceOptimizer 实例化

       if config.optimizer_cpu_offload:
            # ... 配置 CPU/GPU 优化器类 ...
            optimizer = HybridDeviceOptimizer(
                param_groups,
                offload_fraction=config.optimizer_offload_fraction,
                cpu_optimizer_cls=cpu_optimizer_cls,
                gpu_optimizer_cls=gpu_optimizer_cls,
                overlap_cpu_optimizer_d2h_h2d=config.overlap_cpu_optimizer_d2h_h2d,
                pin_cpu_grads=config.pin_cpu_grads,
                pin_cpu_params=config.pin_cpu_params,
            )

Bridge CUDA 图保护机制

       assert not config.cpu_offloading 且 config.recompute_granularity 为 None, "不支持 CUDA 图"

在 PEFT 中桥接激活卸载

       if self.config.cpu_offloading 且 self.config.cpu_offloading_activations:
            x.activation_offloading = True
        x, _ = self.linear_in(x)
        x = self.activation(x)
        if self.config.cpu_offloading and self.config.cpu_offloading_activations:
            x.activation_offloading = True
        x, _ = self.linear_out(x)

故障诊断

症状 可能原因 如何确认 解决方法
目前不支持带 CPU 卸载功能的管道并行处理 激活卸载 + PP > 1 检查pipeline_model_parallel_size 将 PP 设为 1 或使用优化器卸载
当启用激活值重新计算时,CPU卸载无法正常工作 激活值卸载 + 重新计算 检查recompute_granularity 将recompute_granularity设为null
在启用 cpu_offloading 时无法启用 fine_grained_activation_offloading 两种卸载模式均已启用 检查这两个标志 请选择其中一种
CPU 卸载不支持 CUDA 图 CUDA 图 + 激活卸载 检查cuda_graph_impl 将cuda_graph_impl设置为"none"
启用激活卸载时发生内存不足(OOM) 模型过大,不适合 PP=1 检查已分配内存是否超过 80 GB 当 PP > 1 时使用优化器卸载
运行速度极度变慢(超过4倍) 100% 优化器卸载,CPU 成为 Adam 的瓶颈 比较不同比例下的迭代时间 降低分数或启用overlap_cpu_optimizer_d2h_h2d
部分优化器卸载时发生内存不足(OOM) 此配置下卸载不足 检查不同分数下的内存情况 增加分率或添加 PP

已知限制

  • 激活卸载要求 PP=1,这使得该方案对于需要流水线并行的大型模型 (30B+ MoE)而言并不实用。
  • 优化器卸载的吞吐量开销呈线性增长(以 Qwen3-30B-A3B 为例,25% 时约为 1.9 倍, 100% 时约为 4.2 倍)。
  • D2H/H2D 重叠仅提供约 7% 的加速,因为 CPU Adam 计算是 主要的瓶颈。
  • fine_grained_activation_offloading是一种独立的模块级方法, 在 PP > 1 时有效,但无法与层级的 cpu_offloading 结合使用。
在 GitHub 上查看
---
name: nemo-mbridge-perf-cpu-offloading
description: Configure and validate CPU offloading for Megatron Bridge training, including activation offloading and optimizer state offloading with HybridDeviceOptimizer.
license: Apache-2.0
---

# CPU Offloading

## References

- Stable docs: @docs/training/cpu-offloading.md
- Structured metadata: @skills/nemo-mbridge-perf-cpu-offloading/card.yaml

## What It Is

Two independent mechanisms to move data from GPU to CPU memory:

| Mechanism | Config namespace | What gets offloaded | PP restriction |
|---|---|---|---|
| Activation offloading | `model.cpu_offloading*` | Activations (and optionally weights) per transformer layer | PP must be 1 |
| Optimizer offloading | `optimizer.optimizer_cpu_offload` | Adam optimizer states (momentum + variance) via `HybridDeviceOptimizer` | None |

## Quick Decision

| Situation | Recommendation |
|---|---|
| Large MoE model (30B+), needs PP > 1 | Optimizer offloading — activation offloading is blocked by PP=1 |
| Small/medium model, PP=1 fits, activation memory dominates | Activation offloading |
| Want tunable memory-speed tradeoff | Optimizer offloading with fractional `optimizer_offload_fraction` |
| Throughput is top priority | Don't enable — offloading always adds overhead |
| CUDA graphs are needed | Only optimizer offloading — activation offloading is incompatible |
| Memory pressure is moderate | Optimizer offload at 25–50% fraction for best efficiency |

## Enablement

### Optimizer CPU offloading (recommended for large models)

```python
cfg.optimizer.optimizer_cpu_offload = True
cfg.optimizer.optimizer_offload_fraction = 1.0
cfg.optimizer.overlap_cpu_optimizer_d2h_h2d = True
```

CLI overrides:

```bash
optimizer.optimizer_cpu_offload=True \
optimizer.optimizer_offload_fraction=0.5 \
optimizer.overlap_cpu_optimizer_d2h_h2d=True
```

### Activation CPU offloading (small/medium models only)

```python
cfg.model.cpu_offloading = True
cfg.model.cpu_offloading_num_layers = 16
cfg.model.cpu_offloading_activations = True
cfg.model.cpu_offloading_weights = False

cfg.model.pipeline_model_parallel_size = 1
cfg.model.recompute_granularity = None
cfg.model.cuda_graph_impl = "none"
```

## Config Parameter Reference

### Optimizer offloading

| Parameter | Default | Description |
|-----------|---------|-------------|
| `optimizer_cpu_offload` | `False` | Master switch |
| `optimizer_offload_fraction` | `0.0` | Fraction of optimizer states on CPU (0.0–1.0) |
| `overlap_cpu_optimizer_d2h_h2d` | `False` | Overlap GPU↔CPU transfers with compute |
| `use_torch_optimizer_for_cpu_offload` | `False` | Use `torch.optim` instead of fused optimizer for CPU portion |

### Activation offloading

| Parameter | Default | Description |
|-----------|---------|-------------|
| `cpu_offloading` | `False` | Master switch |
| `cpu_offloading_num_layers` | `0` | Number of transformer layers to offload (0 to num_layers-1) |
| `cpu_offloading_activations` | `True` | Offload activations |
| `cpu_offloading_weights` | `False` | Offload weights |
| `cpu_offloading_double_buffering` | `False` | Double-buffer across layers while reloading |

## Compatibility And Constraints

### Activation offloading

- `pipeline_model_parallel_size` must be 1
- `recompute_granularity` must be `None`
- Cannot combine with `fine_grained_activation_offloading`
- Cannot combine with CUDA graphs
- `cpu_offloading_num_layers` must be in `[0, num_layers-1)`

### Optimizer offloading

- Requires `use_distributed_optimizer = True` (default in most recipes)
- No PP, recompute, or CUDA graph restrictions
- `optimizer_offload_fraction` must be in `[0.0, 1.0]`

### Practical: large MoE models

Activation offloading is blocked for Qwen3-30B-A3B and similar large MoE
models. The PP=1 constraint means each GPU holds all 48 layers; model
weights + optimizer states alone (~70 GB) exceed H100 80 GB capacity.

## Minimal Runnable Command

```bash
uv run python scripts/training/run_recipe.py \
  --recipe qwen3_30b_a3b_pretrain_config \
  optimizer.optimizer_cpu_offload=True \
  optimizer.optimizer_offload_fraction=0.5 \
  train.train_iters=20 \
  train.global_batch_size=8 \
  train.micro_batch_size=1
```

## Verification

### Unit tests

```bash
uv run python -m pytest \
  tests/unit_tests/models/test_gpt_full_te_layer_autocast_spec.py -k "cpu_offload" \
  tests/unit_tests/peft/test_utils.py -k "cpu_offload" -q
```

### Success criteria

- Config validation passes for the selected offloading mode
- Training completes without OOM or NCCL errors
- Loss matches the non-offloaded baseline (max delta < 0.001)
- Memory usage drops proportionally to offload fraction

## Code Anchors

### MCore activation offload constraints

```1296:1310:3rdparty/Megatron-LM/megatron/core/transformer/transformer_config.py
        if self.cpu_offloading and (
            self.cpu_offloading_num_layers < 0 or self.cpu_offloading_num_layers >= self.num_layers
        ):
            raise ValueError(...)

        if self.cpu_offloading and self.pipeline_model_parallel_size > 1:
            raise ValueError(
                "Currently there is no support for Pipeline parallelism with CPU offloading"
            )

        if self.cpu_offloading and self.recompute_granularity is not None:
            raise ValueError(
                "CPU offloading does not work when activation recomputation is enabled"
            )
```

### MCore CUDA graph incompatibility

```1943:1944:3rdparty/Megatron-LM/megatron/core/transformer/transformer_config.py
            if self.cpu_offloading:
                raise ValueError("CUDA graphs not supported with CPU offloading.")
```

### MCore fine-grained offloading mutual exclusion

```1427:1430:3rdparty/Megatron-LM/megatron/core/transformer/transformer_config.py
        if self.fine_grained_activation_offloading:
            assert (
                not self.cpu_offloading
            ), "fine_grained_activation_offloading cannot be enabled with cpu_offloading."
```

### MCore HybridDeviceOptimizer instantiation

```480:518:3rdparty/Megatron-LM/megatron/core/optimizer/__init__.py
        if config.optimizer_cpu_offload:
            # ... setup cpu/gpu optimizer classes ...
            optimizer = HybridDeviceOptimizer(
                param_groups,
                offload_fraction=config.optimizer_offload_fraction,
                cpu_optimizer_cls=cpu_optimizer_cls,
                gpu_optimizer_cls=gpu_optimizer_cls,
                overlap_cpu_optimizer_d2h_h2d=config.overlap_cpu_optimizer_d2h_h2d,
                pin_cpu_grads=config.pin_cpu_grads,
                pin_cpu_params=config.pin_cpu_params,
            )
```

### Bridge CUDA graph guard

```232:234:src/megatron/bridge/models/gpt_full_te_layer_autocast_spec.py
        assert not config.cpu_offloading and config.recompute_granularity is None, "Cudagraphs not supported"
```

### Bridge activation offloading in PEFT

```621:631:src/megatron/bridge/peft/utils.py
        if self.config.cpu_offloading and self.config.cpu_offloading_activations:
            x.activation_offloading = True
        x, _ = self.linear_in(x)
        x = self.activation(x)
        if self.config.cpu_offloading and self.config.cpu_offloading_activations:
            x.activation_offloading = True
        x, _ = self.linear_out(x)
```

## Failure Diagnosis

| Symptom | Likely Cause | How To Confirm | Fix |
|---|---|---|---|
| `Currently there is no support for Pipeline parallelism with CPU offloading` | Activation offload + PP > 1 | Check `pipeline_model_parallel_size` | Set PP=1 or use optimizer offloading |
| `CPU offloading does not work when activation recomputation is enabled` | Activation offload + recompute | Check `recompute_granularity` | Set `recompute_granularity=null` |
| `fine_grained_activation_offloading cannot be enabled with cpu_offloading` | Both offloading modes enabled | Check both flags | Use one or the other |
| `CUDA graphs not supported with CPU offloading` | CUDA graphs + activation offload | Check `cuda_graph_impl` | Set `cuda_graph_impl="none"` |
| OOM with activation offloading | Model too large for PP=1 | Check allocated memory vs 80 GB | Use optimizer offloading with PP > 1 |
| Extreme slowdown (>4x) | 100% optimizer offload, CPU Adam bottleneck | Compare iter time at different fractions | Reduce fraction or enable `overlap_cpu_optimizer_d2h_h2d` |
| OOM at partial optimizer offload | Insufficient offload for this config | Check memory at different fractions | Increase fraction or add PP |

## Known Limitations

- Activation offloading requires PP=1, making it impractical for large models
  (30B+ MoE) that need pipeline parallelism.
- Optimizer offloading throughput penalty scales linearly (~1.9x at 25%,
  ~4.2x at 100% for Qwen3-30B-A3B).
- D2H/H2D overlap provides only ~7% speedup because CPU Adam compute is
  the dominant bottleneck.
- `fine_grained_activation_offloading` is a separate module-level approach
  that works with PP > 1 but cannot be combined with layer-level
  `cpu_offloading`.

安装 nemo-mbridge-perf-cpu-offloading

下载技能文件并将其解压到 .claude/skills/ 目录中。

下载ZIP

克隆仓库并复制技能文件到您的项目中。

git clone https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-cpu-offloading # Copy SKILL.md to your .claude/skills/ directory

复制 复制
快速设置: 将技能文件夹复制到 .claude/skills/ Claude 将自动检测并使用该技能
仓库 NVIDIA/skills

相关技能

web-search
更新时间 2026-06-29
webapp-testing
更新时间 2026-06-29
lark-base
更新时间 2026-07-05
agentmail
更新时间 2026-06-29
OR