nemo-mbridge-perf-cpu-offloading
NVIDIA/skills
Megatron BridgeのトレーニングにおけるCPUオフロード(HybridDeviceOptimizerによる活性化オフロードおよびオプティマイザ状態のオフロードを含む)を設定し、検証する。
...すべて拡張しますCPUオフロード
参考文献
- 安定版ドキュメント: @docs/training/cpu-offloading.md
- 構造化メタデータ: @skills/nemo-mbridge-perf-cpu-offloading/card.yaml
概要
GPUからCPUメモリへデータを移動させる2つの独立したメカニズム:
| メカニズム | 設定ネームスペース | オフロードされる対象 | PPの制限 |
|---|---|---|---|
| アクティベーションのオフロード | model.cpu_offloading* |
トランスフォーマー層ごとのアクティベーション(およびオプションで重み) | PPは1でなければならない |
| オプティマイザーのオフロード | optimizer.optimizer_cpu_offload |
HybridDeviceOptimizerによる Adam オプティマイザーの状態(モメンタム + 分散) |
なし |
クイック決定
| 状況 | 推奨事項 |
|---|---|
| 大規模なMoEモデル(30B以上)、PP > 1が必要 | オプティマイザーのオフロード — PP=1のため、アクティベーションのオフロードはブロックされる |
| 中小規模のモデルで、PP=1が適切であり、活性化メモリが主な要因となる | 活性化オフロード |
| メモリと速度のトレードオフを調整可能にしたい | optimizer_offload_fractionによるオプティマイザーのオフロード |
| スループットを最優先とする | 有効にしない — オフロードは常にオーバーヘッドを生じる |
| CUDAグラフが必要 | オプティマイザのオフロードのみ — アクティベーションのオフロードとは互換性がない |
| メモリ負荷は中程度 | 最高の効率を得るには、オプティマイザー・オフロードの割合を25~50%に設定する |
有効化
オプティマイザの CPU オフロード(大規模モデルに推奨)
cfg.optimizer.optimizer_cpu_offload = True
cfg.optimizer.optimizer_offload_fraction = 1.0
cfg.optimizer.overlap_cpu_optimizer_d2h_h2d = True
CLIによる上書き設定:
optimizer.optimizer_cpu_offload=True \
optimizer.optimizer_offload_fraction=0.5 \
optimizer.overlap_cpu_optimizer_d2h_h2d=True
活性化関数の CPU オフロード(小・中規模モデルのみ)
cfg.model.cpu_offloading = True
cfg.model.cpu_offloading_num_layers = 16
cfg.model.cpu_offloading_activations = True
cfg.model.cpu_offloading_weights = False
cfg.model.pipeline_model_parallel_size = 1
cfg.model.recompute_granularity = None
cfg.model.cuda_graph_impl = "none"
設定パラメータリファレンス
オプティマイザーのオフロード
| パラメータ | デフォルト | 説明 |
|---|---|---|
optimizer_cpu_offload |
False |
マスタースイッチ |
optimizer_offload_fraction |
0.0 |
CPU上のオプティマイザ状態の割合 (0.0–1.0) |
overlap_cpu_optimizer_d2h_h2d |
False |
GPU↔CPU間の転送を演算とオーバーラップさせる |
use_torch_optimizer_for_cpu_offload |
False |
CPU 部分では、fused オプティマイザーの代わりにtorch.optimを使用する |
活性化関数のオフロード
| パラメータ | デフォルト | 説明 |
|---|---|---|
cpu_offloading |
False |
マスタースイッチ |
cpu_offloading_num_layers |
0 |
オフロードするトランスフォーマー層の数 (0 ~ num_layers-1) |
cpu_offloading_activations |
True |
活性化関数のオフロード |
cpu_offloading_weights |
False |
重みのオフロード |
cpu_offloading_double_buffering |
False |
再読み込み時のレイヤー間ダブルバッファリング |
互換性と制約
活性化のオフロード
pipeline_model_parallel_size は1 である必要があるrecompute_granularity はNoneでなければならないfine_grained_activation_offloadingとは組み合わせられない- CUDAグラフとの併用はできません
- `
cpu_offloading_num_layers` は[0, num_layers-1)の範囲でなければなりません
オプティマイザーのオフロード
use_distributed_optimizer = Trueである必要があります(ほとんどのリシピでデフォルト)- PP、再計算、または CUDA グラフに関する制限はありません
- `
optimizer_offload_fraction` は[0.0, 1.0]の範囲でなければなりません
実用例:大規模なMoEモデル
Qwen3-30B-A3B および同様の大規模 MoE モデルでは、活性化関数のオフロードはブロックされます。PP=1 の制約により、各 GPU が 48 層すべてを保持することになり、モデル の重みとオプティマイザの状態のみでも (~70 GB) が H100 の 80 GB の容量を超過します。
最小限の実行可能コマンド
uv run python scripts/training/run_recipe.py \
--recipe qwen3_30b_a3b_pretrain_config \
optimizer.optimizer_cpu_offload=True \
optimizer.optimizer_offload_fraction=0.5 \
train.train_iters=20 \
train.global_batch_size=8 \
train.micro_batch_size=1
検証
単体テスト
uv run python -m pytest \
tests/unit_tests/models/test_gpt_full_te_layer_autocast_spec.py -k "cpu_offload" \
tests/unit_tests/peft/test_utils.py -k "cpu_offload" -q
成功基準
- 選択したオフロードモードに対して、構成の検証に合格すること
- OOM エラーや NCCL エラーが発生することなくトレーニングが完了する
- 損失がオフロードを行わないベースラインと一致する(最大偏差 < 0.001)
- メモリ使用量は、オフロードの割合に比例して減少する
コードアンカー
MCoreの活性化オフロードに関する制約
if self.cpu_offloading かつ (
self.cpu_offloading_num_layers< 0 or self.cpu_offloading_num_layers >= self.num_layers
):
raise ValueError(...)
if self.cpu_offloading かつ self.pipeline_model_parallel_size > 1:
raise ValueError(
"現在、CPUオフロードとパイプライン並列処理の併用はサポートされていません"
)
if self.cpu_offloading かつ self.recompute_granularity が None でない:
raise ValueError(
"活性化関数の再計算が有効な場合、CPUオフロードは機能しません"
)
MCore CUDA グラフの非互換性
if self.cpu_offloading:
raise ValueError("CPUオフロードではCUDAグラフはサポートされていません。")
MCoreの細粒度オフロードにおける排他制御
if self.fine_grained_activation_offloading:
assert (
not self.cpu_offloading
), "cpu_offloading が有効な場合、fine_grained_activation_offloading を有効にすることはできません。"
MCore HybridDeviceOptimizerのインスタンス化
if config.optimizer_cpu_offload:
# ... CPU/GPU オプティマイザクラスのセットアップ ...
optimizer = HybridDeviceOptimizer(
param_groups,
offload_fraction=config.optimizer_offload_fraction,
cpu_optimizer_cls=cpu_optimizer_cls,
gpu_optimizer_cls=gpu_optimizer_cls,
overlap_cpu_optimizer_d2h_h2d=config.overlap_cpu_optimizer_d2h_h2d,
pin_cpu_grads=config.pin_cpu_grads,
pin_cpu_params=config.pin_cpu_params,
)
Bridge CUDA グラフガード
assert not config.cpu_offloading かつ config.recompute_granularity が None, "Cudagraphs はサポートされていません"
PEFT におけるアクティベーションのオフロードのブリッジ
if self.config.cpu_offloading かつ self.config.cpu_offloading_activations:
x.activation_offloading = True
x, _ = self.linear_in(x)
x = self.activation(x)
if self.config.cpu_offloading and self.config.cpu_offloading_activations:
x.activation_offloading = True
x, _ = self.linear_out(x)
障害診断
| 症状 | 考えられる原因 | 確認方法 | 修正方法 |
|---|---|---|---|
現在、CPUオフロードによるパイプライン並列処理はサポートされていません |
アクティベーションのオフロード + PP > 1 | pipeline_model_parallel_sizeを確認してください |
PP=1に設定するか、オプティマイザによるオフロードを使用してください |
アクティベーションの再計算が有効になっている場合、CPUオフロードは機能しません |
活性化オフロード + 再計算 | recompute_granularityを確認してください |
recompute_granularity を nullに設定してください |
cpu_offloading を有効にしている場合、fine_grained_activation_offloading を有効にすることはできません |
両方のオフロードモードが有効 | 両方のフラグを確認してください | どちらか一方を使用してください |
CPUオフロードではCUDAグラフはサポートされていません |
CUDAグラフ + アクティベーションオフロード | cuda_graph_implを確認してください |
cuda_graph_impl="none"に設定してください |
| アクティベーションのオフロード時にOOMが発生 | PP=1 に対してモデルが大きすぎる | 割り当てられたメモリと80 GBを比較してください | PP > 1 の場合、オプティマイザー・オフロードを使用してください |
| 極端な処理速度の低下(4倍以上) | 100%のオプティマイザ・オフロード、CPUがAdamのボトルネック | 異なるフラクションでの反復時間を比較 | fractionを縮小するか、overlap_cpu_optimizer_d2h_h2dを有効にする |
| オプティマイザーのオフロードが部分的な状態でのOOM | この構成ではオフロードが不十分 | 各フラクションにおけるメモリ容量を確認 | フラクションを増やすか、PPを追加する |
既知の制限事項
- アクティベーションのオフロードには PP=1 が必要であるため、パイプライン並列処理を必要とする大規模モデル (30B+ MoE)では実用的ではありません。
- オプティマイザーのオフロードによるスループットの低下は線形に増加します(Qwen3-30B-A3Bの場合、25%で約1.9倍、 100%で約4.2倍)。
- D2H/H2Dのオーバーラップによる速度向上は約7%にとどまります。これは、CPUでのAdam計算が 主なボトルネックとなっているためです。
fine_grained_activation_offloadingは、PP > 1で機能する別のモジュールレベルのアプローチですが、 レイヤーレベルのcpu_offloadingとは組み合わせることができません。
---
name: nemo-mbridge-perf-cpu-offloading
description: Configure and validate CPU offloading for Megatron Bridge training, including activation offloading and optimizer state offloading with HybridDeviceOptimizer.
license: Apache-2.0
---
# CPU Offloading
## References
- Stable docs: @docs/training/cpu-offloading.md
- Structured metadata: @skills/nemo-mbridge-perf-cpu-offloading/card.yaml
## What It Is
Two independent mechanisms to move data from GPU to CPU memory:
| Mechanism | Config namespace | What gets offloaded | PP restriction |
|---|---|---|---|
| Activation offloading | `model.cpu_offloading*` | Activations (and optionally weights) per transformer layer | PP must be 1 |
| Optimizer offloading | `optimizer.optimizer_cpu_offload` | Adam optimizer states (momentum + variance) via `HybridDeviceOptimizer` | None |
## Quick Decision
| Situation | Recommendation |
|---|---|
| Large MoE model (30B+), needs PP > 1 | Optimizer offloading — activation offloading is blocked by PP=1 |
| Small/medium model, PP=1 fits, activation memory dominates | Activation offloading |
| Want tunable memory-speed tradeoff | Optimizer offloading with fractional `optimizer_offload_fraction` |
| Throughput is top priority | Don't enable — offloading always adds overhead |
| CUDA graphs are needed | Only optimizer offloading — activation offloading is incompatible |
| Memory pressure is moderate | Optimizer offload at 25–50% fraction for best efficiency |
## Enablement
### Optimizer CPU offloading (recommended for large models)
```python
cfg.optimizer.optimizer_cpu_offload = True
cfg.optimizer.optimizer_offload_fraction = 1.0
cfg.optimizer.overlap_cpu_optimizer_d2h_h2d = True
```
CLI overrides:
```bash
optimizer.optimizer_cpu_offload=True \
optimizer.optimizer_offload_fraction=0.5 \
optimizer.overlap_cpu_optimizer_d2h_h2d=True
```
### Activation CPU offloading (small/medium models only)
```python
cfg.model.cpu_offloading = True
cfg.model.cpu_offloading_num_layers = 16
cfg.model.cpu_offloading_activations = True
cfg.model.cpu_offloading_weights = False
cfg.model.pipeline_model_parallel_size = 1
cfg.model.recompute_granularity = None
cfg.model.cuda_graph_impl = "none"
```
## Config Parameter Reference
### Optimizer offloading
| Parameter | Default | Description |
|-----------|---------|-------------|
| `optimizer_cpu_offload` | `False` | Master switch |
| `optimizer_offload_fraction` | `0.0` | Fraction of optimizer states on CPU (0.0–1.0) |
| `overlap_cpu_optimizer_d2h_h2d` | `False` | Overlap GPU↔CPU transfers with compute |
| `use_torch_optimizer_for_cpu_offload` | `False` | Use `torch.optim` instead of fused optimizer for CPU portion |
### Activation offloading
| Parameter | Default | Description |
|-----------|---------|-------------|
| `cpu_offloading` | `False` | Master switch |
| `cpu_offloading_num_layers` | `0` | Number of transformer layers to offload (0 to num_layers-1) |
| `cpu_offloading_activations` | `True` | Offload activations |
| `cpu_offloading_weights` | `False` | Offload weights |
| `cpu_offloading_double_buffering` | `False` | Double-buffer across layers while reloading |
## Compatibility And Constraints
### Activation offloading
- `pipeline_model_parallel_size` must be 1
- `recompute_granularity` must be `None`
- Cannot combine with `fine_grained_activation_offloading`
- Cannot combine with CUDA graphs
- `cpu_offloading_num_layers` must be in `[0, num_layers-1)`
### Optimizer offloading
- Requires `use_distributed_optimizer = True` (default in most recipes)
- No PP, recompute, or CUDA graph restrictions
- `optimizer_offload_fraction` must be in `[0.0, 1.0]`
### Practical: large MoE models
Activation offloading is blocked for Qwen3-30B-A3B and similar large MoE
models. The PP=1 constraint means each GPU holds all 48 layers; model
weights + optimizer states alone (~70 GB) exceed H100 80 GB capacity.
## Minimal Runnable Command
```bash
uv run python scripts/training/run_recipe.py \
--recipe qwen3_30b_a3b_pretrain_config \
optimizer.optimizer_cpu_offload=True \
optimizer.optimizer_offload_fraction=0.5 \
train.train_iters=20 \
train.global_batch_size=8 \
train.micro_batch_size=1
```
## Verification
### Unit tests
```bash
uv run python -m pytest \
tests/unit_tests/models/test_gpt_full_te_layer_autocast_spec.py -k "cpu_offload" \
tests/unit_tests/peft/test_utils.py -k "cpu_offload" -q
```
### Success criteria
- Config validation passes for the selected offloading mode
- Training completes without OOM or NCCL errors
- Loss matches the non-offloaded baseline (max delta < 0.001)
- Memory usage drops proportionally to offload fraction
## Code Anchors
### MCore activation offload constraints
```1296:1310:3rdparty/Megatron-LM/megatron/core/transformer/transformer_config.py
if self.cpu_offloading and (
self.cpu_offloading_num_layers < 0 or self.cpu_offloading_num_layers >= self.num_layers
):
raise ValueError(...)
if self.cpu_offloading and self.pipeline_model_parallel_size > 1:
raise ValueError(
"Currently there is no support for Pipeline parallelism with CPU offloading"
)
if self.cpu_offloading and self.recompute_granularity is not None:
raise ValueError(
"CPU offloading does not work when activation recomputation is enabled"
)
```
### MCore CUDA graph incompatibility
```1943:1944:3rdparty/Megatron-LM/megatron/core/transformer/transformer_config.py
if self.cpu_offloading:
raise ValueError("CUDA graphs not supported with CPU offloading.")
```
### MCore fine-grained offloading mutual exclusion
```1427:1430:3rdparty/Megatron-LM/megatron/core/transformer/transformer_config.py
if self.fine_grained_activation_offloading:
assert (
not self.cpu_offloading
), "fine_grained_activation_offloading cannot be enabled with cpu_offloading."
```
### MCore HybridDeviceOptimizer instantiation
```480:518:3rdparty/Megatron-LM/megatron/core/optimizer/__init__.py
if config.optimizer_cpu_offload:
# ... setup cpu/gpu optimizer classes ...
optimizer = HybridDeviceOptimizer(
param_groups,
offload_fraction=config.optimizer_offload_fraction,
cpu_optimizer_cls=cpu_optimizer_cls,
gpu_optimizer_cls=gpu_optimizer_cls,
overlap_cpu_optimizer_d2h_h2d=config.overlap_cpu_optimizer_d2h_h2d,
pin_cpu_grads=config.pin_cpu_grads,
pin_cpu_params=config.pin_cpu_params,
)
```
### Bridge CUDA graph guard
```232:234:src/megatron/bridge/models/gpt_full_te_layer_autocast_spec.py
assert not config.cpu_offloading and config.recompute_granularity is None, "Cudagraphs not supported"
```
### Bridge activation offloading in PEFT
```621:631:src/megatron/bridge/peft/utils.py
if self.config.cpu_offloading and self.config.cpu_offloading_activations:
x.activation_offloading = True
x, _ = self.linear_in(x)
x = self.activation(x)
if self.config.cpu_offloading and self.config.cpu_offloading_activations:
x.activation_offloading = True
x, _ = self.linear_out(x)
```
## Failure Diagnosis
| Symptom | Likely Cause | How To Confirm | Fix |
|---|---|---|---|
| `Currently there is no support for Pipeline parallelism with CPU offloading` | Activation offload + PP > 1 | Check `pipeline_model_parallel_size` | Set PP=1 or use optimizer offloading |
| `CPU offloading does not work when activation recomputation is enabled` | Activation offload + recompute | Check `recompute_granularity` | Set `recompute_granularity=null` |
| `fine_grained_activation_offloading cannot be enabled with cpu_offloading` | Both offloading modes enabled | Check both flags | Use one or the other |
| `CUDA graphs not supported with CPU offloading` | CUDA graphs + activation offload | Check `cuda_graph_impl` | Set `cuda_graph_impl="none"` |
| OOM with activation offloading | Model too large for PP=1 | Check allocated memory vs 80 GB | Use optimizer offloading with PP > 1 |
| Extreme slowdown (>4x) | 100% optimizer offload, CPU Adam bottleneck | Compare iter time at different fractions | Reduce fraction or enable `overlap_cpu_optimizer_d2h_h2d` |
| OOM at partial optimizer offload | Insufficient offload for this config | Check memory at different fractions | Increase fraction or add PP |
## Known Limitations
- Activation offloading requires PP=1, making it impractical for large models
(30B+ MoE) that need pipeline parallelism.
- Optimizer offloading throughput penalty scales linearly (~1.9x at 25%,
~4.2x at 100% for Qwen3-30B-A3B).
- D2H/H2D overlap provides only ~7% speedup because CPU Adam compute is
the dominant bottleneck.
- `fine_grained_activation_offloading` is a separate module-level approach
that works with PP > 1 but cannot be combined with layer-level
`cpu_offloading`.
すべてのファイル
6件のファイルnemo-mbridge-perf-cpu-offloadingをインストール
スキルファイルをダウンロードし、.claude/skills/ ディレクトリに解凍してください。
ZIPをダウンロードリポジトリをクローンし、スキルファイルをプロジェクトにコピーしてください。
git clone https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-cpu-offloading # Copy SKILL.md to your .claude/skills/ directory
コピー





家
