nemo-mbridge-perf-cpu-offloading
NVIDIA/skills
針對 Megatron Bridge 訓練,設定並驗證 CPU 卸載功能,包括啟用函數卸載,以及透過 HybridDeviceOptimizer 進行的優化器狀態卸載。
...展開全部CPU 卸載
參考資料
- 穩定版文件:@docs/training/cpu-offloading.md
- 結構化元資料:@skills/nemo-mbridge-perf-cpu-offloading/card.yaml
是什麼
兩種將資料從 GPU 移至 CPU 記憶體的獨立機制:
| 機制 | 配置命名空間 | 哪些資料會被卸載 | PP 限制 |
|---|---|---|---|
| 激活卸載 | model.cpu_offloading* |
每個變換器層的激發(以及可選的權重) | PP 必須為 1 |
| 優化器卸載 | optimizer.optimizer_cpu_offload |
透過HybridDeviceOptimizer傳遞 Adam 優化器的狀態(動量 + 方差) |
無 |
快速決策
| 情況 | 建議 |
|---|---|
| 大型 MoE 模型(30B+),需 PP > 1 | 優化器卸載 — 由於 PP=1,激活函數卸載受阻 |
| 中小型模型,PP=1 適用,且活化記憶體佔主導地位 | 激發函數卸載 |
| 希望能夠調整記憶體與速度之間的權衡 | 透過optimizer_offload_fraction參數進行優化器卸載 |
| 吞吐量為首要考量 | 請勿啟用 — 卸載總是會增加開銷 |
| 需要 CUDA 圖 | 僅限優化器卸載 — 啟用卸載與此不相容 |
| 記憶體壓力屬中等 | 為達到最佳效率,應將優化器卸載比例設定在 25–50% 之間 |
啟用設定
最佳化器 CPU 卸載(建議用於大型模型)
cfg.optimizer.optimizer_cpu_offload = True
cfg.optimizer.optimizer_offload_fraction = 1.0
cfg.optimizer.overlap_cpu_optimizer_d2h_h2d = True
命令列介面(CLI)覆寫設定:
optimizer.optimizer_cpu_offload=True \
optimizer.optimizer_offload_fraction=0.5 \
optimizer.overlap_cpu_optimizer_d2h_h2d=True
活化函數 CPU 卸載(僅限小型/中型模型)
cfg.model.cpu_offloading = True
cfg.model.cpu_offloading_num_layers = 16
cfg.model.cpu_offloading_activations = True
cfg.model.cpu_offloading_weights = False
cfg.model.pipeline_model_parallel_size = 1
cfg.model.recompute_granularity = None
cfg.model.cuda_graph_impl = "none"
配置參數參考
優化器卸載
| 參數 | 預設值 | 說明 |
|---|---|---|
optimizer_cpu_offload |
False |
主開關 |
optimizer_offload_fraction |
0.0 |
CPU 上優化器狀態的占比 (0.0–1.0) |
overlap_cpu_optimizer_d2h_h2d |
false |
讓 GPU↔CPU 資料傳輸與運算重疊 |
使用 Torch 優化器進行 CPU 卸載 |
False |
針對 CPU 部分,使用torch.optim取代融合式優化器 |
活化函數卸載
| 參數 | 預設值 | 說明 |
|---|---|---|
cpu_offloading |
False |
主開關 |
cpu_offloading_num_layers |
0 |
要卸載的變壓器層數(0 至 num_layers-1) |
cpu_offloading_activations |
True |
卸載激發函數 |
cpu_offloading_weights |
False |
卸載權重 |
cpu_offloading_double_buffering |
False |
重新載入時跨層進行雙緩衝 |
相容性與限制
激發函數卸載
pipeline_model_parallel_size必須為 1recompute_granularity必須為None- 無法與
fine_grained_activation_offloading結合使用 - 無法與 CUDA 圖結合
cpu_offloading_num_layers必須位於[0, num_layers-1)範圍內
優化器卸載
- 需設定
use_distributed_optimizer = True(多數範例中預設為此值) - 無 PP、重新計算或 CUDA 圖的限制
optimizer_offload_fraction必須位於[0.0, 1.0]之間
實務應用:大型 MoE 模型
對於 Qwen3-30B-A3B 及類似的大型 MoE 模型,激發函數卸載功能將被停用。由於 PP=1 的限制,意味著每張 GPU 必須承載全部 48 層;僅模型 權重與優化器狀態(約 70 GB)的總和,便已超過 H100 的 80 GB 容量。
最簡可執行命令
uv run python scripts/training/run_recipe.py \
--recipe qwen3_30b_a3b_pretrain_config \
optimizer.optimizer_cpu_offload=True \
optimizer.optimizer_offload_fraction=0.5 \
train.train_iters=20 \
train.global_batch_size=8 \
train.micro_batch_size=1
驗證
單元測試
uv run python -m pytest \
tests/unit_tests/models/test_gpt_full_te_layer_autocast_spec.py -k "cpu_offload" \
tests/unit_tests/peft/test_utils.py -k "cpu_offload" -q
成功準則
- 所選卸載模式的配置驗證通過
- 訓練完成且未發生 OOM 或 NCCL 錯誤
- 損失值與未卸載的基準相符(最大偏差 < 0.001)
- 記憶體使用量隨卸載比例成比例下降
程式碼錨點
MCore 激發函數卸載限制
if self.cpu_offloading 且 (
self.cpu_offloading_num_layers< 0 or self.cpu_offloading_num_layers >=self.num_layers
):
raise ValueError(...)
if self.cpu_offloading 且 self.pipeline_model_parallel_size > 1:
raise ValueError(
"目前不支援搭配 CPU 卸載的管線並行處理"
)
if self.cpu_offloading 且 self.recompute_granularity 非 None:
raise ValueError(
"當啟用活化函數重新計算時,CPU 卸載功能無法運作"
)
MCore 與 CUDA 圖的不相容性
if self.cpu_offloading:
raise ValueError("不支援在 CPU 卸載情況下使用 CUDA 圖。")
MCore 細粒度卸載互斥
if self.fine_grained_activation_offloading:
assert (
not self.cpu_offloading
), "在啟用 cpu_offloading 的情況下,無法啟用 fine_grained_activation_offloading。"
MCore HybridDeviceOptimizer 實例化
if config.optimizer_cpu_offload:
# ... 設定 CPU/GPU 最佳化器類別 ...
optimizer = HybridDeviceOptimizer(
param_groups,
offload_fraction=config.optimizer_offload_fraction,
cpu_optimizer_cls=cpu_optimizer_cls,
gpu_optimizer_cls=gpu_optimizer_cls,
overlap_cpu_optimizer_d2h_h2d=config.overlap_cpu_optimizer_d2h_h2d,
pin_cpu_grads=config.pin_cpu_grads,
pin_cpu_params=config.pin_cpu_params,
)
Bridge CUDA 圖護衛
assert not config.cpu_offloading 且 config.recompute_granularity 是 None, "不支援 CudaGraph"
在 PEFT 中橋接活化卸載
if self.config.cpu_offloading 且 self.config.cpu_offloading_activations:
x.activation_offloading = True
x, _ = self.linear_in(x)
x = self.activation(x)
if self.config.cpu_offloading and self.config.cpu_offloading_activations:
x.activation_offloading = True
x, _ = self.linear_out(x)
故障診斷
| 症狀 | 可能原因 | 如何確認 | 解決方法 |
|---|---|---|---|
目前不支援搭配 CPU 卸載的 Pipeline 並行處理 |
激發器卸載 + PP > 1 | 檢查pipeline_model_parallel_size |
將 PP 設為 1 或使用優化器卸載 |
當啟用激活值重新計算時,CPU 卸載功能無法運作 |
激發值卸載 + 重新計算 | 檢查recompute_granularity |
將recompute_granularity設定為null |
無法在啟用 cpu_offloading 的情況下啟用 fine_grained_activation_offloading |
兩種卸載模式皆已啟用 | 檢查兩個標誌 | 請選擇其中一種模式 |
CPU 卸載不支援 CUDA 圖 |
CUDA 圖 + 激活卸載 | 檢查cuda_graph_impl |
將cuda_graph_impl設定為"none" |
| 啟用啟用函式卸載時發生 OOM 錯誤 | 模型過大,不適用於 PP=1 | 檢查已分配記憶體是否超過 80 GB | 當 PP > 1 時,請使用優化器卸載 |
| 極度變慢(>4 倍) | 100% 優化器卸載,CPU 成為 Adam 的瓶頸 | 比較不同分數下的迭代時間 | 降低分數或啟用overlap_cpu_optimizer_d2h_h2d |
| 部分優化器卸載時發生 OOM | 此設定下的卸載不足 | 檢查不同比例下的記憶體狀況 | 增加分數或新增 PP |
已知限制
- 激發函數卸載需要 PP=1,因此對於需要管線並行處理的大型模型 (300 億+ MoE)而言並不切實可行。
- 優化器卸載的吞吐量開銷呈線性增長(以 Qwen3-30B-A3B 為例,25% 時約 1.9 倍, 100% 時約 4.2 倍)。
- D2H/H2D 重疊僅能提供約 7% 的加速效果,因為 CPU 上的 Adam 運算 是主要的瓶頸。
fine_grained_activation_offloading是一種獨立的模組級方法, 在 PP > 1 的情況下可行,但無法與層級的cpu_offloading結合使用。
---
name: nemo-mbridge-perf-cpu-offloading
description: Configure and validate CPU offloading for Megatron Bridge training, including activation offloading and optimizer state offloading with HybridDeviceOptimizer.
license: Apache-2.0
---
# CPU Offloading
## References
- Stable docs: @docs/training/cpu-offloading.md
- Structured metadata: @skills/nemo-mbridge-perf-cpu-offloading/card.yaml
## What It Is
Two independent mechanisms to move data from GPU to CPU memory:
| Mechanism | Config namespace | What gets offloaded | PP restriction |
|---|---|---|---|
| Activation offloading | `model.cpu_offloading*` | Activations (and optionally weights) per transformer layer | PP must be 1 |
| Optimizer offloading | `optimizer.optimizer_cpu_offload` | Adam optimizer states (momentum + variance) via `HybridDeviceOptimizer` | None |
## Quick Decision
| Situation | Recommendation |
|---|---|
| Large MoE model (30B+), needs PP > 1 | Optimizer offloading — activation offloading is blocked by PP=1 |
| Small/medium model, PP=1 fits, activation memory dominates | Activation offloading |
| Want tunable memory-speed tradeoff | Optimizer offloading with fractional `optimizer_offload_fraction` |
| Throughput is top priority | Don't enable — offloading always adds overhead |
| CUDA graphs are needed | Only optimizer offloading — activation offloading is incompatible |
| Memory pressure is moderate | Optimizer offload at 25–50% fraction for best efficiency |
## Enablement
### Optimizer CPU offloading (recommended for large models)
```python
cfg.optimizer.optimizer_cpu_offload = True
cfg.optimizer.optimizer_offload_fraction = 1.0
cfg.optimizer.overlap_cpu_optimizer_d2h_h2d = True
```
CLI overrides:
```bash
optimizer.optimizer_cpu_offload=True \
optimizer.optimizer_offload_fraction=0.5 \
optimizer.overlap_cpu_optimizer_d2h_h2d=True
```
### Activation CPU offloading (small/medium models only)
```python
cfg.model.cpu_offloading = True
cfg.model.cpu_offloading_num_layers = 16
cfg.model.cpu_offloading_activations = True
cfg.model.cpu_offloading_weights = False
cfg.model.pipeline_model_parallel_size = 1
cfg.model.recompute_granularity = None
cfg.model.cuda_graph_impl = "none"
```
## Config Parameter Reference
### Optimizer offloading
| Parameter | Default | Description |
|-----------|---------|-------------|
| `optimizer_cpu_offload` | `False` | Master switch |
| `optimizer_offload_fraction` | `0.0` | Fraction of optimizer states on CPU (0.0–1.0) |
| `overlap_cpu_optimizer_d2h_h2d` | `False` | Overlap GPU↔CPU transfers with compute |
| `use_torch_optimizer_for_cpu_offload` | `False` | Use `torch.optim` instead of fused optimizer for CPU portion |
### Activation offloading
| Parameter | Default | Description |
|-----------|---------|-------------|
| `cpu_offloading` | `False` | Master switch |
| `cpu_offloading_num_layers` | `0` | Number of transformer layers to offload (0 to num_layers-1) |
| `cpu_offloading_activations` | `True` | Offload activations |
| `cpu_offloading_weights` | `False` | Offload weights |
| `cpu_offloading_double_buffering` | `False` | Double-buffer across layers while reloading |
## Compatibility And Constraints
### Activation offloading
- `pipeline_model_parallel_size` must be 1
- `recompute_granularity` must be `None`
- Cannot combine with `fine_grained_activation_offloading`
- Cannot combine with CUDA graphs
- `cpu_offloading_num_layers` must be in `[0, num_layers-1)`
### Optimizer offloading
- Requires `use_distributed_optimizer = True` (default in most recipes)
- No PP, recompute, or CUDA graph restrictions
- `optimizer_offload_fraction` must be in `[0.0, 1.0]`
### Practical: large MoE models
Activation offloading is blocked for Qwen3-30B-A3B and similar large MoE
models. The PP=1 constraint means each GPU holds all 48 layers; model
weights + optimizer states alone (~70 GB) exceed H100 80 GB capacity.
## Minimal Runnable Command
```bash
uv run python scripts/training/run_recipe.py \
--recipe qwen3_30b_a3b_pretrain_config \
optimizer.optimizer_cpu_offload=True \
optimizer.optimizer_offload_fraction=0.5 \
train.train_iters=20 \
train.global_batch_size=8 \
train.micro_batch_size=1
```
## Verification
### Unit tests
```bash
uv run python -m pytest \
tests/unit_tests/models/test_gpt_full_te_layer_autocast_spec.py -k "cpu_offload" \
tests/unit_tests/peft/test_utils.py -k "cpu_offload" -q
```
### Success criteria
- Config validation passes for the selected offloading mode
- Training completes without OOM or NCCL errors
- Loss matches the non-offloaded baseline (max delta < 0.001)
- Memory usage drops proportionally to offload fraction
## Code Anchors
### MCore activation offload constraints
```1296:1310:3rdparty/Megatron-LM/megatron/core/transformer/transformer_config.py
if self.cpu_offloading and (
self.cpu_offloading_num_layers < 0 or self.cpu_offloading_num_layers >= self.num_layers
):
raise ValueError(...)
if self.cpu_offloading and self.pipeline_model_parallel_size > 1:
raise ValueError(
"Currently there is no support for Pipeline parallelism with CPU offloading"
)
if self.cpu_offloading and self.recompute_granularity is not None:
raise ValueError(
"CPU offloading does not work when activation recomputation is enabled"
)
```
### MCore CUDA graph incompatibility
```1943:1944:3rdparty/Megatron-LM/megatron/core/transformer/transformer_config.py
if self.cpu_offloading:
raise ValueError("CUDA graphs not supported with CPU offloading.")
```
### MCore fine-grained offloading mutual exclusion
```1427:1430:3rdparty/Megatron-LM/megatron/core/transformer/transformer_config.py
if self.fine_grained_activation_offloading:
assert (
not self.cpu_offloading
), "fine_grained_activation_offloading cannot be enabled with cpu_offloading."
```
### MCore HybridDeviceOptimizer instantiation
```480:518:3rdparty/Megatron-LM/megatron/core/optimizer/__init__.py
if config.optimizer_cpu_offload:
# ... setup cpu/gpu optimizer classes ...
optimizer = HybridDeviceOptimizer(
param_groups,
offload_fraction=config.optimizer_offload_fraction,
cpu_optimizer_cls=cpu_optimizer_cls,
gpu_optimizer_cls=gpu_optimizer_cls,
overlap_cpu_optimizer_d2h_h2d=config.overlap_cpu_optimizer_d2h_h2d,
pin_cpu_grads=config.pin_cpu_grads,
pin_cpu_params=config.pin_cpu_params,
)
```
### Bridge CUDA graph guard
```232:234:src/megatron/bridge/models/gpt_full_te_layer_autocast_spec.py
assert not config.cpu_offloading and config.recompute_granularity is None, "Cudagraphs not supported"
```
### Bridge activation offloading in PEFT
```621:631:src/megatron/bridge/peft/utils.py
if self.config.cpu_offloading and self.config.cpu_offloading_activations:
x.activation_offloading = True
x, _ = self.linear_in(x)
x = self.activation(x)
if self.config.cpu_offloading and self.config.cpu_offloading_activations:
x.activation_offloading = True
x, _ = self.linear_out(x)
```
## Failure Diagnosis
| Symptom | Likely Cause | How To Confirm | Fix |
|---|---|---|---|
| `Currently there is no support for Pipeline parallelism with CPU offloading` | Activation offload + PP > 1 | Check `pipeline_model_parallel_size` | Set PP=1 or use optimizer offloading |
| `CPU offloading does not work when activation recomputation is enabled` | Activation offload + recompute | Check `recompute_granularity` | Set `recompute_granularity=null` |
| `fine_grained_activation_offloading cannot be enabled with cpu_offloading` | Both offloading modes enabled | Check both flags | Use one or the other |
| `CUDA graphs not supported with CPU offloading` | CUDA graphs + activation offload | Check `cuda_graph_impl` | Set `cuda_graph_impl="none"` |
| OOM with activation offloading | Model too large for PP=1 | Check allocated memory vs 80 GB | Use optimizer offloading with PP > 1 |
| Extreme slowdown (>4x) | 100% optimizer offload, CPU Adam bottleneck | Compare iter time at different fractions | Reduce fraction or enable `overlap_cpu_optimizer_d2h_h2d` |
| OOM at partial optimizer offload | Insufficient offload for this config | Check memory at different fractions | Increase fraction or add PP |
## Known Limitations
- Activation offloading requires PP=1, making it impractical for large models
(30B+ MoE) that need pipeline parallelism.
- Optimizer offloading throughput penalty scales linearly (~1.9x at 25%,
~4.2x at 100% for Qwen3-30B-A3B).
- D2H/H2D overlap provides only ~7% speedup because CPU Adam compute is
the dominant bottleneck.
- `fine_grained_activation_offloading` is a separate module-level approach
that works with PP > 1 but cannot be combined with layer-level
`cpu_offloading`.
所有檔案
6 個檔案安裝 nemo-mbridge-perf-cpu-offloading
請下載並將技能檔案解壓縮至您的 .claude/skills/ 目錄中。
下載 ZIP複製儲存庫並將技能檔案複製到您的專案中。
git clone https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-cpu-offloading # Copy SKILL.md to your .claude/skills/ directory
複製





首頁
