옵션
집집 Skill 데이터 과학 및 ML nemo-mbridge-perf-cpu-offloading

nemo-mbridge-perf-cpu-offloading

NVIDIA/skills NVIDIA/skills

HybridDeviceOptimizer를 활용한 활성화 오프로딩 및 최적화기 상태 오프로딩을 포함하여, Megatron Bridge 훈련을 위한 CPU 오프로딩을 구성하고 검증합니다.

...모든 것을 확장하십시오
2
업데이트 된 시간 2026년 9월 28일

CPU 부하 분산

참고 문헌

  • 안정화 문서: @docs/training/cpu-offloading.md
  • 구조화된 메타데이터: @skills/nemo-mbridge-perf-cpu-offloading/card.yaml

개요

GPU에서 CPU 메모리로 데이터를 이동시키는 두 가지 독립적인 메커니즘:

메커니즘 구성 네임스페이스 오프로드되는 항목 PP 제한
활성화 오프로딩 model.cpu_offloading* 트랜스포머 레이어별 활성화(및 선택적으로 가중치) PP는 1이어야 함
최적화기 오프로딩 optimizer.optimizer_cpu_offload HybridDeviceOptimizer를 통한 Adam 최적화기 상태(모멘텀 + 분산) 없음

신속한 결정

상황 권장 사항
대규모 MoE 모델(30B+), PP > 1 필요 최적화기 오프로딩 — PP=1로 인해 활성화 오프로딩이 차단됨
소형/중형 모델, PP=1이 적합하며, 활성화 메모리 사용량이 주를 이룸 활성화 오프로딩
메모리-속도 트레이드오프를 조정할 수 있기를 원함 optimizer_offload_fraction 매개변수를 사용한 최적화기 오프로딩
처리량이 최우선 순위 활성화하지 마십시오 — 오프로딩은 항상 오버헤드를 발생시킵니다
CUDA 그래프가 필요합니다 오직 최적화기 오프로딩만 가능 — 활성화 오프로딩은 호환되지 않음
메모리 부하가 중간 수준입니다 최적의 효율을 위해 최적화기 오프로딩 비율을 25~50%로 설정하십시오

활성화

최적화기 CPU 오프로딩 (대규모 모델에 권장)

cfg.optimizer.optimizer_cpu_offload = True
cfg.optimizer.optimizer_offload_fraction = 1.0
cfg.optimizer.overlap_cpu_optimizer_d2h_h2d = True

CLI 오버라이드:

optimizer.optimizer_cpu_offload=True \
optimizer.optimizer_offload_fraction=0.5 \
optimizer.overlap_cpu_optimizer_d2h_h2d=True

활성화 함수 CPU 오프로딩 (소형/중형 모델에만 해당)

cfg.model.cpu_offloading = True
cfg.model.cpu_offloading_num_layers = 16
cfg.model.cpu_offloading_activations = True
cfg.model.cpu_offloading_weights = False

cfg.model.pipeline_model_parallel_size = 1
cfg.model.recompute_granularity = None
cfg.model.cuda_graph_impl = "none"

구성 매개변수 참조

최적화기 오프로딩

매개변수 기본값 설명
optimizer_cpu_offload False 마스터 스위치
optimizer_offload_fraction 0.0 CPU에서 실행되는 최적화기 상태의 비율 (0.0–1.0)
overlap_cpu_optimizer_d2h_h2d False GPU↔CPU 전송을 연산과 중첩
use_torch_optimizer_for_cpu_offload False CPU 부분에서 융합 최적화기 대신 torch.optim 사용

활성화 함수 오프로딩

매개변수 기본값 설명
cpu_offloading False 마스터 스위치
cpu_offloading_num_layers 0 오프로딩할 트랜스포머 레이어 수 (0부터 num_layers-1까지)
cpu_offloading_activations True 활성화 함수 오프로딩
cpu_offloading_weights False 가중치 오프로딩
cpu_offloading_double_buffering False 재로드 시 레이어 간 더블 버퍼링

호환성 및 제약 사항

활성화 오프로딩

  • pipeline_model_parallel_size는 1이어야 함
  • recompute_granularity는 None이어야 함
  • fine_grained_activation_offloading과 결합할 수 없음
  • CUDA 그래프와 함께 사용할 수 없음
  • cpu_offloading_num_layers는 [0, num_layers-1) 범위여야 합니다

최적화기 오프로딩

  • use_distributed_optimizer = True 여야 함(대부분의 레시피에서 기본값)
  • PP, 재계산 또는 CUDA 그래프에 대한 제한 없음
  • optimizer_offload_fraction은 [0.0, 1.0] 범위 내에 있어야 합니다

실무 적용: 대규모 MoE 모델

Qwen3-30B-A3B 및 이와 유사한 대규모 MoE 모델의 경우 활성화 함수 오프로딩이 차단됩니다. PP=1 제약 조건으로 인해 각 GPU가 48개 레이어를 모두 보유하게 되며, 모델 가중치와 최적화기 상태만으로도 (~70 GB) H100의 80 GB 용량을 초과합니다.

최소 실행 가능 명령어

uv run python scripts/training/run_recipe.py \
  --recipe qwen3_30b_a3b_pretrain_config \
  optimizer.optimizer_cpu_offload=True \
  optimizer.optimizer_offload_fraction=0.5 \
  train.train_iters=20 \
  train.global_batch_size=8 \
  train.micro_batch_size=1

검증

단위 테스트

uv run python -m pytest \
  tests/unit_tests/models/test_gpt_full_te_layer_autocast_spec.py -k "cpu_offload" \
  tests/unit_tests/peft/test_utils.py -k "cpu_offload" -q

성공 기준

  • 선택한 오프로딩 모드에 대한 구성 유효성 검사가 통과됨
  • OOM 또는 NCCL 오류 없이 훈련이 완료됩니다
  • 손실 값이 오프로딩되지 않은 기준값과 일치함(최대 차이 < 0.001)
  • 메모리 사용량이 오프로딩 비율에 비례하여 감소합니다

코드 앵커

MCore 활성화 오프로딩 제약 조건

       if self.cpu_offloading and (
            self.cpu_offloading_num_layers < 0 or self.cpu_offloading_num_layers >=self.num_layers
        ):
            raise ValueError(...)

        if self.cpu_offloading and self.pipeline_model_parallel_size > 1:
            raise ValueError(
                "현재 CPU 오프로딩과 파이프라인 병렬 처리를 동시에 지원하는 기능은 제공되지 않습니다"
            )

        if self.cpu_offloading and self.recompute_granularity is not None:
            raise ValueError(
                "활성화 재계산이 활성화된 상태에서는 CPU 오프로딩이 작동하지 않습니다"
            )

MCore CUDA 그래프 비호환성

           if self.cpu_offloading:
                raise ValueError("CPU 오프로딩 시 CUDA 그래프가 지원되지 않습니다.")

MCore 세분화된 오프로딩 상호 배제

       if self.fine_grained_activation_offloading:
            assert (
                not self.cpu_offloading
            ), "cpu_offloading이 활성화된 상태에서는 fine_grained_activation_offloading을 활성화할 수 없습니다."

MCore HybridDeviceOptimizer 인스턴스화

       if config.optimizer_cpu_offload:
            # ... CPU/GPU 최적화기 클래스 설정 ...
            optimizer = HybridDeviceOptimizer(
                param_groups,
                offload_fraction=config.optimizer_offload_fraction,
                cpu_optimizer_cls=cpu_optimizer_cls,
                gpu_optimizer_cls=gpu_optimizer_cls,
                overlap_cpu_optimizer_d2h_h2d=config.overlap_cpu_optimizer_d2h_h2d,
                pin_cpu_grads=config.pin_cpu_grads,
                pin_cpu_params=config.pin_cpu_params,
            )

Bridge CUDA 그래프 가드

       assert config.cpu_offloading이 아니며 config.recompute_granularity가 None인 경우, "Cudagraphs 미지원"

PEFT에서 활성화 오프로딩 연결

       if self.config.cpu_offloading and self.config.cpu_offloading_activations:
            x.activation_offloading = True
        x, _ = self.linear_in(x)
        x = self.activation(x)
        if self.config.cpu_offloading and self.config.cpu_offloading_activations:
            x.activation_offloading = True
        x, _ = self.linear_out(x)

오류 진단

증상 가능한 원인 확인 방법 해결 방법
현재 CPU 오프로딩을 활용한 파이프라인 병렬 처리는 지원되지 않습니다 활성화 오프로딩 + PP > 1 pipeline_model_parallel_size 확인 PP=1로 설정하거나 최적화기 오프로딩을 사용하십시오
활성화 재계산이 활성화된 경우 CPU 오프로딩이 작동하지 않습니다 활성화 오프로드 + 재계산 recompute_granularity 확인 recompute_granularity를 null로 설정하십시오
cpu_offloading이 활성화된 상태에서는 fine_grained_activation_offloading을 활성화할 수 없습니다 두 오프로딩 모드 모두 활성화됨 두 플래그 모두 확인 둘 중 하나만 사용하십시오
CPU 오프로딩에서는 CUDA 그래프가 지원되지 않습니다 CUDA 그래프 + 활성화 오프로딩 cuda_graph_impl 확인 cuda_graph_impl="none"으로 설정하십시오
활성화 오프로드 시 OOM 발생 PP=1에 비해 모델이 너무 큽니다 할당된 메모리를 80 GB와 비교하여 확인하십시오 PP > 1일 때 최적화기 오프로딩 사용
극심한 속도 저하 (>4배) 100% 최적화기 오프로딩, CPU Adam 병목 현상 서로 다른 분율에서 반복 시간 비교 분율을 줄이거나 overlap_cpu_optimizer_d2h_h2d를 활성화하십시오
부분 최적화기 오프로드 시 OOM 발생 이 구성에 대한 오프로드 부족 각기 다른 분율에서 메모리 상태 확인 분수를 높이거나 PP를 추가하십시오

알려진 제한 사항

  • 활성화 오프로딩에는 PP=1이 필요하므로, 파이프라인 병렬 처리가 필요한 대규모 모델 (30B+ MoE)의 경우 실용적이지 않습니다.
  • 최적화기 오프로딩의 처리량 손실은 선형적으로 증가합니다(Qwen3-30B-A3B의 경우 25%에서 ~1.9배, 100%에서 ~4.2배).
  • D2H/H2D 중첩은 CPU Adam 연산이 주요 병목 현상이기 때문에 약 7%의 속도 향상만 제공합니다.
  • fine_grained_activation_offloading은 PP > 1일 때 작동하는 별도의 모듈 수준 접근 방식이지만, 레이어 수준의 cpu_offloading과는 결합할 수 없습니다.
GitHub에서 보기
---
name: nemo-mbridge-perf-cpu-offloading
description: Configure and validate CPU offloading for Megatron Bridge training, including activation offloading and optimizer state offloading with HybridDeviceOptimizer.
license: Apache-2.0
---

# CPU Offloading

## References

- Stable docs: @docs/training/cpu-offloading.md
- Structured metadata: @skills/nemo-mbridge-perf-cpu-offloading/card.yaml

## What It Is

Two independent mechanisms to move data from GPU to CPU memory:

| Mechanism | Config namespace | What gets offloaded | PP restriction |
|---|---|---|---|
| Activation offloading | `model.cpu_offloading*` | Activations (and optionally weights) per transformer layer | PP must be 1 |
| Optimizer offloading | `optimizer.optimizer_cpu_offload` | Adam optimizer states (momentum + variance) via `HybridDeviceOptimizer` | None |

## Quick Decision

| Situation | Recommendation |
|---|---|
| Large MoE model (30B+), needs PP > 1 | Optimizer offloading — activation offloading is blocked by PP=1 |
| Small/medium model, PP=1 fits, activation memory dominates | Activation offloading |
| Want tunable memory-speed tradeoff | Optimizer offloading with fractional `optimizer_offload_fraction` |
| Throughput is top priority | Don't enable — offloading always adds overhead |
| CUDA graphs are needed | Only optimizer offloading — activation offloading is incompatible |
| Memory pressure is moderate | Optimizer offload at 25–50% fraction for best efficiency |

## Enablement

### Optimizer CPU offloading (recommended for large models)

```python
cfg.optimizer.optimizer_cpu_offload = True
cfg.optimizer.optimizer_offload_fraction = 1.0
cfg.optimizer.overlap_cpu_optimizer_d2h_h2d = True
```

CLI overrides:

```bash
optimizer.optimizer_cpu_offload=True \
optimizer.optimizer_offload_fraction=0.5 \
optimizer.overlap_cpu_optimizer_d2h_h2d=True
```

### Activation CPU offloading (small/medium models only)

```python
cfg.model.cpu_offloading = True
cfg.model.cpu_offloading_num_layers = 16
cfg.model.cpu_offloading_activations = True
cfg.model.cpu_offloading_weights = False

cfg.model.pipeline_model_parallel_size = 1
cfg.model.recompute_granularity = None
cfg.model.cuda_graph_impl = "none"
```

## Config Parameter Reference

### Optimizer offloading

| Parameter | Default | Description |
|-----------|---------|-------------|
| `optimizer_cpu_offload` | `False` | Master switch |
| `optimizer_offload_fraction` | `0.0` | Fraction of optimizer states on CPU (0.0–1.0) |
| `overlap_cpu_optimizer_d2h_h2d` | `False` | Overlap GPU↔CPU transfers with compute |
| `use_torch_optimizer_for_cpu_offload` | `False` | Use `torch.optim` instead of fused optimizer for CPU portion |

### Activation offloading

| Parameter | Default | Description |
|-----------|---------|-------------|
| `cpu_offloading` | `False` | Master switch |
| `cpu_offloading_num_layers` | `0` | Number of transformer layers to offload (0 to num_layers-1) |
| `cpu_offloading_activations` | `True` | Offload activations |
| `cpu_offloading_weights` | `False` | Offload weights |
| `cpu_offloading_double_buffering` | `False` | Double-buffer across layers while reloading |

## Compatibility And Constraints

### Activation offloading

- `pipeline_model_parallel_size` must be 1
- `recompute_granularity` must be `None`
- Cannot combine with `fine_grained_activation_offloading`
- Cannot combine with CUDA graphs
- `cpu_offloading_num_layers` must be in `[0, num_layers-1)`

### Optimizer offloading

- Requires `use_distributed_optimizer = True` (default in most recipes)
- No PP, recompute, or CUDA graph restrictions
- `optimizer_offload_fraction` must be in `[0.0, 1.0]`

### Practical: large MoE models

Activation offloading is blocked for Qwen3-30B-A3B and similar large MoE
models. The PP=1 constraint means each GPU holds all 48 layers; model
weights + optimizer states alone (~70 GB) exceed H100 80 GB capacity.

## Minimal Runnable Command

```bash
uv run python scripts/training/run_recipe.py \
  --recipe qwen3_30b_a3b_pretrain_config \
  optimizer.optimizer_cpu_offload=True \
  optimizer.optimizer_offload_fraction=0.5 \
  train.train_iters=20 \
  train.global_batch_size=8 \
  train.micro_batch_size=1
```

## Verification

### Unit tests

```bash
uv run python -m pytest \
  tests/unit_tests/models/test_gpt_full_te_layer_autocast_spec.py -k "cpu_offload" \
  tests/unit_tests/peft/test_utils.py -k "cpu_offload" -q
```

### Success criteria

- Config validation passes for the selected offloading mode
- Training completes without OOM or NCCL errors
- Loss matches the non-offloaded baseline (max delta < 0.001)
- Memory usage drops proportionally to offload fraction

## Code Anchors

### MCore activation offload constraints

```1296:1310:3rdparty/Megatron-LM/megatron/core/transformer/transformer_config.py
        if self.cpu_offloading and (
            self.cpu_offloading_num_layers < 0 or self.cpu_offloading_num_layers >= self.num_layers
        ):
            raise ValueError(...)

        if self.cpu_offloading and self.pipeline_model_parallel_size > 1:
            raise ValueError(
                "Currently there is no support for Pipeline parallelism with CPU offloading"
            )

        if self.cpu_offloading and self.recompute_granularity is not None:
            raise ValueError(
                "CPU offloading does not work when activation recomputation is enabled"
            )
```

### MCore CUDA graph incompatibility

```1943:1944:3rdparty/Megatron-LM/megatron/core/transformer/transformer_config.py
            if self.cpu_offloading:
                raise ValueError("CUDA graphs not supported with CPU offloading.")
```

### MCore fine-grained offloading mutual exclusion

```1427:1430:3rdparty/Megatron-LM/megatron/core/transformer/transformer_config.py
        if self.fine_grained_activation_offloading:
            assert (
                not self.cpu_offloading
            ), "fine_grained_activation_offloading cannot be enabled with cpu_offloading."
```

### MCore HybridDeviceOptimizer instantiation

```480:518:3rdparty/Megatron-LM/megatron/core/optimizer/__init__.py
        if config.optimizer_cpu_offload:
            # ... setup cpu/gpu optimizer classes ...
            optimizer = HybridDeviceOptimizer(
                param_groups,
                offload_fraction=config.optimizer_offload_fraction,
                cpu_optimizer_cls=cpu_optimizer_cls,
                gpu_optimizer_cls=gpu_optimizer_cls,
                overlap_cpu_optimizer_d2h_h2d=config.overlap_cpu_optimizer_d2h_h2d,
                pin_cpu_grads=config.pin_cpu_grads,
                pin_cpu_params=config.pin_cpu_params,
            )
```

### Bridge CUDA graph guard

```232:234:src/megatron/bridge/models/gpt_full_te_layer_autocast_spec.py
        assert not config.cpu_offloading and config.recompute_granularity is None, "Cudagraphs not supported"
```

### Bridge activation offloading in PEFT

```621:631:src/megatron/bridge/peft/utils.py
        if self.config.cpu_offloading and self.config.cpu_offloading_activations:
            x.activation_offloading = True
        x, _ = self.linear_in(x)
        x = self.activation(x)
        if self.config.cpu_offloading and self.config.cpu_offloading_activations:
            x.activation_offloading = True
        x, _ = self.linear_out(x)
```

## Failure Diagnosis

| Symptom | Likely Cause | How To Confirm | Fix |
|---|---|---|---|
| `Currently there is no support for Pipeline parallelism with CPU offloading` | Activation offload + PP > 1 | Check `pipeline_model_parallel_size` | Set PP=1 or use optimizer offloading |
| `CPU offloading does not work when activation recomputation is enabled` | Activation offload + recompute | Check `recompute_granularity` | Set `recompute_granularity=null` |
| `fine_grained_activation_offloading cannot be enabled with cpu_offloading` | Both offloading modes enabled | Check both flags | Use one or the other |
| `CUDA graphs not supported with CPU offloading` | CUDA graphs + activation offload | Check `cuda_graph_impl` | Set `cuda_graph_impl="none"` |
| OOM with activation offloading | Model too large for PP=1 | Check allocated memory vs 80 GB | Use optimizer offloading with PP > 1 |
| Extreme slowdown (>4x) | 100% optimizer offload, CPU Adam bottleneck | Compare iter time at different fractions | Reduce fraction or enable `overlap_cpu_optimizer_d2h_h2d` |
| OOM at partial optimizer offload | Insufficient offload for this config | Check memory at different fractions | Increase fraction or add PP |

## Known Limitations

- Activation offloading requires PP=1, making it impractical for large models
  (30B+ MoE) that need pipeline parallelism.
- Optimizer offloading throughput penalty scales linearly (~1.9x at 25%,
  ~4.2x at 100% for Qwen3-30B-A3B).
- D2H/H2D overlap provides only ~7% speedup because CPU Adam compute is
  the dominant bottleneck.
- `fine_grained_activation_offloading` is a separate module-level approach
  that works with PP > 1 but cannot be combined with layer-level
  `cpu_offloading`.

nemo-mbridge-perf-cpu-offloading 설치

스킬 파일을 다운로드하여 .claude/skills/ 디렉터리에 압축을 풀어주세요.

ZIP 다운로드

저장소를 클론하고 스킬 파일을 프로젝트에 복사하세요.

git clone https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-cpu-offloading # Copy SKILL.md to your .claude/skills/ directory

복사 복사
빠른 설정: 스킬 폴더를 .claude/skills/로 복사하세요. Claude가 해당 스킬을 자동으로 감지하여 사용합니다.
저장소 NVIDIA/skills

관련 스킬

web-search
업데이트 된 시간 2026년 6월 29일
webapp-testing
업데이트 된 시간 2026년 6월 29일
lark-base
업데이트 된 시간 2026년 7월 5일
agentmail
업데이트 된 시간 2026년 6월 29일
OR