옵션
집집 Skill 데이터 과학 및 ML nemo-mbridge-perf-moe-vlm-training

nemo-mbridge-perf-moe-vlm-training

NVIDIA/skills NVIDIA/skills

Megatron Bridge에서 전문가 혼합(Mixture-of-Experts) 비전-언어 모델을 훈련하기 위한 실용적인 지침을 제공하며, FSDP 및 3D-병렬 접근법을 비교하고 최근 다모달 실험에서 얻은 교훈을 제시합니다.

...모든 것을 확장하십시오
1
업데이트 된 시간 2026년 9월 28일

MoE VLM 훈련

안정화 문서: @docs/training/moe-optimization.md 카드: @skills/nemo-mbridge-perf-moe-vlm-training/card.yaml

FSDP 대 3D 병렬

접근 방식 장점 가장 적합한
FSDP 다중 모드 실행을 위한 가장 간단한 방법 초기 가동, 메모리 우선 튜닝, 부자연스러운 PP 경계
3D 병렬 튜닝 후 더 높은 성능 한계 깔끔한 PP 레이아웃을 갖춘 안정적인 모델과 심층적인 탐색을 위한 시간 확보

MoE VLM의 경우, 실제 워크플로는 대개 다음과 같습니다:

  1. FSDP를 사용하여 첫 번째 신뢰할 수 있는 실행 결과 확보
  2. 실제 데이터 입력을 안정화하고, 재계산하며, 메모리 동작을 확인
  3. 처리량 여유분이 추가 작업에 비해 가치가 있을 때만 3D 병렬 처리로 전환

최근 VLM 실행 결과 요약

Qwen3-VL 클래스 모델

트래커 전반에 걸쳐 주요 패턴은 일관되었습니다:

  • GB200급 시스템에서 FSDP는 비교적 간단한 설정만으로도 이미 10% 후반대의 양호한 활용도를 달성할 수 있음
  • B200 FSDP 실행은 가능하지만, 재계산 선택 및 고정된 비전 설정에 더 민감합니다
  • 3D 병렬 처리는 유사하거나 더 나은 작동 지점으로 회복될 수 있으나, MBS, 재계산, 실제 비전 경로를 함께 튜닝한 후에야 가능합니다

실제 데이터 대 모의 데이터

모의 데이터를 사용한 VLM 실행 결과는 신뢰할 만한 성능 지표가 아닙니다. 실험 결과, 이미지가 포함되지 않은 모의 실행은 실제 다중 모달 입력과 비교했을 때 "약간 낙관적인" 수준보다는 "대략 두 배 정도 빠른" 수준에 더 가깝게 나타났습니다.

VLM 처리량에 대한 결론을 내리기 전에는 실제 또는 현실적인 이미지 페이로드를 사용해야 합니다.

소규모 다모달 MoE 실행

Qwen3.5 스타일의 소규모 다중 모달 실험에서도 동일한 교훈이 확인됩니다:

  • HybridEP는 GB200에서 안정적인 기본 설정입니다.
  • 훈련 루프가 안정화되면 TE 기반 CUDA 그래프가 도움이 됩니다
  • 더 큰 MBS는 효과가 있을 수 있지만, 비전 인코더가 다음 병목 현상이 되지 않는 경우에만 해당됩니다

결정 가이드

다음과 같은 경우에는 FSDP를 선택하십시오.

  • 새로운 VLM을 처음 구동할 때
  • 모델의 임베딩, 비전, 디코더 단계 간 경계가 복잡할 때
  • 절대 처리량보다 메모리 적합성이 더 중요한 경우
  • 디코더 중심 튜닝 중에 비전 스택을 일시 정지해야 할 수 있는 경우

다음과 같은 경우에는 3D 병렬을 선택하십시오.

  • 모델이 FSDP 환경에서 이미 안정화된 경우
  • PP 레이아웃이 명확하고 재현 가능할 때
  • MBS를 스윕하고, 재계산하며, CUDA 그래프 스코프를 함께 확인할 수 있는 경우
  • 목표가 가장 쉬운 초기 구동이 아닌, 최상의 정상 상태 처리량인 경우

주요 튜닝 요소

  1. 적절한 경우 비전 스택을 고정하십시오: 작업이 디코더에 집중되어 있다면, 비전 측을 고정하는 것이 작지만 실질적인 처리량 향상을 가져오고 메모리 부하를 줄여줍니다.

  2. MBS를 적극적으로 스윕하십시오: 비전 경로가 연산 대 오버헤드 균형을 변화시키기 때문에, VLMs는 텍스트 전용 MoE 실행보다 MBS에 더 민감합니다.

  3. 모델이 잘 맞으면 선택적 재계산을 선호하세요: 전체 재계산은 가동 초기 단계에서 유용한 도구이지만, 선택적 재계산이 일반적으로 더 나은 정상 상태를 보장합니다.

  4. CUDA 그래프 범위를 워크로드에 맞추십시오: attn moe_router moe_preprocess 는 더 안전한 MoE 기본 설정이지만, 통제된 실험을 위해서는 더 좁은 범위도 여전히 유용할 수 있습니다.

  5. EP만으로는 불충분한 경우에만 ETP를 사용하십시오: ETP는 레이아웃의 잠재력을 끌어낼 수 있지만, 동시에 더 많은 통신과 더 많은 튜닝 영역을 유발합니다.

대표적인 구성 계열

FSDP 우선 GB200 경로

TP=1  CP=1  PP=1
EP 크기는 전문가 토폴로지에 맞춰지며, 대개 큼
디스패처: GB200급 시스템에서 하이브리드 EP 사용
재계산: 전체 재계산으로 시작하여, 이후 선택적 재계산으로 완화

3D 병렬 GB200 경로

TP=1  CP=1  PP=1 또는 적당한 수준의 PP
EP 및 ETP는 전문가 토폴로지에 맞춰 규모 설정
디스패처: HybridEP
CUDA 그래프: 좁게 시작하여, 실제 데이터 경로가 안정화된 후에만 확장

호환성

기능 FSDP 3D 병렬
GB200에서의 하이브리드 EP 강한 기본값 토폴로지가 안정화되면 강력한 기본값
CUDA 그래프 가동 후 유용함 유용하지만, 적용 범위에 더 민감함
비전 동결 자연스럽게 어울림 가능하지만, 주요 성능 최적화 경로로는 덜 사용됨
선택적 재계산 권장 권장

주의할 점

  1. 모의 다중 모달 데이터는 오해를 불러일으킬 수 있습니다: 이로 인해 디코더가 실제 엔드투엔드 VLM 경로보다 훨씬 더 양호한 것처럼 보일 수 있습니다.

  2. 비전 인코더가 예상치 못하게 큰 비중을 차지할 수 있습니다: 모든 것을 디스패처의 탓으로 돌리기 전에 인코더, 프로젝터, 디코더를 각각 별도로 프로파일링하십시오.

  3. 유효 작업량이 다른 FSDP 및 3D-병렬 실행 결과를 비교하지 마십시오: 단순히 단계 시간뿐만 아니라 유용한 토큰과 워크로드 형태를 기준으로 정규화하십시오.

  4. ETP는 무료가 아닙니다: 기본값으로 사용하기보다는 모델 적합도 또는 토폴로지 조정 도구로 활용하십시오.

  5. 재계산과 CUDA 그래프 선택은 상호 연관되어 있습니다: 모델이 적합하게 맞도록 하는 설정이 반드시 최상의 정상 상태 속도를 제공하는 설정은 아닙니다.

GitHub에서 보기
---
name: nemo-mbridge-perf-moe-vlm-training
description: Provides practical guidance for training Mixture-of-Experts Vision-Language Models in Megatron Bridge, comparing FSDP and 3D-parallel approaches with lessons from recent multimodal experiments.
license: Apache-2.0
---

# MoE VLM Training

Stable docs: @docs/training/moe-optimization.md
Card: @skills/nemo-mbridge-perf-moe-vlm-training/card.yaml

## FSDP vs 3D Parallel

| Approach | Strength | Best fit |
|---|---|---|
| FSDP | Simplest path to a working multimodal run | first bring-up, memory-first tuning, awkward PP boundaries |
| 3D parallel | Higher ceiling after tuning | stable models with a clean PP layout and time for deeper sweeps |

For MoE VLMs, the practical workflow is usually:

1. get the first reliable run with FSDP
2. stabilize real-data input, recompute, and memory behavior
3. move to 3D parallel only if the throughput headroom is worth the extra work

## Rounded Findings From Recent VLM Runs

### Qwen3-VL class models

The main patterns were consistent across the tracker:

- FSDP on GB200-class systems can already reach healthy high-teens utilization
  with a comparatively simple setup
- B200 FSDP runs are viable, but more sensitive to recompute choice and frozen
  vision settings
- 3D parallel can recover to a similar or better operating point, but only after
  tuning MBS, recompute, and the real vision path together

### Real data vs mock data

Mock-data VLM runs are not trustworthy performance proxies. In the experiments,
image-free mock runs looked closer to "roughly twice as fast" than "slightly
optimistic" when compared with real multimodal input.

Use real or realistic image payloads before drawing any conclusion about VLM
throughput.

### Smaller multimodal MoE runs

The smaller Qwen3.5-style multimodal experiments reinforce the same lessons:

- HybridEP is a solid default on GB200
- TE-scoped CUDA graphs help once the training loop is stable
- larger MBS can pay off, but only if the vision encoder does not become the
  next bottleneck

## Decision Guide

### Choose FSDP when

- you are bringing up a new VLM for the first time
- the model has awkward stage boundaries across embedding, vision, and decoder
- memory fit matters more than absolute throughput
- you may freeze the vision stack during decoder-focused tuning

### Choose 3D parallel when

- the model is already stable under FSDP
- the PP layout is clear and repeatable
- you can sweep MBS, recompute, and CUDA-graph scope together
- the goal is best steady-state throughput, not easiest bring-up

## Key Tuning Knobs

1. **Freeze the vision stack when appropriate**: if the work is decoder-focused,
   freezing the vision side often gives a small but real throughput gain and
   reduces memory pressure.

2. **Sweep MBS aggressively**: VLMs are more MBS-sensitive than text-only MoE
   runs because the vision path changes the compute-to-overhead balance.

3. **Prefer selective recompute once the model fits**: full recompute is a
   useful bring-up tool, but selective recompute is usually the better steady
   state.

4. **Match CUDA-graph scope to the workload**: `attn moe_router moe_preprocess`
   is the safer MoE default, while narrower scopes can still be useful for
   controlled experiments.

5. **Use ETP only when EP alone is insufficient**: it can unlock a layout, but
   it also introduces more communication and more tuning surface.

## Representative Config Families

### FSDP-first GB200 path

```text
TP=1  CP=1  PP=1
EP sized to the expert topology, often large
Dispatcher: HybridEP on GB200-class systems
Recompute: start with full, then relax toward selective recompute
```

### 3D-parallel GB200 path

```text
TP=1  CP=1  PP=1 or modest PP
EP and ETP sized to the expert topology
Dispatcher: HybridEP
CUDA Graph: start narrow, then widen only after the real-data path is stable
```

## Compatibility

| Feature | FSDP | 3D parallel |
|---|---|---|
| HybridEP on GB200 | strong default | strong default once topology is stable |
| CUDA graphs | useful after bring-up | useful, but more scope-sensitive |
| Freeze vision | natural fit | possible, but less often used as the headline perf path |
| Selective recompute | recommended | recommended |

## Pitfalls

1. **Mock multimodal data is misleading**: it can make the decoder look much
   healthier than the real end-to-end VLM path.

2. **The vision encoder can dominate unexpectedly**: profile encoder, projector,
   and decoder separately before attributing everything to the dispatcher.

3. **Do not compare FSDP and 3D-parallel runs with different effective work**:
   normalize by useful tokens and workload shape, not only by step time.

4. **ETP is not free**: use it as a fit or topology tool, not as the default.

5. **Recompute and CUDA-graph choices are coupled**: the setting that gets the
   model to fit is often not the setting that gives the best steady-state speed.

모든 파일

1개 파일

nemo-mbridge-perf-moe-vlm-training 설치

스킬 파일을 다운로드하여 .claude/skills/ 디렉터리에 압축을 풀어주세요.

ZIP 다운로드

저장소를 클론하고 스킬 파일을 프로젝트에 복사하세요.

git clone https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-moe-vlm-training # Copy SKILL.md to your .claude/skills/ directory

복사 복사
빠른 설정: 스킬 폴더를 .claude/skills/로 복사하세요. Claude가 해당 스킬을 자동으로 감지하여 사용할 것입니다.
저장소 NVIDIA/skills

관련 스킬

web-search
업데이트 된 시간 2026년 6월 29일
webapp-testing
업데이트 된 시간 2026년 6월 29일
lark-base
업데이트 된 시간 2026년 7월 5일
agentmail
업데이트 된 시간 2026년 6월 29일
OR