選項
首頁首頁 Skill 數據科學與機器學習 nemo-mbridge-perf-moe-vlm-training

nemo-mbridge-perf-moe-vlm-training

NVIDIA/skills NVIDIA/skills

針對在 Megatron Bridge 中訓練專家混合式視覺語言模型提供實用指引,並透過近期多模態實驗的經驗教訓,比較 FSDP 與 3D 平行方法。

...展開全部
1
更新時間 2026-09-28

MoE VLM 訓練

穩定版文件:@docs/training/moe-optimization.md 學習卡:@skills/nemo-mbridge-perf-moe-vlm-training/card.yaml

FSDP 與 3D 平行計算的比較

方法 優勢 最適合
FSDP 實現可運作的多模態執行最簡便的路徑 首次啟動、記憶體優先調校、棘手的 PP 邊界
3D 平行處理 調校後具有更高潛力 具備整潔 PP 佈局且有充足時間進行深度掃描的穩定模型

對於 MoE VLMs,實際的工作流程通常如下:

  1. 使用 FSDP 獲得首次可靠的運行結果
  2. 穩定真實數據輸入、重新計算,並分析記憶體行為
  3. 僅當吞吐量餘裕足以抵銷額外工作量時,才轉向 3D 並行運算

近期 VLM 運行所得的綜合發現

Qwen3-VL 類別模型

在各追蹤器中,主要模式均一致:

  • 在 GB200 級系統上執行 FSDP,即使採用相對簡單的配置, 利用率已能達到相當可觀的 10% 多一點
  • B200 上的 FSDP 運行雖可行,但對重新計算選項及固定 視覺設定的敏感度較高
  • 3D 平行運算可恢復至相似或更佳的運作點,但必須先 共同調整 MBS、重新計算及實際視覺路徑

真實資料與模擬資料

基於模擬資料的 VLM 執行結果並非可靠的效能指標。在實驗中, 相較於真實的多模態輸入,無影像的模擬執行結果看起來更接近「速度約為兩倍」, 而非「略顯樂觀」。

在對 VLM 吞吐量做出任何結論之前,請使用真實或逼真的影像載荷。

規模較小的多模態 MoE 運行

規模較小的 Qwen3.5 風格多模態實驗也印證了相同的教訓:

  • HybridEP 是 GB200 上可靠的預設選項
  • 一旦訓練迴圈趨於穩定,採用 TE 範圍的 CUDA 圖便能發揮作用
  • 更大的 MBS 可能帶來效益,但前提是視覺編碼器不會成為 下一個瓶頸

決策指南

在以下情況下選擇 FSDP:

  • 首次啟動新的 VLM 時
  • 模型在嵌入層、視覺層與解碼器層之間的階段邊界安排較為彆扭時
  • 記憶體適配性比絕對吞吐量更為重要
  • 您可能在進行以解碼器為重點的調優時需要凍結視覺堆疊

在以下情況下選擇 3D 並行模式

  • 當模型在 FSDP 下已趨於穩定時
  • PP 佈局清晰且可重複
  • 您可以同時掃描 MBS、重新計算以及 CUDA 圖譜範圍
  • 目標是最佳穩態吞吐量,而非最簡易的啟動流程

關鍵調整參數

  1. 在適當情況下凍結視覺堆疊:若工作重點在於解碼器, 凍結視覺端通常能帶來微小但確實的吞吐量提升,並 減輕記憶體壓力。

  2. 積極掃描 MBS:相較於僅處理文字的 MoE 運行,VLMs 對 MBS 的敏感度更高, 因為視覺路徑會改變運算與開銷之間的平衡。

  3. 模型擬合完成後,優先採用選擇性重新計算:全量重新計算雖是 有用的啟動工具,但選擇性重新計算通常是更佳的 穩態方案。

  4. 將 CUDA 圖範圍與工作負載相匹配:attn moe_router moe_preprocess 是較安全的 MoE 預設設定,而範圍較窄的設定對於 受控實驗仍可能有所助益。

  5. 僅在 EP 單獨不足時才使用 ETP:它雖能解鎖佈局,但 也會引入更多通訊開銷及更多調優面。

代表性配置家族

FSDP 優先的 GB200 路徑

TP=1  CP=1  PP=1
EP 尺寸依專家拓撲調整,通常較大
調度器:在 GB200 級系統上採用 HybridEP
重新計算:先從全量開始,再放寬至選擇性重新計算

3D 平行 GB200 路徑

TP=1  CP=1  PP=1 或適度 PP
EP 與 ETP 規模依專家拓撲調整
調度器:HybridEP
CUDA 圖:從窄範圍開始,待真實資料路徑穩定後再擴寬

相容性

功能 FSDP 3D 並行
GB200 上的 HybridEP 強預設值 拓撲穩定後採用強力預設
CUDA 圖形 在系統上線後頗具實用性 相當實用,但對作用域較為敏感
畫面凍結 與系統天然契合 雖有可能,但較少被用作主要效能路徑
選擇性重新計算 建議 建議

常見陷阱

  1. 模擬的多模態數據具有誤導性:它可能會使解碼器看起來比 實際的端到端 VLM 路徑更為「健康」。

  2. 視覺編碼器可能會出乎意料地佔據主導地位:在將所有問題歸咎於調度器之前,應分別對編碼器、投影器 和解碼器進行效能分析。

  3. 切勿將有效工作量不同的 FSDP 與 3D-parallel 執行結果相互比較: 應根據有效標記與工作負載形狀進行正規化,而不僅僅是依據步驟時間。

  4. ETP 並非免費:應將其用作擬合或拓撲工具,而非預設選項。

  5. 重新計算與 CUDA 圖的選擇是耦合的:能讓 模型擬合的設定,往往並非能提供最佳穩態速度的設定。

在 GitHub 上查看
---
name: nemo-mbridge-perf-moe-vlm-training
description: Provides practical guidance for training Mixture-of-Experts Vision-Language Models in Megatron Bridge, comparing FSDP and 3D-parallel approaches with lessons from recent multimodal experiments.
license: Apache-2.0
---

# MoE VLM Training

Stable docs: @docs/training/moe-optimization.md
Card: @skills/nemo-mbridge-perf-moe-vlm-training/card.yaml

## FSDP vs 3D Parallel

| Approach | Strength | Best fit |
|---|---|---|
| FSDP | Simplest path to a working multimodal run | first bring-up, memory-first tuning, awkward PP boundaries |
| 3D parallel | Higher ceiling after tuning | stable models with a clean PP layout and time for deeper sweeps |

For MoE VLMs, the practical workflow is usually:

1. get the first reliable run with FSDP
2. stabilize real-data input, recompute, and memory behavior
3. move to 3D parallel only if the throughput headroom is worth the extra work

## Rounded Findings From Recent VLM Runs

### Qwen3-VL class models

The main patterns were consistent across the tracker:

- FSDP on GB200-class systems can already reach healthy high-teens utilization
  with a comparatively simple setup
- B200 FSDP runs are viable, but more sensitive to recompute choice and frozen
  vision settings
- 3D parallel can recover to a similar or better operating point, but only after
  tuning MBS, recompute, and the real vision path together

### Real data vs mock data

Mock-data VLM runs are not trustworthy performance proxies. In the experiments,
image-free mock runs looked closer to "roughly twice as fast" than "slightly
optimistic" when compared with real multimodal input.

Use real or realistic image payloads before drawing any conclusion about VLM
throughput.

### Smaller multimodal MoE runs

The smaller Qwen3.5-style multimodal experiments reinforce the same lessons:

- HybridEP is a solid default on GB200
- TE-scoped CUDA graphs help once the training loop is stable
- larger MBS can pay off, but only if the vision encoder does not become the
  next bottleneck

## Decision Guide

### Choose FSDP when

- you are bringing up a new VLM for the first time
- the model has awkward stage boundaries across embedding, vision, and decoder
- memory fit matters more than absolute throughput
- you may freeze the vision stack during decoder-focused tuning

### Choose 3D parallel when

- the model is already stable under FSDP
- the PP layout is clear and repeatable
- you can sweep MBS, recompute, and CUDA-graph scope together
- the goal is best steady-state throughput, not easiest bring-up

## Key Tuning Knobs

1. **Freeze the vision stack when appropriate**: if the work is decoder-focused,
   freezing the vision side often gives a small but real throughput gain and
   reduces memory pressure.

2. **Sweep MBS aggressively**: VLMs are more MBS-sensitive than text-only MoE
   runs because the vision path changes the compute-to-overhead balance.

3. **Prefer selective recompute once the model fits**: full recompute is a
   useful bring-up tool, but selective recompute is usually the better steady
   state.

4. **Match CUDA-graph scope to the workload**: `attn moe_router moe_preprocess`
   is the safer MoE default, while narrower scopes can still be useful for
   controlled experiments.

5. **Use ETP only when EP alone is insufficient**: it can unlock a layout, but
   it also introduces more communication and more tuning surface.

## Representative Config Families

### FSDP-first GB200 path

```text
TP=1  CP=1  PP=1
EP sized to the expert topology, often large
Dispatcher: HybridEP on GB200-class systems
Recompute: start with full, then relax toward selective recompute
```

### 3D-parallel GB200 path

```text
TP=1  CP=1  PP=1 or modest PP
EP and ETP sized to the expert topology
Dispatcher: HybridEP
CUDA Graph: start narrow, then widen only after the real-data path is stable
```

## Compatibility

| Feature | FSDP | 3D parallel |
|---|---|---|
| HybridEP on GB200 | strong default | strong default once topology is stable |
| CUDA graphs | useful after bring-up | useful, but more scope-sensitive |
| Freeze vision | natural fit | possible, but less often used as the headline perf path |
| Selective recompute | recommended | recommended |

## Pitfalls

1. **Mock multimodal data is misleading**: it can make the decoder look much
   healthier than the real end-to-end VLM path.

2. **The vision encoder can dominate unexpectedly**: profile encoder, projector,
   and decoder separately before attributing everything to the dispatcher.

3. **Do not compare FSDP and 3D-parallel runs with different effective work**:
   normalize by useful tokens and workload shape, not only by step time.

4. **ETP is not free**: use it as a fit or topology tool, not as the default.

5. **Recompute and CUDA-graph choices are coupled**: the setting that gets the
   model to fit is often not the setting that gives the best steady-state speed.

所有檔案

1 個檔案

安裝 nemo-mbridge-perf-moe-vlm-training

請下載並將技能檔案解壓縮至您的 .claude/skills/ 目錄中。

下載 ZIP

複製儲存庫並將技能檔案複製到您的專案中。

git clone https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-moe-vlm-training # Copy SKILL.md to your .claude/skills/ directory

複製 複製
快速設定: 將技能資料夾複製到 .claude/skills/ Claude 會自動偵測並使用該技能
儲存庫 NVIDIA/skills

相關技能

web-search
更新時間 2026-06-29
webapp-testing
更新時間 2026-06-29
lark-base
更新時間 2026-07-05
agentmail
更新時間 2026-06-29
OR