選項
首頁首頁 Skill 數據科學與機器學習 tao-finetune-huggingface-model

tao-finetune-huggingface-model

NVIDIA/skills NVIDIA/skills

使用 NGC PyTorch 容器在本地 NVIDIA GPU 上微調 HuggingFace 的 CV、VLM 或 LLM 模型,支援全量或 LoRA 訓練、資料集處理,以及可選的模型推送到 Hub。

...展開全部
1
更新時間 2026-09-29

tao-finetune-huggingface-model

針對 HuggingFace 模型在本地 NVIDIA GPU 上進行微調,以實時獲取的文件為基礎,並以經過策劃的參考資料作為備用安全網。一個 NGC 容器,幾個專注的指令碼,一次推送到 HF Hub。請遵循本檔案中的規則;不要自行發揮。

權威順序(優先順序最高在前):

  1. 使用者輸入 — 明確的 model_id、dataset_id、training_method、config.yaml 覆蓋。
  2. 實時研究 — 模型卡片、HF 倉庫示例、作者微調指令碼、HF 任務文件、論文;始終獲取(第 3 步 + references/research-priorities.md)。
  3. 策劃的參考資料(references/*.md) — 當實時研究沉默或模稜兩可時的後備方案。
  4. 你的訓練資料記憶 — 最後手段;存疑,需與 (2)/(3) 交叉核對。

(2) 和 (3) 之間的衝突解決以及源行差異說明位於 references/research-priorities.md。

輸入

必需:

  • model_id — HuggingFace 模型 ID,例如 google/vit-base-patch16-224

條件憑據(從會話環境讀取,在啟動前匯出,如果存在):

  • HF_TOKEN — 僅當模型/資料集是受限(讀取)或 push_to_hub 開啟(寫入)時;公開 + 公開 + push_to_hub: false 不需要。值從不讀取 — 僅透過 [ -n "$HF_TOKEN" ] 檢查存在性。
  • WANDB_API_KEY, WANDB_PROJECT — 僅當啟用 WandB 時;WANDB_MODE=disabled 選擇退出。

資料集 — 恰好一個:

  • dataset_id — HuggingFace 資料集 ID (來源:hf)
  • local_dataset_path — 本地資料夾或檔案 (來源:local);可選 local_dataset_format ∈ {auto, imagefolder, coco, voc, jsonl, arrow, parquet, csv}(預設:自動檢測)。
  • (省略) — 代理推薦流行資料集 (來源:recommend)

可選(有預設值):

  • task_type — 從配置 + 模型卡片自動檢測
  • n_train=10000, n_eval=1000, n_epochs=3, lora_r=16
  • output_dir=./output/<model_short_name></model_short_name>
  • hf_model_repo — 推送目標;如果未設定且 HF_TOKEN 具有寫入許可權,自動派生為 <whoami>/<model_short_name>-finetuned</model_short_name></whoami>。
  • push_to_hub=True — 設定為 False 以跳過
  • skip_baseline=False — 跳過零樣本基線評估

可選交付物(預設關閉):

emit_progress_log: false   # output_dir/PROGRESS.md(逐步日誌)
emit_report:       false   # reports/report.{pdf,html} 包含曲線和樣本
emit_unit_tests:   false   # tests/ 包含假資料異構批次測試

所有值位於 output_dir/config.yaml。切勿在 Python 中硬編碼。

執行平臺

此技能編排執行什麼;平臺技能擁有在 GPU 主機上執行如何 — 請先閱讀它們。

關注點權威技能
GPU 主機執行時(驅動 580,CUDA Toolkit 13.0,NVIDIA Container Toolkit 1.19.0)`tao-skill-bank:tao-setup-nvidia-gpu-host`
`docker run` 標誌、NGC 認證、掛載、環境變數傳遞`tao-skill-bank:tao-run-on-docker`
本地 Docker 作業預檢(守護程序、GPU 冒煙測試)`tao-skill-bank:tao-run-on-local-docker`

預設平臺: local-docker — 構建一次性映象(run-<short>:latest</short>)並在本地 Docker 守護程序上執行它。僅當使用者明確需要不同後端(Brev 遠端 GPU、SLURM/Kubernetes)時才詢問;然後先執行該平臺的預檢,並透過它將第 4–5 步的 docker run 命令路由到其中。GPU 執行時和僅存在性憑據預檢(值從不讀取)、標準的 docker run 標誌集、list_tao_platforms.py 選擇命令以及工作流特定標誌(--entrypoint /bin/bash -lc、PYTORCH_CUDA_ALLOC_CONF、--name hft_train)位於 references/workflow-intake-preflight.md。

參考資料 — 備用安全網

僅在實時研究沉默、模稜兩可或不可用時查閱;實時文件對於特定模型和當前 API 始終優先。每個步驟連結其所需的參考資料;完整目錄在 references/detailed-workflow.md 中。

始終開啟:core-rules.md、error-playbook.md、compat-workarounds.md、model-discovery.md、dataset-recommendations.md、dataset-sources.md、dataset-patterns.md、hardware-container.md、research-priorities.md、cv-scripts.md、vlm-scripts.md、docker-runs.md、hub-push.md、pipeline-skill-template.md、deliverables.md。可選(當其標誌/需求適用時):progress-tracking.md、testing.md、reporting.md、workflow-intake-preflight.md、workflow-generate-train.md、workflow-push-rerun.md。

規則: 在回退之前,記錄你嘗試過的實時源及其不足的原因(config.yaml 中的 notes:,如果啟用則記錄 PROGRESS.md)。cv-scripts.md / vlm-scripts.md 中的 [FETCH LIVE] 標記是研究清單,不是內聯程式碼 — 如果某個塊沒有第 3 步的發現,請重新獲取列出的 URL。

核心規則

不可協商的行為。簡短版本(完整列舉 — 幻覺匯入列表、未經批准絕不允許列表、完整錯誤恢復和硬體調整大小表 — 在 references/core-rules.md 中,在任何訓練時決策前諮詢):

  • 你的 HF 庫知識已過時。 在編寫任何 ML 程式碼之前獲取實時文件(模型卡片、HF 倉庫示例、任務文件) — 不要從記憶中生成訓練器引數 / 整理器 / 變換(第 3 步)。
  • 使用 --max_steps 1 在真實資料上進行冒煙測試 在任何完整執行之前;未經驗證的冒煙測試不得批次啟動。
  • 切勿靜默替換 model_id、dataset_id 或 training_method — 如果使用者請求的內容無法載入,停止並詢問。
  • 錯誤恢復是最小更改。 OOM → 減半批次,加倍 grad_accum,啟用梯度檢查點(未經批准不切換 LoRA);NaN → 將 LR 減少 10 倍;平坦損失 → 檢查整理器;相同錯誤 3 次 → 停止並詢問。不要迴圈。
  • 資料集列在整理器之前驗證 — 在 prepare_data.py 中重新命名;需要重構 → 停止並詢問。
  • 硬體調整大小經驗法則(bf16): ≤3B → 24 GB,7–13B → 80 GB,30B+ → 多 GPU 或 1× 80 GB 上的 LoRA,70B+ → 8× 80 GB 或 LoRA。完整微調無法容納且未請求 LoRA → 切換前詢問。

工作流 — 6 個步驟

單次傳遞,順序執行;每個步驟在下一個開始之前都有明確的門控。

第 1 步 — 檢查與資格認定

目標: 決定是否繼續。探測模型 + 資料集,應用接受/拒絕,註冊適用的相容修復,編寫初始 config.yaml。

先決條件:MODEL_ID、可選 DATASET_ID / local_dataset_path、可選 HF_TOKEN、OUTPUT_DIR(預設 ./output/<model_short_name></model_short_name>)。探測在僅 CPU 的 python:3.12-slim Docker 容器中執行(繫結掛載 .probe/scratch),因此主機不需要虛擬環境 — Docker 必須首先存在。Docker 存在性保護、容器環境、完整探測呼叫以及模型/資料集探測指令碼位於 references/workflow-intake-preflight.md、references/model-discovery.md 和 references/dataset-sources.md。

探測要求:

  • 模型:載入 AutoConfig,讀取模型卡片標籤,從 architectures + 標籤 + 卡片示例中檢測任務(回退日誌記錄在 model-discovery.md 中)。
  • 資料集:對於推薦的資料集,首先從 dataset-recommendations.md 展示 3-5 個選擇;對於本地資料,只讀繫結掛載並使用 dataset-sources.md 格式檢測。
  • 如果模型配置失敗、任務超出範圍、沒有配方源或資料集無法載入/匹配任務模式,則提前拒絕。
  • 根據模型/任務評估 compat-workarounds.md;將硬體相關規則推遲到第 2 步。

編寫初始 config.yaml(model_id、task、dataset_id 或 local_dataset_path、research_sources: [] 在第 3 步填充、applicable_workarounds: 來自第 1 步、notes: [] 用於參考資料回退、push_to_hub: true 預設值 — 註釋模板在 references/workflow-intake-preflight.md 中)。門控滿足後可選 rm -rf "$OUTPUT_DIR/.probe"。

門控: config.yaml 存在,包含模型、資料集、任務、applicable_workarounds;如果任何欄位缺失,不得繼續。

第 2 步 — 硬體審計與 NGC 映象

目標: 驗證 Docker + GPU + 磁碟,實時選擇 NGC PyTorch 映象,最終確定硬體相關的相容規則。

2a. 審計(硬門控) — 三個檢查(命令在 references/workflow-intake-preflight.md 中):

  1. GPU 主機執行時 — tao-setup-nvidia-gpu-host 的 setup-nvidia-gpu-host.sh --backend docker --check-only;如果失敗,請求批准然後使用 --install --yes 重新執行。
  2. 空閒磁碟軟警告 — 透過 MIN_DISK_GB 覆蓋(預設 100 GB);建議 ≥ 100 GB 用於 NGC 基礎(~20 GB)+ HF 快取 + 檢查點 + 資料。
  3. 條件憑據存在性(來自會話環境,值從不讀取) — HF_TOKEN 僅在受限或 push_to_hub 開啟時;WANDB_* 僅在 WandB 開啟時。

硬失敗時不得進入第 4 步 — 第 4 步的 docker build 拉取 20+ GB NGC 基礎,而缺失的 nvidia-container-toolkit 稍後才會顯示為 could not select device driver "" with capabilities: [[gpu]]。在 config.yaml 中記錄 gpu_count、gpu_name、driver_major、vram_gb_per_gpu。

2b. 選擇 NGC 映象(實時): 從 NVIDIA 深度學習框架支援矩陣(https://docs.nvidia.com/deeplearning/frameworks/support-matrix/index.html),PyTorch NGC 容器部分,選擇版本號最高的映象,其中 Min driver ≤ 檢測到的 driver_major 且容器 CUDA ≤ 主機 CUDA Toolkit(儘可能匹配以便 cuDNN / TensorRT 對齊)。不要因 aN/bN/rcN PyTorch 標籤而拒絕映象 — NGC 驗證整個映象;選擇最新的 CUDA 對齊版本,讓 compat-workarounds.md 處理每個版本的問題。如果矩陣不可達,使用 references/hardware-container.md 中的回退;預設 nvcr.io/nvidia/pytorch:24.09-py3(驅動 ≥ 545;SDPA+GQA 錯誤 — 如果 num_key_value_heads , setattn_implementation: "eager")。在config.yaml中記錄ngc_image`。

2c. 重新評估硬體相關的相容規則: 重新執行 compat-workarounds.md 遍歷,針對其 detect 需要 hw 的條目;就地更新 applicable_workarounds:。

2d. 模型適配檢查: 估計 param_bytes ≈ 2×param_count(bf16);如果

60% of vram_gb_per_gpu × 1e9,在面向使用者的摘要中推薦 LoRA。

門控: config.yaml 包含 ngc_image、gpu_count、gpu_name、driver_major、vram_gb_per_gpu;記錄了硬體相關的相容修復。

第 3 步 — 研究配方

目標: 獲取實時配方 — transformers/trl/peft 的訓練資料知識存疑,因此第 3 步是不可協商的。按優先順序順序遍歷 references/research-priorities.md(優先順序 1 → 6);一旦你為檢測到的任務獲得以下內容,即停止:

  • AutoModel / 處理器類
  • 訓練 + 評估變換
  • 整理器
  • compute_metrics
  • 超引數提示(LR、批次大小、epochs、排程器)

將發現記錄在 meta/recipe.md 中,將源 URL 追加到 config.yaml: research_sources:。沒有實時發現的槽位回退到匹配的腳手架(cv-scripts.md / vlm-scripts.md),在 notes: 下記錄為“回退到腳手架 — 沒有實時源用於”。衝突解決規則在 references/research-priorities.md 中。

門控: 每個必需的槽位已填充,帶有源 URL 或腳手架回退說明。

第 4 步 — 生成專案與冒煙測試

目標: 編寫所有指令碼,構建映象,準備資料,在真實資料上執行 1 步冒煙測試(一個 docker build,兩個 docker run)。

4a. 在專案檔案 output_dir/ 中生成:config.yaml、Dockerfile、requirements.txt、prepare_data.py、train.py、run_eval.py、infer.py、可選 merge_lora.py、可選 tests/、.gitignore。實時第 3 步研究是權威;cv-scripts.md / vlm-scripts.md 僅提供腳手架形狀。將每個 applicable_workarounds 條目應用為 Dockerfile 塊、要求固定、配置覆蓋或執行時環境變數。硬規則:run_eval.py 保持確切檔名(避免與 HF evaluate 包衝突);每個生成的 .py 以 NVIDIA Apache-2.0 版權標頭開頭,任何發射器在缺少時失敗;emit_unit_tests: true 根據 references/testing.md 生成並執行測試。指令碼主體、Dockerfile 形狀和發射器合同在 references/workflow-generate-train.md 中。

4b. 構建、準備、冒煙 — docker build -t run-<short>:latest .</short>,然後 prepare_data 和 --smoke --max_steps 1 執行(references/docker-runs.md§1-3)。冒煙透過標準(在 logs/smoke.log 中):

  • 無異常
  • 損失是有限的(不是 0.0,不是 NaN)
  • 步驟 1 時 grad_norm > 0

如果 emit_unit_tests: true,還在容器中執行 pytest tests/。任何失敗 → 停止。

4c. 預檢摘要 — 在完整訓練之前,列印並驗證:參考 URL、資料集列、Hub 目標、監控目標、NGC 映象、硬體、冒煙損失/梯度範數。

門控: 專案檔案已寫入,映象已構建,冒煙透過,預檢沒有空白欄位。

第 5 步 — 訓練、評估、推理

目標: 基線評估、完整訓練、訓練後評估、可選 LoRA 合併、5 個推理樣本(所有命令:references/docker-runs.md §4-8)。

子步驟docker-runs.md跳過條件
5a. 基線評估(零樣本)§4`skip_baseline: true`
5b. 完整訓練(分離)§5—
5c. LoRA 合併§6非 VLM+LoRA
5d. 訓練後評估§7—
5e. 推理(5 個樣本)§8—

多 GPU:在 python train.py 前新增 torchrun --nproc_per_node=$gpu_count。

訓練流式傳輸時,觀察 docker logs -f hft_train:損失應在 10-20 步內下降;平坦損失(整理器/標籤掩碼錯誤)、NaN(LR 過高)和 OOM 都會停止執行 — 恢復在 references/core-rules.md 中。如果 emit_report: true,在第 5e 步後根據 references/reporting.md 執行 report.py。

門控: 全部:

  • checkpoints/final/(或 LoRA 的 checkpoints/merged/)存在
  • reports/eval_results.json 具有數值主指標
  • reports/baseline_results.json 存在(除非跳過)
  • reports/inference_samples/ 有 5 個樣本
  • wandb URL 顯示下降的損失

第 6 步 — 推送與發射重跑技能

目標: 釋出執行並使其可重現,無需重新研究。

根據 references/hub-push.md 推送(權重、模型卡片、評估/基線 JSON、config.yaml、Dockerfile、requirements.txt、推理樣本、發射的報告),除非明確 push_to_hub: false。從 references/pipeline-skill-template.md 發射 <output_dir>/skills/run-<short>/SKILL.md</short></output_dir> — 替換每個佔位符,包含完整的 YAML 後設資料 + NVIDIA 版權 HTML 註釋,如果缺少這些內容則使任何發射器失敗。

門控(完成標準): 全部:

  • 第 5 步門控滿足
  • HF Hub 倉庫存在於解析的 URL,包含權重 + 卡片 + results/(除非 push_to_hub: false)
  • <output_dir>/skills/run-<short>/SKILL.md</short></output_dir> 存在,沒有 <placeholder></placeholder> 遺留,帶有 pipeline-skill-template.md 規定的後設資料 + 版權 HTML 註釋

最終訊息:wandb URL、HF Hub URL、基線 -> 微調主指標、reports/inference_samples/ 以及重跑技能路徑。

錯誤劇本

在已知的執行時錯誤上,諮詢 references/error-playbook.md 中的症狀 → 最小修復表(NGC 入口點、PyTorch/Transformers 迴歸、numpy ABI、Albumentations bbox、PEFT/檢查點、LoRA 目標廣度、CV 增強差距、步驟 0 時的 OOM)在重新設計任何內容之前。當某一行在跨執行中觸發兩次時,將其提升為 compat-workaroutines.md 中的 detect 規則 — 在第 1 步中自動應用,在錯誤觸發之前。

溝通風格

  • 簡潔。無填充詞,無重述請求;適當時使用單字回答。
  • 引用工件時始終包含直接的 Hub 和 wandb URL。
  • 出錯時:說明出了什麼問題、原因、你做了什麼更改 — 無選單。
  • 對於有明確答案的請求,切勿呈現“選項 A/B/C”。行動。

示例管道

  • tao-rerun-convnext-cifar10
  • tao-rerun-detr-cppe5
  • tao-rerun-segformer-foodseg103
  • tao-rerun-smolvlm-vqav2
在 GitHub 上查看
---
name: tao-finetune-huggingface-model
description: Fine-tune HuggingFace CV, VLM, or LLM models on local NVIDIA GPUs using an NGC PyTorch container, with support for full or LoRA training, dataset handling, and optional model push to the Hub.
license: Apache-2.0
---
<!-- Copyright (c) 2026, NVIDIA CORPORATION. All rights reserved. Licensed under the Apache License, Version 2.0; see http://www.apache.org/licenses/LICENSE-2.0 -->

# tao-finetune-huggingface-model

Local NVIDIA GPU fine-tuning for HuggingFace models, grounded in live-fetched
documentation with curated references as a fallback safety net. One NGC container,
a few focused scripts, one push to HF Hub. Follow the rules in this file; don't
improvise.

**Order of authority (highest first):**

1. **User input** — explicit `model_id`, `dataset_id`, `training_method`, `config.yaml` overrides.
2. **Live research** — model card, HF repo example, author finetune script, HF task docs, paper; always fetched (Step 3 + `references/research-priorities.md`).
3. **Curated references** (`references/*.md`) — fallback when live research is silent/ambiguous.
4. **Your training-data memory** — last resort; suspect, cross-check against (2)/(3).

Conflict resolution between (2) and (3) and the source-line discrepancy note are
in `references/research-priorities.md`.

---

## Inputs

**Required:**
- `model_id` — HuggingFace model ID, e.g. `google/vit-base-patch16-224`

**Conditional credentials (read from the session environment, exported before launching when present):**
- `HF_TOKEN` — only when the model/dataset is **gated** (read) or `push_to_hub` is on (write); public + public + `push_to_hub: false` needs none. Value never read — presence-only via `[ -n "$HF_TOKEN" ]`.
- `WANDB_API_KEY`, `WANDB_PROJECT` — only when WandB is enabled; `WANDB_MODE=disabled` opts out.

**Dataset — exactly one:**
- `dataset_id` — HuggingFace dataset ID *(source: `hf`)*
- `local_dataset_path` — local folder or file *(source: `local`)*; optional
  `local_dataset_format` ∈ {auto, imagefolder, coco, voc, jsonl, arrow, parquet,
  csv} (default: auto-detect).
- *(omit)* — agent recommends popular datasets *(source: `recommend`)*

**Optional (have defaults):**
- `task_type` — auto-detected from config + model card
- `n_train=10000`, `n_eval=1000`, `n_epochs=3`, `lora_r=16`
- `output_dir=./output/<model_short_name>`
- `hf_model_repo` — push target; if unset and HF_TOKEN has write access,
  auto-derived as `<whoami>/<model_short_name>-finetuned`.
- `push_to_hub=True` — set to `False` to skip
- `skip_baseline=False` — skip zero-shot baseline eval

**Optional deliverables (off by default):**
```yaml
emit_progress_log: false   # output_dir/PROGRESS.md (per-step journal)
emit_report:       false   # reports/report.{pdf,html} with curves & samples
emit_unit_tests:   false   # tests/ with fake-data heterogeneous-batch tests
```

All values live in `output_dir/config.yaml`. Never hardcode in Python.

---

## Execution platform

This skill orchestrates *what* to run; the platform skills own *how* to run it on
a GPU host — read them first.

| Concern | Authoritative skill |
|---|---|
| GPU host runtime (driver 580, CUDA Toolkit 13.0, NVIDIA Container Toolkit 1.19.0) | [`tao-skill-bank:tao-setup-nvidia-gpu-host`](../../platform/tao-setup-nvidia-gpu-host/SKILL.md) |
| `docker run` flags, NGC auth, mounts, env passthrough | [`tao-skill-bank:tao-run-on-docker`](../../platform/tao-run-on-docker/SKILL.md) |
| Local Docker job preflight (daemon, GPU smoke) | [`tao-skill-bank:tao-run-on-local-docker`](../../platform/tao-run-on-local-docker/SKILL.md) |

**Default platform:** `local-docker` — build a one-off image (`run-<short>:latest`)
and run it on the local Docker daemon. Ask only when the user explicitly needs a
different backend (Brev remote GPU, SLURM/Kubernetes); then run that platform's
Preflight first and route the Steps 4–5 `docker run` commands through it. The
GPU-runtime and presence-only credential preflights (values never read), the
canonical `docker run` flag set, the `list_tao_platforms.py` selection command, and
the workflow-specific flags (`--entrypoint /bin/bash -lc`, `PYTORCH_CUDA_ALLOC_CONF`,
`--name hft_train`) are in `references/workflow-intake-preflight.md`.

---

## References — fallback safety net

Consulted **only** when live research is silent, ambiguous, or unavailable; live
docs always win for the specific model and current API. Each step links the
references it needs; full catalog in `references/detailed-workflow.md`.

Always-on: `core-rules.md`, `error-playbook.md`, `compat-workarounds.md`,
`model-discovery.md`, `dataset-recommendations.md`, `dataset-sources.md`,
`dataset-patterns.md`, `hardware-container.md`, `research-priorities.md`,
`cv-scripts.md`, `vlm-scripts.md`, `docker-runs.md`, `hub-push.md`,
`pipeline-skill-template.md`, `deliverables.md`. Opt-in (when their flag/need
applies): `progress-tracking.md`, `testing.md`, `reporting.md`,
`workflow-intake-preflight.md`, `workflow-generate-train.md`, `workflow-push-rerun.md`.

**Rule:** before falling back, log the live source you tried and why it was
insufficient (`config.yaml` `notes:`, and PROGRESS.md if enabled). `[FETCH LIVE]`
markers in `cv-scripts.md` / `vlm-scripts.md` are a research checklist, not code to
inline — refetch the listed URL if a block has no Step 3 finding.

---

## Core rules

Non-negotiable behaviors. **Short version** (full enumeration —
hallucinated-imports list, never-without-approval list, full error-recovery and
hardware-sizing tables — in `references/core-rules.md`, consult before any
training-time decision):

- **Your HF-library knowledge is outdated.** Fetch live docs (model card, HF
  repo example, task doc) before writing any ML code — don't generate trainer
  args / collator / transforms from memory (Step 3).
- **Smoke-test on real data with `--max_steps 1`** before any full run; no batch
  launches without a verified smoke.
- **Never silently substitute** model_id, dataset_id, or training_method — if
  what the user asked for doesn't load, stop and ask.
- **Error recovery is minimal-change.** OOM → halve batch, double grad_accum,
  enable gradient checkpointing (no LoRA switch without approval); NaN → reduce
  LR 10×; flat loss → inspect collator; same error 3× → stop and ask. Don't loop.
- **Dataset columns verified BEFORE the collator** — rename in `prepare_data.py`;
  restructuring needed → stop and ask.
- **Hardware-sizing thumb (bf16):** ≤3B → 24 GB, 7–13B → 80 GB, 30B+ → multi-GPU
  or LoRA on 1× 80 GB, 70B+ → 8× 80 GB or LoRA. Full finetune won't fit and no
  LoRA requested → ask before switching.

---

## Workflow — 6 steps

Single pass, sequential; each step has a clear gate before the next begins.

### Step 1 — Inspect & qualify

**Goal:** decide whether to proceed. Probe model + dataset, apply accept/reject,
register applicable compat fixes, write the initial `config.yaml`.

Prerequisites: `MODEL_ID`, optional `DATASET_ID` / `local_dataset_path`,
optional `HF_TOKEN`, `OUTPUT_DIR` (default `./output/<model_short_name>`). Probes
run in a CPU-only `python:3.12-slim` Docker container (bind-mounted `.probe/`
scratch) so the host needs no virtualenv — Docker must exist first. Docker-presence
guard, container env, full probe invocation, and the model/dataset probe scripts
are in `references/workflow-intake-preflight.md`, `references/model-discovery.md`,
and `references/dataset-sources.md`.

Probe requirements:

- Model: load `AutoConfig`, read model-card tags, detect task from
  `architectures` + tags + card examples (fallback logging in `model-discovery.md`).
- Dataset: for recommended datasets, first present 3-5 choices from
  `dataset-recommendations.md`; for local data, bind-mount read-only and use
  `dataset-sources.md` format detection.
- Reject early if the model config fails, the task is out of scope, no recipe
  source exists, or the dataset cannot load / match the task schema.
- Evaluate `compat-workarounds.md` against the model/task; defer hardware-dependent
  rules to Step 2.

Write the initial `config.yaml` (`model_id`, `task`, `dataset_id` or
`local_dataset_path`, `research_sources: []` filled in Step 3,
`applicable_workarounds:` from Step 1, `notes: []` for reference fallbacks,
`push_to_hub: true` default — annotated template in
`references/workflow-intake-preflight.md`). Optionally `rm -rf "$OUTPUT_DIR/.probe"`
once the gate is met.

**Gate:** `config.yaml` exists with model, dataset, task, applicable_workarounds;
do not proceed if any field is missing.

---

### Step 2 — Hardware audit & NGC image

**Goal:** verify Docker + GPU + disk, pick the NGC PyTorch image live, finalize
hardware-dependent compat rules.

**2a. Audit (hard gate)** — three checks (commands in
`references/workflow-intake-preflight.md`):
1. GPU host runtime — `tao-setup-nvidia-gpu-host`'s
   `setup-nvidia-gpu-host.sh --backend docker --check-only`; on fail, ask approval
   then re-run with `--install --yes`.
2. Free-disk soft-warn — override via `MIN_DISK_GB` (default 100 GB); recommend
   ≥ 100 GB for NGC base (~20 GB) + HF cache + checkpoints + data.
3. Conditional credential presence (from the session environment, values never
   read) — `HF_TOKEN` only when gated or `push_to_hub` is on; `WANDB_*` only when
   WandB is on.

**Do not proceed to Step 4 on a hard-fail** — Step 4's `docker build` pulls a
20+ GB NGC base, and a missing `nvidia-container-toolkit` only surfaces later as
`could not select device driver "" with capabilities: [[gpu]]`. Record `gpu_count`,
`gpu_name`, `driver_major`, `vram_gb_per_gpu` in `config.yaml`.

**2b. Pick NGC image (live):** from the NVIDIA Deep Learning Frameworks support
matrix (<https://docs.nvidia.com/deeplearning/frameworks/support-matrix/index.html>),
PyTorch NGC container section, pick the highest-versioned image where
`Min driver ≤ detected driver_major` and container CUDA `≤` host CUDA Toolkit
(match closely so cuDNN / TensorRT line up). Do **not** reject an image for an
`aN`/`bN`/`rcN` PyTorch tag — NGC validates the full image; pick the newest
CUDA-aligned one and let `compat-workarounds.md` handle per-version issues. If the
matrix is unreachable, use the fallbacks in `references/hardware-container.md`;
default `nvcr.io/nvidia/pytorch:24.09-py3` (driver ≥ 545; SDPA+GQA bug — if
`num_key_value_heads < num_attention_heads`, set `attn_implementation: "eager"`).
Record `ngc_image` in `config.yaml`.

**2c. Re-evaluate hardware-dependent compat rules:** re-run the
`compat-workarounds.md` walk for entries whose `detect` needs `hw`; update
`applicable_workarounds:` in place.

**2d. Model-fit check:** estimate `param_bytes ≈ 2×param_count` (bf16); if
> 60% of `vram_gb_per_gpu × 1e9`, recommend LoRA in the user-facing summary.

**Gate:** `config.yaml` has `ngc_image`, `gpu_count`, `gpu_name`, `driver_major`,
`vram_gb_per_gpu`; hardware-dependent compat fixes recorded.

---

### Step 3 — Research the recipe

**Goal:** fetch the live recipe — training-data knowledge of
`transformers`/`trl`/`peft` is suspect, so Step 3 is non-negotiable. Walk
`references/research-priorities.md` in priority order (Priority 1 → 6); stop once
you have, for the detected task:

- `AutoModel` / processor class
- Train + eval transforms
- Collator
- `compute_metrics`
- Hyperparameter hints (LR, batch size, epochs, scheduler)

Record findings in `meta/recipe.md`, append source URLs to
`config.yaml: research_sources:`. A slot with no live finding falls back to the
matching scaffold (`cv-scripts.md` / `vlm-scripts.md`), logged as "fallback to
scaffold — no live source for <slot>" under `notes:`. Conflict-resolution rules
are in `references/research-priorities.md`.

**Gate:** every required slot filled, with a source URL or scaffold-fallback note.

---

### Step 4 — Generate project & smoke-test

**Goal:** write all scripts, build the image, prepare data, run a 1-step smoke on
real data (one `docker build`, two `docker run`s).

**4a. Generate project files** in `output_dir/`: `config.yaml`, `Dockerfile`,
`requirements.txt`, `prepare_data.py`, `train.py`, `run_eval.py`, `infer.py`,
optional `merge_lora.py`, optional `tests/`, `.gitignore`. Live Step 3 research is
authority; `cv-scripts.md` / `vlm-scripts.md` give scaffold shape only. Apply every
`applicable_workarounds` entry as a Dockerfile block, requirement pin, config
override, or runtime env var. Hard rules: `run_eval.py` keeps that exact filename
(avoids colliding with the HF `evaluate` package); every generated `.py` starts
with the NVIDIA Apache-2.0 copyright header and any emitter fails when it is
missing; `emit_unit_tests: true` generates and runs tests per
`references/testing.md`. Script bodies, Dockerfile shape, and the emitter contract
are in `references/workflow-generate-train.md`.

**4b. Build, prepare, smoke** — `docker build -t run-<short>:latest .`, then
`prepare_data` and the `--smoke --max_steps 1` run (`references/docker-runs.md`
§1-3). Smoke pass criteria (in `logs/smoke.log`):
- No exception
- Loss is finite (not `0.0`, not `NaN`)
- `grad_norm > 0` at step 1

If `emit_unit_tests: true`, also run `pytest tests/` in the container. Any failure → STOP.

**4c. Preflight summary** — before full training, print and verify: reference URL,
dataset columns, Hub target, monitoring target, NGC image, hardware, smoke loss/grad norm.

**Gate:** project files written, image built, smoke PASSED, preflight has no
blank fields.

---

### Step 5 — Train, evaluate, infer

**Goal:** baseline eval, full training, post-train eval, optional LoRA merge, 5
inference samples (all commands: `references/docker-runs.md` §4-8).

| Sub-step | docker-runs.md | Skip if |
|---|---|---|
| 5a. Baseline eval (zero-shot) | §4 | `skip_baseline: true` |
| 5b. Full training (detached) | §5 | — |
| 5c. LoRA merge | §6 | not VLM+LoRA |
| 5d. Post-train eval | §7 | — |
| 5e. Inference (5 samples) | §8 | — |

Multi-GPU: prepend `torchrun --nproc_per_node=$gpu_count` to `python train.py`.

While training streams, watch `docker logs -f hft_train`: loss should drop within
10-20 steps; flat loss (collator/label-masking bug), NaN (LR too high), and OOM
all stop the run — recovery in `references/core-rules.md`. If `emit_report: true`,
run `report.py` after Step 5e per `references/reporting.md`.

**Gate:** all of:
- `checkpoints/final/` (or `checkpoints/merged/` for LoRA) exists
- `reports/eval_results.json` has a numeric primary metric
- `reports/baseline_results.json` exists (unless skipped)
- `reports/inference_samples/` has 5 samples
- wandb URL shows descending loss

---

### Step 6 — Push & emit rerun skill

**Goal:** publish the run and make it reproducible without re-research.

Push per `references/hub-push.md` (weights, model card, eval/baseline JSONs,
`config.yaml`, `Dockerfile`, `requirements.txt`, inference samples, reports when
emitted) unless `push_to_hub: false` is explicit. Emit
`<output_dir>/skills/run-<short>/SKILL.md` from
`references/pipeline-skill-template.md` — substitute every placeholder, include
full YAML metadata + the NVIDIA copyright HTML comment, and make any emitter fail
if those are missing.

**Gate (Done criteria):** all of:
- Step 5 gate met
- HF Hub repo exists at the resolved URL with weights + card + `results/`
  (unless `push_to_hub: false`)
- `<output_dir>/skills/run-<short>/SKILL.md` exists, no `<placeholder>` left,
  with metadata + copyright HTML comment per `pipeline-skill-template.md`

Final message: wandb URL, HF Hub URL, baseline -> fine-tuned primary metric,
`reports/inference_samples/`, and the rerun skill path.

---

## Error playbook

On a known runtime error, consult the symptom → minimal-fix table in
`references/error-playbook.md` (NGC entrypoint, PyTorch/Transformers regressions,
numpy ABI, Albumentations bbox, PEFT/checkpointing, LoRA target breadth, CV
augmentation gaps, OOM at step 0) before redesigning anything. When a row there
fires twice across runs, lift it into `compat-workarounds.md` with a `detect` rule
— auto-applied in Step 1 before the error can fire.

---

## Communication style

- Terse. No filler, no restating the request; one-word answers when appropriate.
- Always include direct Hub and wandb URLs when referencing artifacts.
- On error: state what went wrong, why, what you changed — no menus.
- Never present "Option A/B/C" for a request with a clear answer. Act.

## Example pipelines

- [tao-rerun-convnext-cifar10](references/tao-rerun-convnext-cifar10.md)
- [tao-rerun-detr-cppe5](references/tao-rerun-detr-cppe5.md)
- [tao-rerun-segformer-foodseg103](references/tao-rerun-segformer-foodseg103.md)
- [tao-rerun-smolvlm-vqav2](references/tao-rerun-smolvlm-vqav2.md)

所有檔案

69 個檔案

安裝 tao-finetune-huggingface-model

將技能檔案下載並解壓至 .claude/skills/ 目錄。

下載 ZIP

複製儲存庫並將技能檔案複製到您的專案中。

git clone https://github.com/NVIDIA/skills/tree/main/skills/tao-finetune-huggingface-model # Copy SKILL.md to your .claude/skills/ directory

複製 複製
快速設定: 將技能資料夾複製到 .claude/skills/ 目錄。Claude 將自動檢測並使用該技能。
儲存庫 NVIDIA/skills

相關技能

web-search
更新時間 2026-06-29
webapp-testing
更新時間 2026-06-29
lark-base
更新時間 2026-07-05
agentmail
更新時間 2026-06-29
OR