オプション
家家 Skill データサイエンスと機械学習 tao-finetune-huggingface-model

tao-finetune-huggingface-model

NVIDIA/skills NVIDIA/skills

NGC PyTorch コンテナを使用して、ローカルの NVIDIA GPU 上で HuggingFace の CV、VLM、または LLM モデルをファインチューニングします。フルトレーニングまたは LoRA トレーニング、データセットの処理、およびオプションでモデルを Hub へのプッシュをサポートしています。

...すべて拡張します
1
更新された時間 2026年9月29日

tao-finetune-huggingface-model

HuggingFaceモデルのローカルNVIDIA GPUでのファインチューニング。ライブで取得したドキュメントを基盤とし、キュレーションされた参照を安全網として使用します。1つのNGCコンテナ、数本の集中スクリプト、HF Hubへの1回のプッシュ。このファイルのルールに従ってください。独自の推測を行わないでください。

権限の順序(最高優先度から):

  1. ユーザー入力 — 明示的な model_id、dataset_id、training_method、config.yaml の上書き。
  2. ライブリサーチ — モデルカード、HFリポジトリの例、著者のファインチューンスクリプト、HFタスクドキュメント、論文;常にフェッチされる(ステップ3 + references/research-priorities.md)。
  3. キュレーションされた参照 (references/*.md) — ライブリサーチが沈黙している場合や曖昧な場合のフォールバック。
  4. あなたのトレーニングデータのメモリ — 最後の手段;疑わしいため、(2)/(3) に対してクロスチェックを行う。

(2) と (3) の間の競合解決およびソース行の不一致に関する注記は references/research-priorities.md にあります。

入力

必須:

  • model_id — HuggingFaceモデルID、例:google/vit-base-patch16-224

条件付き認証情報(セッション環境から読み取り、起動前にエクスポートされる場合):

  • HF_TOKEN — モデル/データセットがゲートされている(読み取り)場合、または push_to_hub がオンの場合(書き取り)のみ;公開 + 公開 + push_to_hub: false は不要。値は読み取られない — [ -n "$HF_TOKEN" ] による存在のみ。
  • WANDB_API_KEY、WANDB_PROJECT — WandBが有効な場合のみ;WANDB_MODE=disabled でオプトアウト可能。

データセット — 正確に1つ:

  • dataset_id — HuggingFaceデータセットID (ソース: hf)
  • local_dataset_path — ローカルフォルダまたはファイル (ソース: local);オプションの local_dataset_format ∈ {auto, imagefolder, coco, voc, jsonl, arrow, parquet, csv}(デフォルト:自動検出)。
  • (省略) — エージェントが人気データセットを推奨 (ソース: recommend)

オプション(デフォルト値あり):

  • task_type — 設定 + モデルカードから自動検出
  • n_train=10000、n_eval=1000、n_epochs=3、lora_r=16
  • output_dir=./output/<model_short_name></model_short_name>
  • hf_model_repo — プッシュターゲット;未設定かつ HF_TOKEN に書き取りアクセスがある場合、<whoami>/<model_short_name>-finetuned</model_short_name></whoami> として自動導出。
  • push_to_hub=True — スキップするには False に設定
  • skip_baseline=False — ゼロショットベースライン評価をスキップ

オプションの成果物(デフォルトでオフ):

emit_progress_log: false   # output_dir/PROGRESS.md(ステップごとのジャーナル)
emit_report:       false   # 曲線とサンプルを含む reports/report.{pdf,html}
emit_unit_tests:   false   # フェイクデータを使用したヘテロジニアスバッチテストを含む tests/

すべての値は output_dir/config.yaml にあります。Python内でハードコードしないでください。

実行プラットフォーム

このスキルは 何を 実行するかを調整します;プラットフォームスキルは GPUホスト上で どのように 実行するかを所有します — 最初にそれらを読み取ってください。

懸念事項権威あるスキル
GPUホストランタイム(ドライバー580、CUDA Toolkit 13.0、NVIDIA Container Toolkit 1.19.0)`tao-skill-bank:tao-setup-nvidia-gpu-host`
`docker run` フラグ、NGC認証、マウント、環境変数の渡しま`tao-skill-bank:tao-run-on-docker`
ローカルDockerジョブの事前チェック(デーモン、GPUスモークテスト)`tao-skill-bank:tao-run-on-local-docker`

デフォルトプラットフォーム: local-docker — 単発イメージ(run-<short>:latest</short>)をビルドし、ローカルDockerデーモンで実行します。ユーザーが明示的に異なるバックエンド(BrevリモートGPU、SLURM/Kubernetes)を必要とする場合にのみ問い合わせてください;その場合、そのプラットフォームの事前チェックを最初に実行し、ステップ4–5の docker run コマンドをそれを通じてルーティングします。GPUランタイムおよび存在のみ認証情報の事前チェック(値は読み取られない)、標準的な docker run フラグセット、list_tao_platforms.py 選択コマンド、およびワークフロー固有のフラグ(--entrypoint /bin/bash -lc、PYTORCH_CUDA_ALLOC_CONF、--name hft_train)は references/workflow-intake-preflight.md にあります。

参照 — フォールバック安全網

ライブリサーチが沈黙している場合、曖昧な場合、または利用できない場合にのみ参照されます;ライブドキュメントは特定のモデルおよび現在のAPIに対して常に優先されます。各ステップは必要な参照へのリンクを提供します;完全なカタログは references/detailed-workflow.md にあります。

常時有効:core-rules.md、error-playbook.md、compat-workarounds.md、model-discovery.md、dataset-recommendations.md、dataset-sources.md、dataset-patterns.md、hardware-container.md、research-priorities.md、cv-scripts.md、vlm-scripts.md、docker-runs.md、hub-push.md、pipeline-skill-template.md、deliverables.md。オプトイン(フラグ/必要性が適用される場合):progress-tracking.md、testing.md、reporting.md、workflow-intake-preflight.md、workflow-generate-train.md、workflow-push-rerun.md。

ルール: フォールバックする前に、試したライブソースおよびそれが不十分だった理由をログに記録する(config.yaml の notes:、および有効な場合は PROGRESS.md)。cv-scripts.md / vlm-scripts.md 内の [FETCH LIVE] マーカーはインライン化するコードではなく、リサーチチェックリストです — ブロックにステップ3の発見がない場合、リストされたURLを再フェッチしてください。

コアルール

譲れない動作。短縮版(完全な列挙 — 幻覚インポートリスト、承認なしで決して使用しないリスト、完全なエラー回復およびハードウェアサイズ表 — は references/core-rules.md にあります、トレーニング時の意思決定の前に参照してください):

  • あなたのHFライブラリの知識は古くなっています。 任意のMLコードを書く前にライブドキュメント(モデルカード、HFリポジトリの例、タスクドキュメント)をフェッチしてください — メモリからトレーナー引数 / コレクター / 変換を生成しないでください(ステップ3)。
  • 完全な実行の前に --max_steps 1 で実際のデータ上でスモークテストを実行する — 検証されたスモークなしでバッチ起動を行わない。
  • モデル_id、データセット_id、またはトレーニング_methodを黙って置換しない — ユーザーが要求したものがロードできない場合、停止して問い合わせる。
  • エラー回復は最小限の変更。 OOM → バッチを半減、grad_accumを2倍、勾配チェックポイント有効化(承認なしでLoRA切り替え不可);NaN → LRを10倍減少;平坦な損失 → コレクターを検査;同じエラーが3回発生 → 停止して問い合わせる。ループしない。
  • データセットの列はコレクターの前に検証する — prepare_data.py で名前を変更;再構築が必要 → 停止して問い合わせる。
  • ハードウェアサイズ表(bf16): ≤3B → 24 GB、7–13B → 80 GB、30B+ → 1× 80 GBでのマルチGPUまたはLoRA、70B+ → 8× 80 GBまたはLoRA。完全なファインチューニングが収まらないかつLoRAが要求されていない場合 → 切り替える前に問い合わせる。

ワークフロー — 6ステップ

単一パス、順序通り;各ステップには次の開始前に明確なゲートがあります。

ステップ1 — 検査および資格付け

目標: 続行するかどうかを決定する。モデル + データセットをプローブし、受け入れ/拒否を適用し、適用可能な互換性修正を登録し、初期 config.yaml を書き込む。

前提条件:MODEL_ID、オプションの DATASET_ID / local_dataset_path、オプションの HF_TOKEN、OUTPUT_DIR(デフォルト ./output/<model_short_name></model_short_name>)。プローブはCPUのみの python:3.12-slim Dockerコンテナ(バインドマウントされた .probe/ スクラッチ)で実行されるため、ホストにはvirtualenvが必要ありません — Dockerが存在する必要があります。Docker存在ガード、コンテナ環境、完全なプローブ呼び出し、およびモデル/データセットプローブスクリプトは references/workflow-intake-preflight.md、references/model-discovery.md、および references/dataset-sources.md にあります。

プローブ要件:

  • モデル:AutoConfig をロードし、モデルカードタグを読み取り、architectures + タグ + カード例からタスクを検出する(フォールバックログは model-discovery.md)。
  • データセット:推奨データセットの場合、まず dataset-recommendations.md から3-5つの選択肢を表示;ローカルデータの場合、バインドマウントして読み取り専用とし、dataset-sources.md のフォーマット検出を使用する。
  • モデル設定が失敗した場合、タスクが範囲外の場合、レシピソースが存在しない場合、またはデータセットがロードできない / タスクスキーマと一致しない場合は早期に拒否する。
  • モデル/タスクに対して compat-workarounds.md を評価する;ハードウェア依存ルールをステップ2に延期する。

初期 config.yaml を書き込む(model_id、task、dataset_id または local_dataset_path、ステップ3で埋められた research_sources: []、ステップ1からの applicable_workarounds:、参照フォールバック用の notes: []、デフォルトの push_to_hub: true — アノテーション付きテンプレートは references/workflow-intake-preflight.md)。ゲートが満たされたらオプションで rm -rf "$OUTPUT_DIR/.probe" を実行。

ゲート: モデル、データセット、タスク、applicable_workarounds を含む config.yaml が存在する;いずれかのフィールドが欠落している場合は続行しない。

ステップ2 — ハードウェア監査およびNGCイメージ

目標: Docker + GPU + ディスクを検証し、ライブでNGC PyTorchイメージを選択し、ハードウェア依存の互換性ルールを確定する。

2a. 監査(ハードゲート) — 3つのチェック(コマンドは references/workflow-intake-preflight.md にあります):

  1. GPUホストランタイム — tao-setup-nvidia-gpu-host の setup-nvidia-gpu-host.sh --backend docker --check-only;失敗した場合、承認を求め、次に --install --yes で再実行。
  2. 空きディスクソフト警告 — MIN_DISK_GB(デフォルト100 GB)で上書き;NGCベース(~20 GB)+ HFキャッシュ + チェックポイント + データ用に ≥ 100 GBを推奨。
  3. 条件付き認証情報の存在(セッション環境から、値は読み取られない)— HF_TOKEN はゲートされている場合または push_to_hub がオンの場合のみ;WANDB_* はWandBがオンの場合のみ。

ハード失敗の場合、ステップ4に進まないこと — ステップ4の docker build は20 GB以上のNGCベースをプルし、欠落した nvidia-container-toolkit は後で could not select device driver "" with capabilities: [[gpu]] として表面化する。config.yaml に gpu_count、gpu_name、driver_major、vram_gb_per_gpu を記録する。

2b. NGCイメージの選択(ライブ): NVIDIAディープラーニングフレームワークサポートマトリックス(https://docs.nvidia.com/deeplearning/frameworks/support-matrix/index.html)のPyTorch NGCコンテナセクションから、Min driver ≤ 検出された driver_major かつコンテナCUDA ≤ ホストCUDA Toolkit(cuDNN / TensorRTが一致するように密接に一致)で最高バージョンのイメージを選択する。aN/bN/rcN PyTorchタグのためにイメージを拒否しないこと — NGCは完全なイメージを検証する;最新のCUDAに一致するものを選択し、compat-workarounds.md がバージョンごとの問題を処理する。マトリックスに到達できない場合、references/hardware-container.md のフォールバックを使用する;デフォルト nvcr.io/nvidia/pytorch:24.09-py3(ドライバー ≥ 545;SDPA+GQAバグ — num_key_value_heads , setattn_implementation: "eager")。config.yamlにngc_image` を記録する。

2c. ハードウェア依存の互換性ルールの再評価: hw が必要なエントリに対して compat-workarounds.md のウォークを再実行する;applicable_workarounds: をその場で更新する。

2d. モデル適合チェック: param_bytes ≈ 2×param_count(bf16)を見積もる;もし

vram_gb_per_gpu × 1e9 の 60%、ユーザー向けサマリーでLoRAを推奨する。

ゲート: config.yaml に ngc_image、gpu_count、gpu_name、driver_major、vram_gb_per_gpu がある;ハードウェア依存の互換性修正が記録されている。

ステップ3 — レシピのリサーチ

目標: ライブレシピをフェッチする — transformers/trl/peft のトレーニングデータの知識は疑わしいため、ステップ3は譲れない。優先度順(優先度1 → 6)で references/research-priorities.md を歩く;検出されたタスクに対して、以下を取得したら停止する:

  • AutoModel / プロセッサクラス
  • トレーニング + 評価変換
  • コレクター
  • compute_metrics
  • ハイパーパラメータヒント(LR、バッチサイズ、エポック、スケジューラ)

meta/recipe.md に発見を記録し、ソースURLを config.yaml: research_sources: に追加する。ライブ発見がないスロットは、一致するスケフォールド(cv-scripts.md / vlm-scripts.md)にフォールバックし、notes: 下に "fallback to scaffold — no live source for " としてログに記録される。競合解決ルールは references/research-priorities.md にあります。

ゲート: 必要なスロットがすべて埋められ、ソースURLまたはスケフォールドフォールバックノートがある。

ステップ4 — プロジェクトの生成およびスモークテスト

目標: 全スクリプトを書き、イメージをビルドし、データを準備し、実際のデータで1ステップのスモークを実行する(1回の docker build、2回の docker run)。

4a. プロジェクトファイルの生成 output_dir/ 内:config.yaml、Dockerfile、requirements.txt、prepare_data.py、train.py、run_eval.py、infer.py、オプションの merge_lora.py、オプションの tests/、.gitignore。ライブのステップ3リサーチが権威;cv-scripts.md / vlm-scripts.md はスケフォールドの形状のみを提供する。すべての applicable_workarounds エントリをDockerfileブロック、要件ピン、設定上書き、またはランタイム環境変数として適用する。ハードルール:run_eval.py はその正確なファイル名を維持する(HF evaluate パッケージとの衝突を避ける);生成されたすべての .py はNVIDIA Apache-2.0著作権ヘッダーで始まり、欠落している場合、任意のエミッターは失敗する;emit_unit_tests: true は references/testing.md に従ってテストを生成および実行する。スクリプト本文、Dockerfileの形状、およびエミッター契約は references/workflow-generate-train.md にあります。

4b. ビルド、準備、スモーク — docker build -t run-<short>:latest .</short>、次に prepare_data および --smoke --max_steps 1 実行(references/docker-runs.md§1-3)。スモークパス基準(logs/smoke.log 内):

  • 例外なし
  • 損失が有限(0.0 ではない、NaN ではない)
  • ステップ1で grad_norm > 0

emit_unit_tests: true の場合、コンテナ内で pytest tests/ も実行する。いずれかの失敗 → 停止。

4c. 事前チェックサマリー — 完全なトレーニングの前に、印刷して検証する:参照URL、データセットの列、Hubターゲット、モニタリングターゲット、NGCイメージ、ハードウェア、スモーク損失/勾配ノルム。

ゲート: プロジェクトファイルが書かれ、イメージがビルドされ、スモークが合格し、事前チェックに空白のフィールドがない。

ステップ5 — トレーニング、評価、推論

目標: ベースライン評価、完全なトレーニング、トレーニング後の評価、オプションのLoRAマージ、5つの推論サンプル(全コマンド:references/docker-runs.md §4-8)。

サブステップdocker-runs.mdスキップ条件
5a. ベースライン評価(ゼロショット)§4`skip_baseline: true`
5b. 完全なトレーニング(デタッチド)§5—
5c. LoRAマージ§6VLM+LoRA 以外
5d. トレーニング後の評価§7—
5e. 推論(5サンプル)§8—

マルチGPU:python train.py の前に torchrun --nproc_per_node=$gpu_count をプレフィックスする。

トレーニングがストリーミングしている間、docker logs -f hft_train を監視する:損失は10-20ステップ以内に減少するはず;平坦な損失(コレクター/ラベルマスキングバグ)、NaN(LRが高すぎる)、およびOOMはすべて実行を停止する — 回復は references/core-rules.md にあります。emit_report: true の場合、ステップ5eの後に references/reporting.md に従って report.py を実行する。

ゲート: 以下のすべて:

  • checkpoints/final/(またはLoRAの場合 checkpoints/merged/)が存在する
  • reports/eval_results.json に数値の主要メトリックがある
  • reports/baseline_results.json が存在する(スキップされていない場合)
  • reports/inference_samples/ に5つのサンプルがある
  • wandb URLが減少する損失を示す

ステップ6 — プッシュおよび再実行スキルのエミッション

目標: 実行を公開し、再リサーチなしで再現可能にする。

push_to_hub: false が明示されていない限り、references/hub-push.md に従ってプッシュする(重み、モデルカード、評価/ベースラインJSON、config.yaml、Dockerfile、requirements.txt、推論サンプル、エミッションされたレポート)。references/pipeline-skill-template.md から <output_dir>/skills/run-<short>/SKILL.md</short></output_dir> をエミッションする — 全プレースホルダーを置換し、完全なYAMLメタデータ + NVIDIA著作権HTMLコメントを含め、欠落している場合、任意のエミッターが失敗するようにする。

ゲート(完了基準): 以下のすべて:

  • ステップ5のゲートが満たされている
  • 解決されたURLにHF Hubリポジトリが存在し、重み + カード + results/ を含む(push_to_hub: false 以外)
  • <output_dir>/skills/run-<short>/SKILL.md</short></output_dir> が存在し、<placeholder></placeholder> が残っておらず、pipeline-skill-template.md に従ってメタデータ + 著作権HTMLコメントを含む

最終メッセージ:wandb URL、HF Hub URL、ベースライン -> ファインチューニングされた主要メトリック、reports/inference_samples/、および再実行スキルのパス。

エラープレイブック

既知のランタイムエラーの場合、再設計する前に references/error-playbook.md の症状 → 最小限の修正表を参照する(NGCエントポイント、PyTorch/Transformersの回帰、numpy ABI、Albumentations bbox、PEFT/チェックポイント、LoRAターゲットの広さ、CV拡張ギャップ、ステップ0のOOM)。そこで1つの行が実行全体で2回発生した場合、detect ルールで compat-workarounds.md に持ち上げる — ステップ1でエラーが発生する前に自動適用される。

コミュニケーションスタイル

  • 簡潔。フィラーなし、リクエストの繰り返しなし;適切な場合は一言回答。
  • 成果物を参照する際は常に直接のHubおよびwandb URLを含める。
  • エラーの場合:何が間違えたか、なぜか、何を変更したかを述べる — メニューなし。
  • 明確な答えがあるリクエストに対して "オプションA/B/C" を提示しない。行動せよ。

例パイプライン

  • tao-rerun-convnext-cifar10
  • tao-rerun-detr-cppe5
  • tao-rerun-segformer-foodseg103
  • tao-rerun-smolvlm-vqav2
GitHubで見る
---
name: tao-finetune-huggingface-model
description: Fine-tune HuggingFace CV, VLM, or LLM models on local NVIDIA GPUs using an NGC PyTorch container, with support for full or LoRA training, dataset handling, and optional model push to the Hub.
license: Apache-2.0
---
<!-- Copyright (c) 2026, NVIDIA CORPORATION. All rights reserved. Licensed under the Apache License, Version 2.0; see http://www.apache.org/licenses/LICENSE-2.0 -->

# tao-finetune-huggingface-model

Local NVIDIA GPU fine-tuning for HuggingFace models, grounded in live-fetched
documentation with curated references as a fallback safety net. One NGC container,
a few focused scripts, one push to HF Hub. Follow the rules in this file; don't
improvise.

**Order of authority (highest first):**

1. **User input** — explicit `model_id`, `dataset_id`, `training_method`, `config.yaml` overrides.
2. **Live research** — model card, HF repo example, author finetune script, HF task docs, paper; always fetched (Step 3 + `references/research-priorities.md`).
3. **Curated references** (`references/*.md`) — fallback when live research is silent/ambiguous.
4. **Your training-data memory** — last resort; suspect, cross-check against (2)/(3).

Conflict resolution between (2) and (3) and the source-line discrepancy note are
in `references/research-priorities.md`.

---

## Inputs

**Required:**
- `model_id` — HuggingFace model ID, e.g. `google/vit-base-patch16-224`

**Conditional credentials (read from the session environment, exported before launching when present):**
- `HF_TOKEN` — only when the model/dataset is **gated** (read) or `push_to_hub` is on (write); public + public + `push_to_hub: false` needs none. Value never read — presence-only via `[ -n "$HF_TOKEN" ]`.
- `WANDB_API_KEY`, `WANDB_PROJECT` — only when WandB is enabled; `WANDB_MODE=disabled` opts out.

**Dataset — exactly one:**
- `dataset_id` — HuggingFace dataset ID *(source: `hf`)*
- `local_dataset_path` — local folder or file *(source: `local`)*; optional
  `local_dataset_format` ∈ {auto, imagefolder, coco, voc, jsonl, arrow, parquet,
  csv} (default: auto-detect).
- *(omit)* — agent recommends popular datasets *(source: `recommend`)*

**Optional (have defaults):**
- `task_type` — auto-detected from config + model card
- `n_train=10000`, `n_eval=1000`, `n_epochs=3`, `lora_r=16`
- `output_dir=./output/<model_short_name>`
- `hf_model_repo` — push target; if unset and HF_TOKEN has write access,
  auto-derived as `<whoami>/<model_short_name>-finetuned`.
- `push_to_hub=True` — set to `False` to skip
- `skip_baseline=False` — skip zero-shot baseline eval

**Optional deliverables (off by default):**
```yaml
emit_progress_log: false   # output_dir/PROGRESS.md (per-step journal)
emit_report:       false   # reports/report.{pdf,html} with curves & samples
emit_unit_tests:   false   # tests/ with fake-data heterogeneous-batch tests
```

All values live in `output_dir/config.yaml`. Never hardcode in Python.

---

## Execution platform

This skill orchestrates *what* to run; the platform skills own *how* to run it on
a GPU host — read them first.

| Concern | Authoritative skill |
|---|---|
| GPU host runtime (driver 580, CUDA Toolkit 13.0, NVIDIA Container Toolkit 1.19.0) | [`tao-skill-bank:tao-setup-nvidia-gpu-host`](../../platform/tao-setup-nvidia-gpu-host/SKILL.md) |
| `docker run` flags, NGC auth, mounts, env passthrough | [`tao-skill-bank:tao-run-on-docker`](../../platform/tao-run-on-docker/SKILL.md) |
| Local Docker job preflight (daemon, GPU smoke) | [`tao-skill-bank:tao-run-on-local-docker`](../../platform/tao-run-on-local-docker/SKILL.md) |

**Default platform:** `local-docker` — build a one-off image (`run-<short>:latest`)
and run it on the local Docker daemon. Ask only when the user explicitly needs a
different backend (Brev remote GPU, SLURM/Kubernetes); then run that platform's
Preflight first and route the Steps 4–5 `docker run` commands through it. The
GPU-runtime and presence-only credential preflights (values never read), the
canonical `docker run` flag set, the `list_tao_platforms.py` selection command, and
the workflow-specific flags (`--entrypoint /bin/bash -lc`, `PYTORCH_CUDA_ALLOC_CONF`,
`--name hft_train`) are in `references/workflow-intake-preflight.md`.

---

## References — fallback safety net

Consulted **only** when live research is silent, ambiguous, or unavailable; live
docs always win for the specific model and current API. Each step links the
references it needs; full catalog in `references/detailed-workflow.md`.

Always-on: `core-rules.md`, `error-playbook.md`, `compat-workarounds.md`,
`model-discovery.md`, `dataset-recommendations.md`, `dataset-sources.md`,
`dataset-patterns.md`, `hardware-container.md`, `research-priorities.md`,
`cv-scripts.md`, `vlm-scripts.md`, `docker-runs.md`, `hub-push.md`,
`pipeline-skill-template.md`, `deliverables.md`. Opt-in (when their flag/need
applies): `progress-tracking.md`, `testing.md`, `reporting.md`,
`workflow-intake-preflight.md`, `workflow-generate-train.md`, `workflow-push-rerun.md`.

**Rule:** before falling back, log the live source you tried and why it was
insufficient (`config.yaml` `notes:`, and PROGRESS.md if enabled). `[FETCH LIVE]`
markers in `cv-scripts.md` / `vlm-scripts.md` are a research checklist, not code to
inline — refetch the listed URL if a block has no Step 3 finding.

---

## Core rules

Non-negotiable behaviors. **Short version** (full enumeration —
hallucinated-imports list, never-without-approval list, full error-recovery and
hardware-sizing tables — in `references/core-rules.md`, consult before any
training-time decision):

- **Your HF-library knowledge is outdated.** Fetch live docs (model card, HF
  repo example, task doc) before writing any ML code — don't generate trainer
  args / collator / transforms from memory (Step 3).
- **Smoke-test on real data with `--max_steps 1`** before any full run; no batch
  launches without a verified smoke.
- **Never silently substitute** model_id, dataset_id, or training_method — if
  what the user asked for doesn't load, stop and ask.
- **Error recovery is minimal-change.** OOM → halve batch, double grad_accum,
  enable gradient checkpointing (no LoRA switch without approval); NaN → reduce
  LR 10×; flat loss → inspect collator; same error 3× → stop and ask. Don't loop.
- **Dataset columns verified BEFORE the collator** — rename in `prepare_data.py`;
  restructuring needed → stop and ask.
- **Hardware-sizing thumb (bf16):** ≤3B → 24 GB, 7–13B → 80 GB, 30B+ → multi-GPU
  or LoRA on 1× 80 GB, 70B+ → 8× 80 GB or LoRA. Full finetune won't fit and no
  LoRA requested → ask before switching.

---

## Workflow — 6 steps

Single pass, sequential; each step has a clear gate before the next begins.

### Step 1 — Inspect & qualify

**Goal:** decide whether to proceed. Probe model + dataset, apply accept/reject,
register applicable compat fixes, write the initial `config.yaml`.

Prerequisites: `MODEL_ID`, optional `DATASET_ID` / `local_dataset_path`,
optional `HF_TOKEN`, `OUTPUT_DIR` (default `./output/<model_short_name>`). Probes
run in a CPU-only `python:3.12-slim` Docker container (bind-mounted `.probe/`
scratch) so the host needs no virtualenv — Docker must exist first. Docker-presence
guard, container env, full probe invocation, and the model/dataset probe scripts
are in `references/workflow-intake-preflight.md`, `references/model-discovery.md`,
and `references/dataset-sources.md`.

Probe requirements:

- Model: load `AutoConfig`, read model-card tags, detect task from
  `architectures` + tags + card examples (fallback logging in `model-discovery.md`).
- Dataset: for recommended datasets, first present 3-5 choices from
  `dataset-recommendations.md`; for local data, bind-mount read-only and use
  `dataset-sources.md` format detection.
- Reject early if the model config fails, the task is out of scope, no recipe
  source exists, or the dataset cannot load / match the task schema.
- Evaluate `compat-workarounds.md` against the model/task; defer hardware-dependent
  rules to Step 2.

Write the initial `config.yaml` (`model_id`, `task`, `dataset_id` or
`local_dataset_path`, `research_sources: []` filled in Step 3,
`applicable_workarounds:` from Step 1, `notes: []` for reference fallbacks,
`push_to_hub: true` default — annotated template in
`references/workflow-intake-preflight.md`). Optionally `rm -rf "$OUTPUT_DIR/.probe"`
once the gate is met.

**Gate:** `config.yaml` exists with model, dataset, task, applicable_workarounds;
do not proceed if any field is missing.

---

### Step 2 — Hardware audit & NGC image

**Goal:** verify Docker + GPU + disk, pick the NGC PyTorch image live, finalize
hardware-dependent compat rules.

**2a. Audit (hard gate)** — three checks (commands in
`references/workflow-intake-preflight.md`):
1. GPU host runtime — `tao-setup-nvidia-gpu-host`'s
   `setup-nvidia-gpu-host.sh --backend docker --check-only`; on fail, ask approval
   then re-run with `--install --yes`.
2. Free-disk soft-warn — override via `MIN_DISK_GB` (default 100 GB); recommend
   ≥ 100 GB for NGC base (~20 GB) + HF cache + checkpoints + data.
3. Conditional credential presence (from the session environment, values never
   read) — `HF_TOKEN` only when gated or `push_to_hub` is on; `WANDB_*` only when
   WandB is on.

**Do not proceed to Step 4 on a hard-fail** — Step 4's `docker build` pulls a
20+ GB NGC base, and a missing `nvidia-container-toolkit` only surfaces later as
`could not select device driver "" with capabilities: [[gpu]]`. Record `gpu_count`,
`gpu_name`, `driver_major`, `vram_gb_per_gpu` in `config.yaml`.

**2b. Pick NGC image (live):** from the NVIDIA Deep Learning Frameworks support
matrix (<https://docs.nvidia.com/deeplearning/frameworks/support-matrix/index.html>),
PyTorch NGC container section, pick the highest-versioned image where
`Min driver ≤ detected driver_major` and container CUDA `≤` host CUDA Toolkit
(match closely so cuDNN / TensorRT line up). Do **not** reject an image for an
`aN`/`bN`/`rcN` PyTorch tag — NGC validates the full image; pick the newest
CUDA-aligned one and let `compat-workarounds.md` handle per-version issues. If the
matrix is unreachable, use the fallbacks in `references/hardware-container.md`;
default `nvcr.io/nvidia/pytorch:24.09-py3` (driver ≥ 545; SDPA+GQA bug — if
`num_key_value_heads < num_attention_heads`, set `attn_implementation: "eager"`).
Record `ngc_image` in `config.yaml`.

**2c. Re-evaluate hardware-dependent compat rules:** re-run the
`compat-workarounds.md` walk for entries whose `detect` needs `hw`; update
`applicable_workarounds:` in place.

**2d. Model-fit check:** estimate `param_bytes ≈ 2×param_count` (bf16); if
> 60% of `vram_gb_per_gpu × 1e9`, recommend LoRA in the user-facing summary.

**Gate:** `config.yaml` has `ngc_image`, `gpu_count`, `gpu_name`, `driver_major`,
`vram_gb_per_gpu`; hardware-dependent compat fixes recorded.

---

### Step 3 — Research the recipe

**Goal:** fetch the live recipe — training-data knowledge of
`transformers`/`trl`/`peft` is suspect, so Step 3 is non-negotiable. Walk
`references/research-priorities.md` in priority order (Priority 1 → 6); stop once
you have, for the detected task:

- `AutoModel` / processor class
- Train + eval transforms
- Collator
- `compute_metrics`
- Hyperparameter hints (LR, batch size, epochs, scheduler)

Record findings in `meta/recipe.md`, append source URLs to
`config.yaml: research_sources:`. A slot with no live finding falls back to the
matching scaffold (`cv-scripts.md` / `vlm-scripts.md`), logged as "fallback to
scaffold — no live source for <slot>" under `notes:`. Conflict-resolution rules
are in `references/research-priorities.md`.

**Gate:** every required slot filled, with a source URL or scaffold-fallback note.

---

### Step 4 — Generate project & smoke-test

**Goal:** write all scripts, build the image, prepare data, run a 1-step smoke on
real data (one `docker build`, two `docker run`s).

**4a. Generate project files** in `output_dir/`: `config.yaml`, `Dockerfile`,
`requirements.txt`, `prepare_data.py`, `train.py`, `run_eval.py`, `infer.py`,
optional `merge_lora.py`, optional `tests/`, `.gitignore`. Live Step 3 research is
authority; `cv-scripts.md` / `vlm-scripts.md` give scaffold shape only. Apply every
`applicable_workarounds` entry as a Dockerfile block, requirement pin, config
override, or runtime env var. Hard rules: `run_eval.py` keeps that exact filename
(avoids colliding with the HF `evaluate` package); every generated `.py` starts
with the NVIDIA Apache-2.0 copyright header and any emitter fails when it is
missing; `emit_unit_tests: true` generates and runs tests per
`references/testing.md`. Script bodies, Dockerfile shape, and the emitter contract
are in `references/workflow-generate-train.md`.

**4b. Build, prepare, smoke** — `docker build -t run-<short>:latest .`, then
`prepare_data` and the `--smoke --max_steps 1` run (`references/docker-runs.md`
§1-3). Smoke pass criteria (in `logs/smoke.log`):
- No exception
- Loss is finite (not `0.0`, not `NaN`)
- `grad_norm > 0` at step 1

If `emit_unit_tests: true`, also run `pytest tests/` in the container. Any failure → STOP.

**4c. Preflight summary** — before full training, print and verify: reference URL,
dataset columns, Hub target, monitoring target, NGC image, hardware, smoke loss/grad norm.

**Gate:** project files written, image built, smoke PASSED, preflight has no
blank fields.

---

### Step 5 — Train, evaluate, infer

**Goal:** baseline eval, full training, post-train eval, optional LoRA merge, 5
inference samples (all commands: `references/docker-runs.md` §4-8).

| Sub-step | docker-runs.md | Skip if |
|---|---|---|
| 5a. Baseline eval (zero-shot) | §4 | `skip_baseline: true` |
| 5b. Full training (detached) | §5 | — |
| 5c. LoRA merge | §6 | not VLM+LoRA |
| 5d. Post-train eval | §7 | — |
| 5e. Inference (5 samples) | §8 | — |

Multi-GPU: prepend `torchrun --nproc_per_node=$gpu_count` to `python train.py`.

While training streams, watch `docker logs -f hft_train`: loss should drop within
10-20 steps; flat loss (collator/label-masking bug), NaN (LR too high), and OOM
all stop the run — recovery in `references/core-rules.md`. If `emit_report: true`,
run `report.py` after Step 5e per `references/reporting.md`.

**Gate:** all of:
- `checkpoints/final/` (or `checkpoints/merged/` for LoRA) exists
- `reports/eval_results.json` has a numeric primary metric
- `reports/baseline_results.json` exists (unless skipped)
- `reports/inference_samples/` has 5 samples
- wandb URL shows descending loss

---

### Step 6 — Push & emit rerun skill

**Goal:** publish the run and make it reproducible without re-research.

Push per `references/hub-push.md` (weights, model card, eval/baseline JSONs,
`config.yaml`, `Dockerfile`, `requirements.txt`, inference samples, reports when
emitted) unless `push_to_hub: false` is explicit. Emit
`<output_dir>/skills/run-<short>/SKILL.md` from
`references/pipeline-skill-template.md` — substitute every placeholder, include
full YAML metadata + the NVIDIA copyright HTML comment, and make any emitter fail
if those are missing.

**Gate (Done criteria):** all of:
- Step 5 gate met
- HF Hub repo exists at the resolved URL with weights + card + `results/`
  (unless `push_to_hub: false`)
- `<output_dir>/skills/run-<short>/SKILL.md` exists, no `<placeholder>` left,
  with metadata + copyright HTML comment per `pipeline-skill-template.md`

Final message: wandb URL, HF Hub URL, baseline -> fine-tuned primary metric,
`reports/inference_samples/`, and the rerun skill path.

---

## Error playbook

On a known runtime error, consult the symptom → minimal-fix table in
`references/error-playbook.md` (NGC entrypoint, PyTorch/Transformers regressions,
numpy ABI, Albumentations bbox, PEFT/checkpointing, LoRA target breadth, CV
augmentation gaps, OOM at step 0) before redesigning anything. When a row there
fires twice across runs, lift it into `compat-workarounds.md` with a `detect` rule
— auto-applied in Step 1 before the error can fire.

---

## Communication style

- Terse. No filler, no restating the request; one-word answers when appropriate.
- Always include direct Hub and wandb URLs when referencing artifacts.
- On error: state what went wrong, why, what you changed — no menus.
- Never present "Option A/B/C" for a request with a clear answer. Act.

## Example pipelines

- [tao-rerun-convnext-cifar10](references/tao-rerun-convnext-cifar10.md)
- [tao-rerun-detr-cppe5](references/tao-rerun-detr-cppe5.md)
- [tao-rerun-segformer-foodseg103](references/tao-rerun-segformer-foodseg103.md)
- [tao-rerun-smolvlm-vqav2](references/tao-rerun-smolvlm-vqav2.md)

すべてのファイル

69件のファイル

tao-finetune-huggingface-modelをインストール

スキルファイルをダウンロードして、.claude/skills/ ディレクトリに展開してください。

ZIPをダウンロード

リポジトリをクローンし、スキルファイルをプロジェクトにコピーしてください。

git clone https://github.com/NVIDIA/skills/tree/main/skills/tao-finetune-huggingface-model # Copy SKILL.md to your .claude/skills/ directory

コピー コピー
クイックセットアップ: スキルフォルダを .claude/skills/ にコピーしてください。Claude はそのスキルを自動的に検出し、使用します。
リポジトリ NVIDIA/skills

関連スキル

web-search
更新された時間 2026年6月29日
webapp-testing
更新された時間 2026年6月29日
lark-base
更新された時間 2026年7月5日
agentmail
更新された時間 2026年6月29日
OR