tao-finetune-huggingface-model
NVIDIA/skills
使用 NGC PyTorch 容器在本地 NVIDIA GPU 上微调 HuggingFace 的 CV、VLM 或 LLM 模型,支持全量或 LoRA 训练、数据集处理,以及可选的模型推送到 Hub。
...展开全部tao-finetune-huggingface-model
针对 HuggingFace 模型在本地 NVIDIA GPU 上进行微调,以实时获取的文档为基础,并以经过策划的参考资料作为备用安全网。一个 NGC 容器,几个专注的脚本,一次推送到 HF Hub。请遵循本文件中的规则;不要自行发挥。
权威顺序(优先级最高在前):
- 用户输入 — 明确的
model_id、dataset_id、training_method、config.yaml覆盖。 - 实时研究 — 模型卡片、HF 仓库示例、作者微调脚本、HF 任务文档、论文;始终获取(第 3 步 +
references/research-priorities.md)。 - 策划的参考资料(
references/*.md) — 当实时研究沉默或模棱两可时的后备方案。 - 你的训练数据记忆 — 最后手段;存疑,需与 (2)/(3) 交叉核对。
(2) 和 (3) 之间的冲突解决以及源行差异说明位于 references/research-priorities.md。
输入
必需:
model_id— HuggingFace 模型 ID,例如google/vit-base-patch16-224
条件凭据(从会话环境读取,在启动前导出,如果存在):
HF_TOKEN— 仅当模型/数据集是受限(读取)或push_to_hub开启(写入)时;公开 + 公开 +push_to_hub: false不需要。值从不读取 — 仅通过[ -n "$HF_TOKEN" ]检查存在性。WANDB_API_KEY,WANDB_PROJECT— 仅当启用 WandB 时;WANDB_MODE=disabled选择退出。
数据集 — 恰好一个:
dataset_id— HuggingFace 数据集 ID (来源:hf)local_dataset_path— 本地文件夹或文件 (来源:local);可选local_dataset_format∈ {auto, imagefolder, coco, voc, jsonl, arrow, parquet, csv}(默认:自动检测)。- (省略) — 代理推荐流行数据集 (来源:
recommend)
可选(有默认值):
task_type— 从配置 + 模型卡片自动检测n_train=10000,n_eval=1000,n_epochs=3,lora_r=16output_dir=./output/<model_short_name></model_short_name>hf_model_repo— 推送目标;如果未设置且 HF_TOKEN 具有写入权限,自动派生为<whoami>/<model_short_name>-finetuned</model_short_name></whoami>。push_to_hub=True— 设置为False以跳过skip_baseline=False— 跳过零样本基线评估
可选交付物(默认关闭):
emit_progress_log: false # output_dir/PROGRESS.md(逐步日志)
emit_report: false # reports/report.{pdf,html} 包含曲线和样本
emit_unit_tests: false # tests/ 包含假数据异构批次测试
所有值位于 output_dir/config.yaml。切勿在 Python 中硬编码。
执行平台
此技能编排运行什么;平台技能拥有在 GPU 主机上运行如何 — 请先阅读它们。
| 关注点 | 权威技能 |
|---|---|
| GPU 主机运行时(驱动 580,CUDA Toolkit 13.0,NVIDIA Container Toolkit 1.19.0) | `tao-skill-bank:tao-setup-nvidia-gpu-host` |
| `docker run` 标志、NGC 认证、挂载、环境变量传递 | `tao-skill-bank:tao-run-on-docker` |
| 本地 Docker 作业预检(守护进程、GPU 冒烟测试) | `tao-skill-bank:tao-run-on-local-docker` |
默认平台: local-docker — 构建一次性镜像(run-<short>:latest</short>)并在本地 Docker 守护进程上运行它。仅当用户明确需要不同后端(Brev 远程 GPU、SLURM/Kubernetes)时才询问;然后先运行该平台的预检,并通过它将第 4–5 步的 docker run 命令路由到其中。GPU 运行时和仅存在性凭据预检(值从不读取)、标准的 docker run 标志集、list_tao_platforms.py 选择命令以及工作流特定标志(--entrypoint /bin/bash -lc、PYTORCH_CUDA_ALLOC_CONF、--name hft_train)位于 references/workflow-intake-preflight.md。
参考资料 — 备用安全网
仅在实时研究沉默、模棱两可或不可用时查阅;实时文档对于特定模型和当前 API 始终优先。每个步骤链接其所需的参考资料;完整目录在 references/detailed-workflow.md 中。
始终开启:core-rules.md、error-playbook.md、compat-workarounds.md、model-discovery.md、dataset-recommendations.md、dataset-sources.md、dataset-patterns.md、hardware-container.md、research-priorities.md、cv-scripts.md、vlm-scripts.md、docker-runs.md、hub-push.md、pipeline-skill-template.md、deliverables.md。可选(当其标志/需求适用时):progress-tracking.md、testing.md、reporting.md、workflow-intake-preflight.md、workflow-generate-train.md、workflow-push-rerun.md。
规则: 在回退之前,记录你尝试过的实时源及其不足的原因(config.yaml 中的 notes:,如果启用则记录 PROGRESS.md)。cv-scripts.md / vlm-scripts.md 中的 [FETCH LIVE] 标记是研究清单,不是内联代码 — 如果某个块没有第 3 步的发现,请重新获取列出的 URL。
核心规则
不可协商的行为。简短版本(完整枚举 — 幻觉导入列表、未经批准绝不允许列表、完整错误恢复和硬件调整大小表 — 在 references/core-rules.md 中,在任何训练时决策前咨询):
- 你的 HF 库知识已过时。 在编写任何 ML 代码之前获取实时文档(模型卡片、HF 仓库示例、任务文档) — 不要从记忆中生成训练器参数 / 整理器 / 变换(第 3 步)。
- 使用
--max_steps 1在真实数据上进行冒烟测试 在任何完整运行之前;未经验证的冒烟测试不得批量启动。 - 切勿静默替换 model_id、dataset_id 或 training_method — 如果用户请求的内容无法加载,停止并询问。
- 错误恢复是最小更改。 OOM → 减半批次,加倍 grad_accum,启用梯度检查点(未经批准不切换 LoRA);NaN → 将 LR 减少 10 倍;平坦损失 → 检查整理器;相同错误 3 次 → 停止并询问。不要循环。
- 数据集列在整理器之前验证 — 在
prepare_data.py中重命名;需要重构 → 停止并询问。 - 硬件调整大小经验法则(bf16): ≤3B → 24 GB,7–13B → 80 GB,30B+ → 多 GPU 或 1× 80 GB 上的 LoRA,70B+ → 8× 80 GB 或 LoRA。完整微调无法容纳且未请求 LoRA → 切换前询问。
工作流 — 6 个步骤
单次传递,顺序执行;每个步骤在下一个开始之前都有明确的门控。
第 1 步 — 检查与资格认定
目标: 决定是否继续。探测模型 + 数据集,应用接受/拒绝,注册适用的兼容修复,编写初始 config.yaml。
先决条件:MODEL_ID、可选 DATASET_ID / local_dataset_path、可选 HF_TOKEN、OUTPUT_DIR(默认 ./output/<model_short_name></model_short_name>)。探测在仅 CPU 的 python:3.12-slim Docker 容器中运行(绑定挂载 .probe/scratch),因此主机不需要虚拟环境 — Docker 必须首先存在。Docker 存在性保护、容器环境、完整探测调用以及模型/数据集探测脚本位于 references/workflow-intake-preflight.md、references/model-discovery.md 和 references/dataset-sources.md。
探测要求:
- 模型:加载
AutoConfig,读取模型卡片标签,从architectures+ 标签 + 卡片示例中检测任务(回退日志记录在model-discovery.md中)。 - 数据集:对于推荐的数据集,首先从
dataset-recommendations.md展示 3-5 个选择;对于本地数据,只读绑定挂载并使用dataset-sources.md格式检测。 - 如果模型配置失败、任务超出范围、没有配方源或数据集无法加载/匹配任务模式,则提前拒绝。
- 根据模型/任务评估
compat-workarounds.md;将硬件相关规则推迟到第 2 步。
编写初始 config.yaml(model_id、task、dataset_id 或 local_dataset_path、research_sources: [] 在第 3 步填充、applicable_workarounds: 来自第 1 步、notes: [] 用于参考资料回退、push_to_hub: true 默认值 — 注释模板在 references/workflow-intake-preflight.md 中)。门控满足后可选 rm -rf "$OUTPUT_DIR/.probe"。
门控: config.yaml 存在,包含模型、数据集、任务、applicable_workarounds;如果任何字段缺失,不得继续。
第 2 步 — 硬件审计与 NGC 镜像
目标: 验证 Docker + GPU + 磁盘,实时选择 NGC PyTorch 镜像,最终确定硬件相关的兼容规则。
2a. 审计(硬门控) — 三个检查(命令在 references/workflow-intake-preflight.md 中):
- GPU 主机运行时 —
tao-setup-nvidia-gpu-host的setup-nvidia-gpu-host.sh --backend docker --check-only;如果失败,请求批准然后使用--install --yes重新运行。 - 空闲磁盘软警告 — 通过
MIN_DISK_GB覆盖(默认 100 GB);建议 ≥ 100 GB 用于 NGC 基础(~20 GB)+ HF 缓存 + 检查点 + 数据。 - 条件凭据存在性(来自会话环境,值从不读取) —
HF_TOKEN仅在受限或push_to_hub开启时;WANDB_*仅在 WandB 开启时。
硬失败时不得进入第 4 步 — 第 4 步的 docker build 拉取 20+ GB NGC 基础,而缺失的 nvidia-container-toolkit 稍后才会显示为 could not select device driver "" with capabilities: [[gpu]]。在 config.yaml 中记录 gpu_count、gpu_name、driver_major、vram_gb_per_gpu。
2b. 选择 NGC 镜像(实时): 从 NVIDIA 深度学习框架支持矩阵(https://docs.nvidia.com/deeplearning/frameworks/support-matrix/index.html),PyTorch NGC 容器部分,选择版本号最高的镜像,其中 Min driver ≤ 检测到的 driver_major 且容器 CUDA ≤ 主机 CUDA Toolkit(尽可能匹配以便 cuDNN / TensorRT 对齐)。不要因 aN/bN/rcN PyTorch 标签而拒绝镜像 — NGC 验证整个镜像;选择最新的 CUDA 对齐版本,让 compat-workarounds.md 处理每个版本的问题。如果矩阵不可达,使用 references/hardware-container.md 中的回退;默认 nvcr.io/nvidia/pytorch:24.09-py3(驱动 ≥ 545;SDPA+GQA 错误 — 如果 num_key_value_heads , setattn_implementation: "eager")。在config.yaml中记录ngc_image`。
2c. 重新评估硬件相关的兼容规则: 重新运行 compat-workarounds.md 遍历,针对其 detect 需要 hw 的条目;就地更新 applicable_workarounds:。
2d. 模型适配检查: 估计 param_bytes ≈ 2×param_count(bf16);如果
60% of
vram_gb_per_gpu × 1e9,在面向用户的摘要中推荐 LoRA。
门控: config.yaml 包含 ngc_image、gpu_count、gpu_name、driver_major、vram_gb_per_gpu;记录了硬件相关的兼容修复。
第 3 步 — 研究配方
目标: 获取实时配方 — transformers/trl/peft 的训练数据知识存疑,因此第 3 步是不可协商的。按优先级顺序遍历 references/research-priorities.md(优先级 1 → 6);一旦你为检测到的任务获得以下内容,即停止:
AutoModel/ 处理器类- 训练 + 评估变换
- 整理器
compute_metrics- 超参数提示(LR、批次大小、epochs、调度器)
将发现记录在 meta/recipe.md 中,将源 URL 追加到 config.yaml: research_sources:。没有实时发现的槽位回退到匹配的脚手架(cv-scripts.md / vlm-scripts.md),在 notes: 下记录为“回退到脚手架 — 没有实时源用于”。冲突解决规则在 references/research-priorities.md 中。
门控: 每个必需的槽位已填充,带有源 URL 或脚手架回退说明。
第 4 步 — 生成项目与冒烟测试
目标: 编写所有脚本,构建镜像,准备数据,在真实数据上运行 1 步冒烟测试(一个 docker build,两个 docker run)。
4a. 在项目文件 output_dir/ 中生成:config.yaml、Dockerfile、requirements.txt、prepare_data.py、train.py、run_eval.py、infer.py、可选 merge_lora.py、可选 tests/、.gitignore。实时第 3 步研究是权威;cv-scripts.md / vlm-scripts.md 仅提供脚手架形状。将每个 applicable_workarounds 条目应用为 Dockerfile 块、要求固定、配置覆盖或运行时环境变量。硬规则:run_eval.py 保持确切文件名(避免与 HF evaluate 包冲突);每个生成的 .py 以 NVIDIA Apache-2.0 版权标头开头,任何发射器在缺少时失败;emit_unit_tests: true 根据 references/testing.md 生成并运行测试。脚本主体、Dockerfile 形状和发射器合同在 references/workflow-generate-train.md 中。
4b. 构建、准备、冒烟 — docker build -t run-<short>:latest .</short>,然后 prepare_data 和 --smoke --max_steps 1 运行(references/docker-runs.md§1-3)。冒烟通过标准(在 logs/smoke.log 中):
- 无异常
- 损失是有限的(不是
0.0,不是NaN) - 步骤 1 时
grad_norm > 0
如果 emit_unit_tests: true,还在容器中运行 pytest tests/。任何失败 → 停止。
4c. 预检摘要 — 在完整训练之前,打印并验证:参考 URL、数据集列、Hub 目标、监控目标、NGC 镜像、硬件、冒烟损失/梯度范数。
门控: 项目文件已写入,镜像已构建,冒烟通过,预检没有空白字段。
第 5 步 — 训练、评估、推理
目标: 基线评估、完整训练、训练后评估、可选 LoRA 合并、5 个推理样本(所有命令:references/docker-runs.md §4-8)。
| 子步骤 | docker-runs.md | 跳过条件 |
|---|---|---|
| 5a. 基线评估(零样本) | §4 | `skip_baseline: true` |
| 5b. 完整训练(分离) | §5 | — |
| 5c. LoRA 合并 | §6 | 非 VLM+LoRA |
| 5d. 训练后评估 | §7 | — |
| 5e. 推理(5 个样本) | §8 | — |
多 GPU:在 python train.py 前添加 torchrun --nproc_per_node=$gpu_count。
训练流式传输时,观察 docker logs -f hft_train:损失应在 10-20 步内下降;平坦损失(整理器/标签掩码错误)、NaN(LR 过高)和 OOM 都会停止运行 — 恢复在 references/core-rules.md 中。如果 emit_report: true,在第 5e 步后根据 references/reporting.md 运行 report.py。
门控: 全部:
checkpoints/final/(或 LoRA 的checkpoints/merged/)存在reports/eval_results.json具有数值主指标reports/baseline_results.json存在(除非跳过)reports/inference_samples/有 5 个样本- wandb URL 显示下降的损失
第 6 步 — 推送与发射重跑技能
目标: 发布运行并使其可重现,无需重新研究。
根据 references/hub-push.md 推送(权重、模型卡片、评估/基线 JSON、config.yaml、Dockerfile、requirements.txt、推理样本、发射的报告),除非明确 push_to_hub: false。从 references/pipeline-skill-template.md 发射 <output_dir>/skills/run-<short>/SKILL.md</short></output_dir> — 替换每个占位符,包含完整的 YAML 元数据 + NVIDIA 版权 HTML 注释,如果缺少这些内容则使任何发射器失败。
门控(完成标准): 全部:
- 第 5 步门控满足
- HF Hub 仓库存在于解析的 URL,包含权重 + 卡片 +
results/(除非push_to_hub: false) <output_dir>/skills/run-<short>/SKILL.md</short></output_dir>存在,没有<placeholder></placeholder>遗留,带有pipeline-skill-template.md规定的元数据 + 版权 HTML 注释
最终消息:wandb URL、HF Hub URL、基线 -> 微调主指标、reports/inference_samples/ 以及重跑技能路径。
错误剧本
在已知的运行时错误上,咨询 references/error-playbook.md 中的症状 → 最小修复表(NGC 入口点、PyTorch/Transformers 回归、numpy ABI、Albumentations bbox、PEFT/检查点、LoRA 目标广度、CV 增强差距、步骤 0 时的 OOM)在重新设计任何内容之前。当某一行在跨运行中触发两次时,将其提升为 compat-workaroutines.md 中的 detect 规则 — 在第 1 步中自动应用,在错误触发之前。
沟通风格
- 简洁。无填充词,无重述请求;适当时使用单字回答。
- 引用工件时始终包含直接的 Hub 和 wandb URL。
- 出错时:说明出了什么问题、原因、你做了什么更改 — 无菜单。
- 对于有明确答案的请求,切勿呈现“选项 A/B/C”。行动。
示例管道
- tao-rerun-convnext-cifar10
- tao-rerun-detr-cppe5
- tao-rerun-segformer-foodseg103
- tao-rerun-smolvlm-vqav2
---
name: tao-finetune-huggingface-model
description: Fine-tune HuggingFace CV, VLM, or LLM models on local NVIDIA GPUs using an NGC PyTorch container, with support for full or LoRA training, dataset handling, and optional model push to the Hub.
license: Apache-2.0
---
<!-- Copyright (c) 2026, NVIDIA CORPORATION. All rights reserved. Licensed under the Apache License, Version 2.0; see http://www.apache.org/licenses/LICENSE-2.0 -->
# tao-finetune-huggingface-model
Local NVIDIA GPU fine-tuning for HuggingFace models, grounded in live-fetched
documentation with curated references as a fallback safety net. One NGC container,
a few focused scripts, one push to HF Hub. Follow the rules in this file; don't
improvise.
**Order of authority (highest first):**
1. **User input** — explicit `model_id`, `dataset_id`, `training_method`, `config.yaml` overrides.
2. **Live research** — model card, HF repo example, author finetune script, HF task docs, paper; always fetched (Step 3 + `references/research-priorities.md`).
3. **Curated references** (`references/*.md`) — fallback when live research is silent/ambiguous.
4. **Your training-data memory** — last resort; suspect, cross-check against (2)/(3).
Conflict resolution between (2) and (3) and the source-line discrepancy note are
in `references/research-priorities.md`.
---
## Inputs
**Required:**
- `model_id` — HuggingFace model ID, e.g. `google/vit-base-patch16-224`
**Conditional credentials (read from the session environment, exported before launching when present):**
- `HF_TOKEN` — only when the model/dataset is **gated** (read) or `push_to_hub` is on (write); public + public + `push_to_hub: false` needs none. Value never read — presence-only via `[ -n "$HF_TOKEN" ]`.
- `WANDB_API_KEY`, `WANDB_PROJECT` — only when WandB is enabled; `WANDB_MODE=disabled` opts out.
**Dataset — exactly one:**
- `dataset_id` — HuggingFace dataset ID *(source: `hf`)*
- `local_dataset_path` — local folder or file *(source: `local`)*; optional
`local_dataset_format` ∈ {auto, imagefolder, coco, voc, jsonl, arrow, parquet,
csv} (default: auto-detect).
- *(omit)* — agent recommends popular datasets *(source: `recommend`)*
**Optional (have defaults):**
- `task_type` — auto-detected from config + model card
- `n_train=10000`, `n_eval=1000`, `n_epochs=3`, `lora_r=16`
- `output_dir=./output/<model_short_name>`
- `hf_model_repo` — push target; if unset and HF_TOKEN has write access,
auto-derived as `<whoami>/<model_short_name>-finetuned`.
- `push_to_hub=True` — set to `False` to skip
- `skip_baseline=False` — skip zero-shot baseline eval
**Optional deliverables (off by default):**
```yaml
emit_progress_log: false # output_dir/PROGRESS.md (per-step journal)
emit_report: false # reports/report.{pdf,html} with curves & samples
emit_unit_tests: false # tests/ with fake-data heterogeneous-batch tests
```
All values live in `output_dir/config.yaml`. Never hardcode in Python.
---
## Execution platform
This skill orchestrates *what* to run; the platform skills own *how* to run it on
a GPU host — read them first.
| Concern | Authoritative skill |
|---|---|
| GPU host runtime (driver 580, CUDA Toolkit 13.0, NVIDIA Container Toolkit 1.19.0) | [`tao-skill-bank:tao-setup-nvidia-gpu-host`](../../platform/tao-setup-nvidia-gpu-host/SKILL.md) |
| `docker run` flags, NGC auth, mounts, env passthrough | [`tao-skill-bank:tao-run-on-docker`](../../platform/tao-run-on-docker/SKILL.md) |
| Local Docker job preflight (daemon, GPU smoke) | [`tao-skill-bank:tao-run-on-local-docker`](../../platform/tao-run-on-local-docker/SKILL.md) |
**Default platform:** `local-docker` — build a one-off image (`run-<short>:latest`)
and run it on the local Docker daemon. Ask only when the user explicitly needs a
different backend (Brev remote GPU, SLURM/Kubernetes); then run that platform's
Preflight first and route the Steps 4–5 `docker run` commands through it. The
GPU-runtime and presence-only credential preflights (values never read), the
canonical `docker run` flag set, the `list_tao_platforms.py` selection command, and
the workflow-specific flags (`--entrypoint /bin/bash -lc`, `PYTORCH_CUDA_ALLOC_CONF`,
`--name hft_train`) are in `references/workflow-intake-preflight.md`.
---
## References — fallback safety net
Consulted **only** when live research is silent, ambiguous, or unavailable; live
docs always win for the specific model and current API. Each step links the
references it needs; full catalog in `references/detailed-workflow.md`.
Always-on: `core-rules.md`, `error-playbook.md`, `compat-workarounds.md`,
`model-discovery.md`, `dataset-recommendations.md`, `dataset-sources.md`,
`dataset-patterns.md`, `hardware-container.md`, `research-priorities.md`,
`cv-scripts.md`, `vlm-scripts.md`, `docker-runs.md`, `hub-push.md`,
`pipeline-skill-template.md`, `deliverables.md`. Opt-in (when their flag/need
applies): `progress-tracking.md`, `testing.md`, `reporting.md`,
`workflow-intake-preflight.md`, `workflow-generate-train.md`, `workflow-push-rerun.md`.
**Rule:** before falling back, log the live source you tried and why it was
insufficient (`config.yaml` `notes:`, and PROGRESS.md if enabled). `[FETCH LIVE]`
markers in `cv-scripts.md` / `vlm-scripts.md` are a research checklist, not code to
inline — refetch the listed URL if a block has no Step 3 finding.
---
## Core rules
Non-negotiable behaviors. **Short version** (full enumeration —
hallucinated-imports list, never-without-approval list, full error-recovery and
hardware-sizing tables — in `references/core-rules.md`, consult before any
training-time decision):
- **Your HF-library knowledge is outdated.** Fetch live docs (model card, HF
repo example, task doc) before writing any ML code — don't generate trainer
args / collator / transforms from memory (Step 3).
- **Smoke-test on real data with `--max_steps 1`** before any full run; no batch
launches without a verified smoke.
- **Never silently substitute** model_id, dataset_id, or training_method — if
what the user asked for doesn't load, stop and ask.
- **Error recovery is minimal-change.** OOM → halve batch, double grad_accum,
enable gradient checkpointing (no LoRA switch without approval); NaN → reduce
LR 10×; flat loss → inspect collator; same error 3× → stop and ask. Don't loop.
- **Dataset columns verified BEFORE the collator** — rename in `prepare_data.py`;
restructuring needed → stop and ask.
- **Hardware-sizing thumb (bf16):** ≤3B → 24 GB, 7–13B → 80 GB, 30B+ → multi-GPU
or LoRA on 1× 80 GB, 70B+ → 8× 80 GB or LoRA. Full finetune won't fit and no
LoRA requested → ask before switching.
---
## Workflow — 6 steps
Single pass, sequential; each step has a clear gate before the next begins.
### Step 1 — Inspect & qualify
**Goal:** decide whether to proceed. Probe model + dataset, apply accept/reject,
register applicable compat fixes, write the initial `config.yaml`.
Prerequisites: `MODEL_ID`, optional `DATASET_ID` / `local_dataset_path`,
optional `HF_TOKEN`, `OUTPUT_DIR` (default `./output/<model_short_name>`). Probes
run in a CPU-only `python:3.12-slim` Docker container (bind-mounted `.probe/`
scratch) so the host needs no virtualenv — Docker must exist first. Docker-presence
guard, container env, full probe invocation, and the model/dataset probe scripts
are in `references/workflow-intake-preflight.md`, `references/model-discovery.md`,
and `references/dataset-sources.md`.
Probe requirements:
- Model: load `AutoConfig`, read model-card tags, detect task from
`architectures` + tags + card examples (fallback logging in `model-discovery.md`).
- Dataset: for recommended datasets, first present 3-5 choices from
`dataset-recommendations.md`; for local data, bind-mount read-only and use
`dataset-sources.md` format detection.
- Reject early if the model config fails, the task is out of scope, no recipe
source exists, or the dataset cannot load / match the task schema.
- Evaluate `compat-workarounds.md` against the model/task; defer hardware-dependent
rules to Step 2.
Write the initial `config.yaml` (`model_id`, `task`, `dataset_id` or
`local_dataset_path`, `research_sources: []` filled in Step 3,
`applicable_workarounds:` from Step 1, `notes: []` for reference fallbacks,
`push_to_hub: true` default — annotated template in
`references/workflow-intake-preflight.md`). Optionally `rm -rf "$OUTPUT_DIR/.probe"`
once the gate is met.
**Gate:** `config.yaml` exists with model, dataset, task, applicable_workarounds;
do not proceed if any field is missing.
---
### Step 2 — Hardware audit & NGC image
**Goal:** verify Docker + GPU + disk, pick the NGC PyTorch image live, finalize
hardware-dependent compat rules.
**2a. Audit (hard gate)** — three checks (commands in
`references/workflow-intake-preflight.md`):
1. GPU host runtime — `tao-setup-nvidia-gpu-host`'s
`setup-nvidia-gpu-host.sh --backend docker --check-only`; on fail, ask approval
then re-run with `--install --yes`.
2. Free-disk soft-warn — override via `MIN_DISK_GB` (default 100 GB); recommend
≥ 100 GB for NGC base (~20 GB) + HF cache + checkpoints + data.
3. Conditional credential presence (from the session environment, values never
read) — `HF_TOKEN` only when gated or `push_to_hub` is on; `WANDB_*` only when
WandB is on.
**Do not proceed to Step 4 on a hard-fail** — Step 4's `docker build` pulls a
20+ GB NGC base, and a missing `nvidia-container-toolkit` only surfaces later as
`could not select device driver "" with capabilities: [[gpu]]`. Record `gpu_count`,
`gpu_name`, `driver_major`, `vram_gb_per_gpu` in `config.yaml`.
**2b. Pick NGC image (live):** from the NVIDIA Deep Learning Frameworks support
matrix (<https://docs.nvidia.com/deeplearning/frameworks/support-matrix/index.html>),
PyTorch NGC container section, pick the highest-versioned image where
`Min driver ≤ detected driver_major` and container CUDA `≤` host CUDA Toolkit
(match closely so cuDNN / TensorRT line up). Do **not** reject an image for an
`aN`/`bN`/`rcN` PyTorch tag — NGC validates the full image; pick the newest
CUDA-aligned one and let `compat-workarounds.md` handle per-version issues. If the
matrix is unreachable, use the fallbacks in `references/hardware-container.md`;
default `nvcr.io/nvidia/pytorch:24.09-py3` (driver ≥ 545; SDPA+GQA bug — if
`num_key_value_heads < num_attention_heads`, set `attn_implementation: "eager"`).
Record `ngc_image` in `config.yaml`.
**2c. Re-evaluate hardware-dependent compat rules:** re-run the
`compat-workarounds.md` walk for entries whose `detect` needs `hw`; update
`applicable_workarounds:` in place.
**2d. Model-fit check:** estimate `param_bytes ≈ 2×param_count` (bf16); if
> 60% of `vram_gb_per_gpu × 1e9`, recommend LoRA in the user-facing summary.
**Gate:** `config.yaml` has `ngc_image`, `gpu_count`, `gpu_name`, `driver_major`,
`vram_gb_per_gpu`; hardware-dependent compat fixes recorded.
---
### Step 3 — Research the recipe
**Goal:** fetch the live recipe — training-data knowledge of
`transformers`/`trl`/`peft` is suspect, so Step 3 is non-negotiable. Walk
`references/research-priorities.md` in priority order (Priority 1 → 6); stop once
you have, for the detected task:
- `AutoModel` / processor class
- Train + eval transforms
- Collator
- `compute_metrics`
- Hyperparameter hints (LR, batch size, epochs, scheduler)
Record findings in `meta/recipe.md`, append source URLs to
`config.yaml: research_sources:`. A slot with no live finding falls back to the
matching scaffold (`cv-scripts.md` / `vlm-scripts.md`), logged as "fallback to
scaffold — no live source for <slot>" under `notes:`. Conflict-resolution rules
are in `references/research-priorities.md`.
**Gate:** every required slot filled, with a source URL or scaffold-fallback note.
---
### Step 4 — Generate project & smoke-test
**Goal:** write all scripts, build the image, prepare data, run a 1-step smoke on
real data (one `docker build`, two `docker run`s).
**4a. Generate project files** in `output_dir/`: `config.yaml`, `Dockerfile`,
`requirements.txt`, `prepare_data.py`, `train.py`, `run_eval.py`, `infer.py`,
optional `merge_lora.py`, optional `tests/`, `.gitignore`. Live Step 3 research is
authority; `cv-scripts.md` / `vlm-scripts.md` give scaffold shape only. Apply every
`applicable_workarounds` entry as a Dockerfile block, requirement pin, config
override, or runtime env var. Hard rules: `run_eval.py` keeps that exact filename
(avoids colliding with the HF `evaluate` package); every generated `.py` starts
with the NVIDIA Apache-2.0 copyright header and any emitter fails when it is
missing; `emit_unit_tests: true` generates and runs tests per
`references/testing.md`. Script bodies, Dockerfile shape, and the emitter contract
are in `references/workflow-generate-train.md`.
**4b. Build, prepare, smoke** — `docker build -t run-<short>:latest .`, then
`prepare_data` and the `--smoke --max_steps 1` run (`references/docker-runs.md`
§1-3). Smoke pass criteria (in `logs/smoke.log`):
- No exception
- Loss is finite (not `0.0`, not `NaN`)
- `grad_norm > 0` at step 1
If `emit_unit_tests: true`, also run `pytest tests/` in the container. Any failure → STOP.
**4c. Preflight summary** — before full training, print and verify: reference URL,
dataset columns, Hub target, monitoring target, NGC image, hardware, smoke loss/grad norm.
**Gate:** project files written, image built, smoke PASSED, preflight has no
blank fields.
---
### Step 5 — Train, evaluate, infer
**Goal:** baseline eval, full training, post-train eval, optional LoRA merge, 5
inference samples (all commands: `references/docker-runs.md` §4-8).
| Sub-step | docker-runs.md | Skip if |
|---|---|---|
| 5a. Baseline eval (zero-shot) | §4 | `skip_baseline: true` |
| 5b. Full training (detached) | §5 | — |
| 5c. LoRA merge | §6 | not VLM+LoRA |
| 5d. Post-train eval | §7 | — |
| 5e. Inference (5 samples) | §8 | — |
Multi-GPU: prepend `torchrun --nproc_per_node=$gpu_count` to `python train.py`.
While training streams, watch `docker logs -f hft_train`: loss should drop within
10-20 steps; flat loss (collator/label-masking bug), NaN (LR too high), and OOM
all stop the run — recovery in `references/core-rules.md`. If `emit_report: true`,
run `report.py` after Step 5e per `references/reporting.md`.
**Gate:** all of:
- `checkpoints/final/` (or `checkpoints/merged/` for LoRA) exists
- `reports/eval_results.json` has a numeric primary metric
- `reports/baseline_results.json` exists (unless skipped)
- `reports/inference_samples/` has 5 samples
- wandb URL shows descending loss
---
### Step 6 — Push & emit rerun skill
**Goal:** publish the run and make it reproducible without re-research.
Push per `references/hub-push.md` (weights, model card, eval/baseline JSONs,
`config.yaml`, `Dockerfile`, `requirements.txt`, inference samples, reports when
emitted) unless `push_to_hub: false` is explicit. Emit
`<output_dir>/skills/run-<short>/SKILL.md` from
`references/pipeline-skill-template.md` — substitute every placeholder, include
full YAML metadata + the NVIDIA copyright HTML comment, and make any emitter fail
if those are missing.
**Gate (Done criteria):** all of:
- Step 5 gate met
- HF Hub repo exists at the resolved URL with weights + card + `results/`
(unless `push_to_hub: false`)
- `<output_dir>/skills/run-<short>/SKILL.md` exists, no `<placeholder>` left,
with metadata + copyright HTML comment per `pipeline-skill-template.md`
Final message: wandb URL, HF Hub URL, baseline -> fine-tuned primary metric,
`reports/inference_samples/`, and the rerun skill path.
---
## Error playbook
On a known runtime error, consult the symptom → minimal-fix table in
`references/error-playbook.md` (NGC entrypoint, PyTorch/Transformers regressions,
numpy ABI, Albumentations bbox, PEFT/checkpointing, LoRA target breadth, CV
augmentation gaps, OOM at step 0) before redesigning anything. When a row there
fires twice across runs, lift it into `compat-workarounds.md` with a `detect` rule
— auto-applied in Step 1 before the error can fire.
---
## Communication style
- Terse. No filler, no restating the request; one-word answers when appropriate.
- Always include direct Hub and wandb URLs when referencing artifacts.
- On error: state what went wrong, why, what you changed — no menus.
- Never present "Option A/B/C" for a request with a clear answer. Act.
## Example pipelines
- [tao-rerun-convnext-cifar10](references/tao-rerun-convnext-cifar10.md)
- [tao-rerun-detr-cppe5](references/tao-rerun-detr-cppe5.md)
- [tao-rerun-segformer-foodseg103](references/tao-rerun-segformer-foodseg103.md)
- [tao-rerun-smolvlm-vqav2](references/tao-rerun-smolvlm-vqav2.md)





首页
