옵션
집집 Skill 데이터 과학 및 ML tao-finetune-huggingface-model

tao-finetune-huggingface-model

NVIDIA/skills NVIDIA/skills

NGC PyTorch 컨테이너를 사용하여 로컬 NVIDIA GPU에서 HuggingFace의 CV, VLM 또는 LLM 모델을 파인튜닝하세요. 전체 학습 또는 LoRA 학습을 지원하며, 데이터셋 처리와 허브로의 모델 푸시(선택 사항)도 가능합니다.

...모든 것을 확장하십시오
1
업데이트 된 시간 2026년 9월 29일

tao-finetune-huggingface-model

HuggingFace 모델을 위한 로컬 NVIDIA GPU 파인튜닝. 실시간으로 가져온 문서를 기반으로 하며, 대체 안전망으로 선별된 참조 자료를 포함합니다. 하나의 NGC 컨테이너, 몇 개의 집중된 스크립트, HF Hub로의 한 번의 푸시. 이 파일의 규칙을 따르십시오. 즉흥적으로 행동하지 마십시오.

권위 순서 (최고 우선순위부터):

  1. 사용자 입력 — 명시적인 model_id, dataset_id, training_method, config.yaml 오버라이드.
  2. 실시간 연구 — 모델 카드, HF 레포지토리 예시, 저자 파인튜닝 스크립트, HF 작업 문서, 논문; 항상 가져옴 (Step 3 + references/research-priorities.md).
  3. 선별된 참조 (references/*.md) — 실시간 연구가 침묵하거나 모호할 때의 대체재.
  4. 당신의 학습 데이터 메모리 — 최후의 수단; 의심스럽고 (2)/(3)에 대해 교차 확인해야 함.

(2)와 (3) 간의 충돌 해결 및 소스 라인 불일치 참고사항은 references/research-priorities.md에 있습니다.

입력값

필수:

  • model_id — HuggingFace 모델 ID, 예: google/vit-base-patch16-224

조건부 자격 증명 (세션 환경에서 읽음, 실행 전 내보낼 때):

  • HF_TOKEN — 모델/데이터셋이 가이트된 경우(읽기) 또는 push_to_hub가 켜진 경우(쓰기)에만 필요; 공개 + 공개 + push_to_hub: false는 필요 없음. 값은 절대 읽히지 않음 — [ -n "$HF_TOKEN" ]을 통한 존재 여부만 확인.
  • WANDB_API_KEY, WANDB_PROJECT — WandB가 활성화된 경우에만 필요; WANDB_MODE=disabled은 opting out(선택 해제).

데이터셋 — 정확히 하나:

  • dataset_id — HuggingFace 데이터셋 ID (소스: hf)
  • local_dataset_path — 로컬 폴더 또는 파일 (소스: local); 선택적 local_dataset_format ∈ {auto, imagefolder, coco, voc, jsonl, arrow, parquet, csv} (기본값: 자동 감지).
  • (생략) — 에이전트가 인기 데이터셋을 추천 (소스: recommend)

선택적 (기본값 있음):

  • task_type — 구성 + 모델 카드에서 자동 감지
  • n_train=10000, n_eval=1000, n_epochs=3, lora_r=16
  • output_dir=./output/<model_short_name></model_short_name>
  • hf_model_repo — 푸시 대상; 설정되지 않고 HF_TOKEN에 쓰기 권한이 있으면 <whoami>/<model_short_name>-finetuned</model_short_name></whoami>로 자동 유도.
  • push_to_hub=True — False로 설정하면 건너뜀
  • skip_baseline=False — 제로샷 베이스라인 평가 건너뛰기

선택적 결과물 (기본값 꺼짐):

emit_progress_log: false   # output_dir/PROGRESS.md (단계별 저널)
emit_report:       false   # 곡선 및 샘플이 포함된 reports/report.{pdf,html}
emit_unit_tests:   false   # 가짜 데이터 이질적 배치 테스트가 포함된 tests/

모든 값은 output_dir/config.yaml에 있습니다. Python에서 하드코딩하지 마십시오.

실행 플랫폼

이 스킬은 무엇을 실행할지 조정합니다; 플랫폼 스킬은 GPU 호스트에서 어떻게 실행할지 소유합니다 — 먼저 이를 읽으십시오.

관심사권위 있는 스킬
GPU 호스트 런타임 (드라이버 580, CUDA Toolkit 13.0, NVIDIA Container Toolkit 1.19.0)`tao-skill-bank:tao-setup-nvidia-gpu-host`
`docker run` 플래그, NGC 인증, 마운트, 환경 전달`tao-skill-bank:tao-run-on-docker`
로컬 Docker 작업 사전 검사 (데몬, GPU 연기 테스트)`tao-skill-bank:tao-run-on-local-docker`

기본 플랫폼: local-docker — 일회용 이미지(run-<short>:latest</short>)를 빌드하고 로컬 Docker 데몬에서 실행합니다. 사용자가 명시적으로 다른 백엔드(Brev 원격 GPU, SLURM/Kubernetes)가 필요할 때만 요청하십시오; 그런 다음 해당 플랫폼의 사전 검사를 먼저 실행하고 Steps 4–5의 docker run 명령을 통해 라우팅하십시오. GPU 런타임 및 존재 여부 자격 증명 사전 검사(값은 절대 읽히지 않음), 표준 docker run 플래그 세트, list_tao_platforms.py 선택 명령 및 워크플로우별 플래그(--entrypoint /bin/bash -lc, PYTORCH_CUDA_ALLOC_CONF, --name hft_train)는 references/workflow-intake-preflight.md에 있습니다.

참조 — 대체 안전망

실시간 연구가 침묵하거나, 모호하거나, 사용할 수 없을 때만 참조됩니다; 특정 모델과 현재 API에 대해서는 실시간 문서가 항상 우선합니다. 각 단계는 필요한 참조를 연결합니다; 전체 카탈로그는 references/detailed-workflow.md에 있습니다.

상시 활성화: core-rules.md, error-playbook.md, compat-workarounds.md, model-discovery.md, dataset-recommendations.md, dataset-sources.md, dataset-patterns.md, hardware-container.md, research-priorities.md, cv-scripts.md, vlm-scripts.md, docker-runs.md, hub-push.md, pipeline-skill-template.md, deliverables.md. 옵트인 (플래그/필요가 적용될 때): progress-tracking.md, testing.md, reporting.md, workflow-intake-preflight.md, workflow-generate-train.md, workflow-push-rerun.md.

규칙: 대체하기 전에 시도한 실시간 소스와 그 불충분한 이유를 기록하십시오 (config.yaml notes:, 및 PROGRESS.md가 활성화된 경우). cv-scripts.md / vlm-scripts.md의 [FETCH LIVE] 마커는 인라인할 코드가 아닌 연구 체크리스트입니다 — 블록에 Step 3 발견 사항이 없는 경우 나열된 URL을 다시 가져오십시오.

핵심 규칙

타협 불가능한 동작. 짧은 버전 (전체 열거 — hallucinated-imports 목록, 승인 없이 절대 사용하지 않는 목록, 전체 오류 복구 및 하드웨어 크기 표 — references/core-rules.md에 있음, 학습 시간 결정 전에 참조):

  • 당신의 HF-library 지식은 낡았습니다. ML 코드를 작성하기 전에 실시간 문서(모델 카드, HF 레포지토리 예시, 작업 문서)를 가져오십시오 — 메모리에서 트레이너 인수 / 콜레이터 / 변환을 생성하지 마십시오 (Step 3).
  • 전체 실행 전에 --max_steps 1로 실제 데이터에서 연기 테스트하십시오 — 검증된 연기 테스트 없이는 배치 실행하지 마십시오.
  • 모델_id, dataset_id, 또는 training_method을 묵시적으로 대체하지 마십시오 — 사용자가 요청한 것이 로드되지 않으면 멈추고 물어보십시오.
  • 오류 복구는 최소 변경입니다. OOM → 배치 크기 절반, grad_accum 두 배, 그래디언트 체크포인트 활성화 (승인 없이 LoRA 전환 없음); NaN → LR 10× 감소; 평탄한 손실 → 콜레이터 검사; 동일한 오류 3× → 멈추고 물어보십시오. 루프를 돌지 마십시오.
  • 데이터셋 열은 콜레이터 전에 검증 — prepare_data.py에서 이름 변경; 재구성이 필요하면 → 멈추고 물어보십시오.
  • 하드웨어 크기 추정 (bf16): ≤3B → 24 GB, 7–13B → 80 GB, 30B+ → 1× 80 GB에서 다중 GPU 또는 LoRA, 70B+ → 8× 80 GB 또는 LoRA. 전체 파인튜닝이 맞지 않고 LoRA 요청 없음 → 전환 전에 물어보십시오.

워크플로우 — 6단계

단일 패스, 순차적; 각 단계는 다음이 시작되기 전에 명확한 게이트가 있습니다.

Step 1 — 검사 및 자격 부여

목표: 계속할지 결정. 모델 + 데이터셋 탐색, 수락/거부 적용, 적용 가능한 호환성 수정 등록, 초기 config.yaml 작성.

전제 조건: MODEL_ID, 선택적 DATASET_ID / local_dataset_path, 선택적 HF_TOKEN, OUTPUT_DIR (기본값 ./output/<model_short_name></model_short_name>). 탐사는 CPU 전용 python:3.12-slim Docker 컨테이너(바인드 마운트된 .probe/스크래치)에서 실행되므로 호스트에 virtualenv가 필요하지 않음 — Docker가 먼저 존재해야 함. Docker 존재 감시자, 컨테이너 환경, 전체 탐사 호출 및 모델/데이터셋 탐사 스크립트는 references/workflow-intake-preflight.md, references/model-discovery.md, references/dataset-sources.md에 있습니다.

탐사 요구사항:

  • 모델: AutoConfig 로드, 모델 카드 태그 읽기, architectures + 태그 + 카드 예시에서 작업 감지 (model-discovery.md에 로깅 대체).
  • 데이터셋: 추천 데이터셋의 경우, 먼저 dataset-recommendations.md에서 3-5개 선택지 제시; 로컬 데이터의 경우 읽기 전용으로 바인드 마운트하고 dataset-sources.md 형식 감지 사용.
  • 모델 구성 실패, 작업이 범위를 벗어남, 레시피 소스가 없음, 또는 데이터셋을 로드/작업 스키마와 일치시킬 수 없는 경우 조기에 거부.
  • 모델/작업에 대해 compat-workarounds.md 평가; 하드웨어 의존적 규칙은 Step 2로 연기.

초기 config.yaml 작성 (model_id, task, dataset_id 또는 local_dataset_path, Step 3에서 채워진 research_sources: [], Step 1에서 applicable_workarounds:, 참조 대체용 notes: [], 기본값 push_to_hub: true — 주석 템플릿은 references/workflow-intake-preflight.md에 있음). 게이트가 충족되면 선택적으로 rm -rf "$OUTPUT_DIR/.probe" 실행.

게이트: 모델, 데이터셋, 작업, applicable_workarounds이 있는 config.yaml 존재; 필드 중 하나가 누락되면 진행하지 마십시오.

Step 2 — 하드웨어 감사 및 NGC 이미지

목표: Docker + GPU + 디스크 확인, NGC PyTorch 이미지를 실시간으로 선택, 하드웨어 의존적 호환성 규칙 최종화.

2a. 감사 (경고 게이트) — 세 가지 확인 (references/workflow-intake-preflight.md의 명령):

  1. GPU 호스트 런타임 — tao-setup-nvidia-gpu-host의 setup-nvidia-gpu-host.sh --backend docker --check-only; 실패 시 승인 요청 후 --install --yes로 다시 실행.
  2. 여유 디스크 경고 — MIN_DISK_GB(기본값 100 GB)로 재정의; NGC 기본(~20 GB) + HF 캐시 + 체크포인트 + 데이터에 대해 ≥ 100 GB 권장.
  3. 조건부 자격 증명 존재 (세션 환경에서, 값은 절대 읽히지 않음) — HF_TOKEN은 가이트된 경우 또는 push_to_hub가 켜진 경우에만; WANDB_*는 WandB가 켜진 경우에만.

경고 실패 시 Step 4로 진행하지 마십시오 — Step 4의 docker build는 20+ GB NGC 기본을 가져오며, 누락된 nvidia-container-toolkit은 나중에 could not select device driver "" with capabilities: [[gpu]]로 나타납니다. config.yaml에 gpu_count, gpu_name, driver_major, vram_gb_per_gpu 기록.

2b. NGC 이미지 선택 (실시간): NVIDIA 딥러닝 프레임워크 지원 행렬(https://docs.nvidia.com/deeplearning/frameworks/support-matrix/index.html)의 PyTorch NGC 컨테이너 섹션에서, Min driver ≤ 감지된 driver_major 및 컨테이너 CUDA ≤ 호스트 CUDA Toolkit인 가장 높은 버전의 이미지를 선택 (cuDNN / TensorRT 정렬을 위해 가깝게 일치). aN/bN/rcN PyTorch 태그로 인해 이미지를 거부하지 마십시오 — NGC는 전체 이미지를 검증합니다; CUDA와 일치하는 최신 버전을 선택하고 compat-workarounds.md가 버전별 문제를 처리하도록 하십시오. 행렬에 도달할 수 없는 경우 references/hardware-container.md의 대체 사용; 기본값 nvcr.io/nvidia/pytorch:24.09-py3 (드라이버 ≥ 545; SDPA+GQA 버그 — num_key_value_heads , setattn_implementation: "eager").config.yaml에ngc_image` 기록.

2c. 하드웨어 의존적 호환성 규칙 재평가: detect에 hw가 필요한 항목에 대해 compat-workarounds.md 워크를 다시 실행; applicable_workarounds:을 인플레이스 업데이트.

2d. 모델 적합성 확인: param_bytes ≈ 2×param_count(bf16) 추정;

vram_gb_per_gpu × 1e9의 60%를 초과하면, 사용자 대상 요약에서 LoRA 권장.

게이트: config.yaml에 ngc_image, gpu_count, gpu_name, driver_major, vram_gb_per_gpu 존재; 하드웨어 의존적 호환성 수정 기록.

Step 3 — 레시피 연구

목표: 실시간 레시피 가져오기 — transformers/trl/peft의 학습 데이터 지식은 의심스럽으므로 Step 3은 타협 불가능. 우선순위 순서(references/research-priorities.md)로 워크 (Priority 1 → 6); 감지된 작업에 대해 다음을 얻으면 중지:

  • AutoModel / 프로세서 클래스
  • 학습 + 평가 변환
  • 콜레이터
  • compute_metrics
  • 하이퍼파라미터 힌트 (LR, 배치 크기, 에포크, 스케줄러)

meta/recipe.md에 발견 사항 기록, config.yaml: research_sources:에 소스 URL 추가. 실시간 발견이 없는 슬롯은 일치하는 스캐폴드(cv-scripts.md / vlm-scripts.md)로 대체되며, notes: 아래 "fallback to scaffold — no live source for "로 로깅됨. 충돌 해결 규칙은 references/research-priorities.md에 있습니다.

게이트: 모든 필수 슬롯 채움, 소스 URL 또는 스캐폴드 대체 노트 포함.

Step 4 — 프로젝트 생성 및 연기 테스트

목표: 모든 스크립트 작성, 이미지 빌드, 데이터 준비, 실제 데이터에서 1단계 연기 실행 (하나의 docker build, 두 개의 docker run).

4a. 프로젝트 파일 생성 output_dir/에서: config.yaml, Dockerfile, requirements.txt, prepare_data.py, train.py, run_eval.py, infer.py, 선택적 merge_lora.py, 선택적 tests/, .gitignore. 실시간 Step 3 연구가 권위; cv-scripts.md / vlm-scripts.md는 스캐폴드 형태만 제공. 모든 applicable_workarounds 항목을 Dockerfile 블록, 요구사항 고정, 구성 오버라이드 또는 런타임 환경 변수로 적용. 하드 규칙: run_eval.py는 해당 정확한 파일명을 유지(HF evaluate 패키지와 충돌 방지); 생성된 모든 .py는 NVIDIA Apache-2.0 저작권 헤더로 시작하며 누락 시 모든 에미터 실패; emit_unit_tests: true는 references/testing.md에 따라 테스트 생성 및 실행. 스크립트 본문, Dockerfile 형태 및 에미터 계약은 references/workflow-generate-train.md에 있습니다.

4b. 빌드, 준비, 연기 — docker build -t run-<short>:latest .</short>, 그런 다음 prepare_data 및 --smoke --max_steps 1 실행 (references/docker-runs.md§1-3). 연기 통과 기준 (logs/smoke.log에서):

  • 예외 없음
  • 손실이 유한함 (0.0 아님, NaN 아님)
  • step 1에서 grad_norm > 0

emit_unit_tests: true인 경우, 컨테이너에서 pytest tests/도 실행. 실패 시 → STOP.

4c. 사전 검사 요약 — 전체 학습 전에 인쇄 및 확인: 참조 URL, 데이터셋 열, Hub 대상, 모니터링 대상, NGC 이미지, 하드웨어, 연기 손실/그래디언트_norm.

게이트: 프로젝트 파일 작성됨, 이미지 빌드됨, 연기 PASSED, 사전 검사에 빈 필드 없음.

Step 5 — 학습, 평가, 추론

목표: 베이스라인 평가, 전체 학습, 학습 후 평가, 선택적 LoRA 병합, 5개 추론 샘플 (모든 명령: references/docker-runs.md §4-8).

하위 단계docker-runs.md다음 경우 건너뛰기
5a. 베이스라인 평가 (제로샷)§4`skip_baseline: true`
5b. 전체 학습 (분리)§5—
5c. LoRA 병합§6VLM+LoRA 아님
5d. 학습 후 평가§7—
5e. 추론 (5 샘플)§8—

다중 GPU: python train.py 앞에 torchrun --nproc_per_node=$gpu_count 추가.

학습 스트리밍 동안 docker logs -f hft_train 모니터링: 손실은 10-20 단계 내에 감소해야 함; 평탄한 손실 (콜레이터/레이블 마스킹 버그), NaN (LR 너무 높음), OOM은 모두 실행 중지 — 복구 방법은 references/core-rules.md에 있음. emit_report: true인 경우, Step 5e 후 references/reporting.md에 따라 report.py 실행.

게이트: 다음 모두:

  • checkpoints/final/ (또는 LoRA의 경우 checkpoints/merged/) 존재
  • reports/eval_results.json에 숫자 주요 지표 있음
  • reports/baseline_results.json 존재 (건너뛴 경우 제외)
  • reports/inference_samples/에 5개 샘플 존재
  • wandb URL이 감소하는 손실 표시

Step 6 — 푸시 및 재실행 스킬 생성

목표: 실행 게시 및 재연구 없이 재현 가능하게 만들기.

명시적 push_to_hub: false가 아닌 한 references/hub-push.md에 따라 푸시 (가중치, 모델 카드, eval/baseline JSON, config.yaml, Dockerfile, requirements.txt, 추론 샘플, 생성된 경우 보고서). references/pipeline-skill-template.md에서 <output_dir>/skills/run-<short>/SKILL.md</short></output_dir> 생성 — 모든 자리 표시자 대체, 전체 YAML 메타데이터 + NVIDIA 저작권 HTML 주석 포함, 누락 시 모든 에미터 실패.

게이트 (완료 기준): 다음 모두:

  • Step 5 게이트 충족
  • HF Hub 레포지토리가 해결된 URL에 가중치 + 카드 + results/(또는 push_to_hub: false) 존재
  • <output_dir>/skills/run-<short>/SKILL.md</short></output_dir> 존재, <placeholder></placeholder> 남지 않음, pipeline-skill-template.md에 따른 메타데이터 + 저작권 HTML 주석 포함

최종 메시지: wandb URL, HF Hub URL, 베이스라인 -> 파인튜닝 주요 지표, reports/inference_samples/, 재실행 스킬 경로.

오류 플레이북

알려진 런타임 오류 시, 재설계하기 전에 references/error-playbook.md(NGC 진입점, PyTorch/Transformers 회귀, numpy ABI, Albumentations bbox, PEFT/체크포인트, LoRA 대상 너비, CV 증강 격차, step 0의 OOM)의 증상 → 최소 수정 표를 참조하십시오. 해당 행이 실행 전체에 걸쳐 두 번 발동하면, 오류가 발동하기 전에 Step 1에서 자동 적용되도록 detect 규칙과 함께 compat-workarounds.md로 승격하십시오.

커뮤니케이션 스타일

  • 간결함. 여백 없음, 요청 재진술 없음; 적절한 경우 한 단어 답변.
  • 아티팩트를 참조할 때 항상 직접 Hub 및 wandb URL 포함.
  • 오류 시: 무엇이 잘못되었는지, 왜, 무엇을 변경했는지 명시 — 메뉴 없음.
  • 명확한 답변이 있는 요청에 대해 "Option A/B/C" 제시하지 않음. 행동하십시오.

예시 파이프라인

  • tao-rerun-convnext-cifar10
  • tao-rerun-detr-cppe5
  • tao-rerun-segformer-foodseg103
  • tao-rerun-smolvlm-vqav2
GitHub에서 보기
---
name: tao-finetune-huggingface-model
description: Fine-tune HuggingFace CV, VLM, or LLM models on local NVIDIA GPUs using an NGC PyTorch container, with support for full or LoRA training, dataset handling, and optional model push to the Hub.
license: Apache-2.0
---
<!-- Copyright (c) 2026, NVIDIA CORPORATION. All rights reserved. Licensed under the Apache License, Version 2.0; see http://www.apache.org/licenses/LICENSE-2.0 -->

# tao-finetune-huggingface-model

Local NVIDIA GPU fine-tuning for HuggingFace models, grounded in live-fetched
documentation with curated references as a fallback safety net. One NGC container,
a few focused scripts, one push to HF Hub. Follow the rules in this file; don't
improvise.

**Order of authority (highest first):**

1. **User input** — explicit `model_id`, `dataset_id`, `training_method`, `config.yaml` overrides.
2. **Live research** — model card, HF repo example, author finetune script, HF task docs, paper; always fetched (Step 3 + `references/research-priorities.md`).
3. **Curated references** (`references/*.md`) — fallback when live research is silent/ambiguous.
4. **Your training-data memory** — last resort; suspect, cross-check against (2)/(3).

Conflict resolution between (2) and (3) and the source-line discrepancy note are
in `references/research-priorities.md`.

---

## Inputs

**Required:**
- `model_id` — HuggingFace model ID, e.g. `google/vit-base-patch16-224`

**Conditional credentials (read from the session environment, exported before launching when present):**
- `HF_TOKEN` — only when the model/dataset is **gated** (read) or `push_to_hub` is on (write); public + public + `push_to_hub: false` needs none. Value never read — presence-only via `[ -n "$HF_TOKEN" ]`.
- `WANDB_API_KEY`, `WANDB_PROJECT` — only when WandB is enabled; `WANDB_MODE=disabled` opts out.

**Dataset — exactly one:**
- `dataset_id` — HuggingFace dataset ID *(source: `hf`)*
- `local_dataset_path` — local folder or file *(source: `local`)*; optional
  `local_dataset_format` ∈ {auto, imagefolder, coco, voc, jsonl, arrow, parquet,
  csv} (default: auto-detect).
- *(omit)* — agent recommends popular datasets *(source: `recommend`)*

**Optional (have defaults):**
- `task_type` — auto-detected from config + model card
- `n_train=10000`, `n_eval=1000`, `n_epochs=3`, `lora_r=16`
- `output_dir=./output/<model_short_name>`
- `hf_model_repo` — push target; if unset and HF_TOKEN has write access,
  auto-derived as `<whoami>/<model_short_name>-finetuned`.
- `push_to_hub=True` — set to `False` to skip
- `skip_baseline=False` — skip zero-shot baseline eval

**Optional deliverables (off by default):**
```yaml
emit_progress_log: false   # output_dir/PROGRESS.md (per-step journal)
emit_report:       false   # reports/report.{pdf,html} with curves & samples
emit_unit_tests:   false   # tests/ with fake-data heterogeneous-batch tests
```

All values live in `output_dir/config.yaml`. Never hardcode in Python.

---

## Execution platform

This skill orchestrates *what* to run; the platform skills own *how* to run it on
a GPU host — read them first.

| Concern | Authoritative skill |
|---|---|
| GPU host runtime (driver 580, CUDA Toolkit 13.0, NVIDIA Container Toolkit 1.19.0) | [`tao-skill-bank:tao-setup-nvidia-gpu-host`](../../platform/tao-setup-nvidia-gpu-host/SKILL.md) |
| `docker run` flags, NGC auth, mounts, env passthrough | [`tao-skill-bank:tao-run-on-docker`](../../platform/tao-run-on-docker/SKILL.md) |
| Local Docker job preflight (daemon, GPU smoke) | [`tao-skill-bank:tao-run-on-local-docker`](../../platform/tao-run-on-local-docker/SKILL.md) |

**Default platform:** `local-docker` — build a one-off image (`run-<short>:latest`)
and run it on the local Docker daemon. Ask only when the user explicitly needs a
different backend (Brev remote GPU, SLURM/Kubernetes); then run that platform's
Preflight first and route the Steps 4–5 `docker run` commands through it. The
GPU-runtime and presence-only credential preflights (values never read), the
canonical `docker run` flag set, the `list_tao_platforms.py` selection command, and
the workflow-specific flags (`--entrypoint /bin/bash -lc`, `PYTORCH_CUDA_ALLOC_CONF`,
`--name hft_train`) are in `references/workflow-intake-preflight.md`.

---

## References — fallback safety net

Consulted **only** when live research is silent, ambiguous, or unavailable; live
docs always win for the specific model and current API. Each step links the
references it needs; full catalog in `references/detailed-workflow.md`.

Always-on: `core-rules.md`, `error-playbook.md`, `compat-workarounds.md`,
`model-discovery.md`, `dataset-recommendations.md`, `dataset-sources.md`,
`dataset-patterns.md`, `hardware-container.md`, `research-priorities.md`,
`cv-scripts.md`, `vlm-scripts.md`, `docker-runs.md`, `hub-push.md`,
`pipeline-skill-template.md`, `deliverables.md`. Opt-in (when their flag/need
applies): `progress-tracking.md`, `testing.md`, `reporting.md`,
`workflow-intake-preflight.md`, `workflow-generate-train.md`, `workflow-push-rerun.md`.

**Rule:** before falling back, log the live source you tried and why it was
insufficient (`config.yaml` `notes:`, and PROGRESS.md if enabled). `[FETCH LIVE]`
markers in `cv-scripts.md` / `vlm-scripts.md` are a research checklist, not code to
inline — refetch the listed URL if a block has no Step 3 finding.

---

## Core rules

Non-negotiable behaviors. **Short version** (full enumeration —
hallucinated-imports list, never-without-approval list, full error-recovery and
hardware-sizing tables — in `references/core-rules.md`, consult before any
training-time decision):

- **Your HF-library knowledge is outdated.** Fetch live docs (model card, HF
  repo example, task doc) before writing any ML code — don't generate trainer
  args / collator / transforms from memory (Step 3).
- **Smoke-test on real data with `--max_steps 1`** before any full run; no batch
  launches without a verified smoke.
- **Never silently substitute** model_id, dataset_id, or training_method — if
  what the user asked for doesn't load, stop and ask.
- **Error recovery is minimal-change.** OOM → halve batch, double grad_accum,
  enable gradient checkpointing (no LoRA switch without approval); NaN → reduce
  LR 10×; flat loss → inspect collator; same error 3× → stop and ask. Don't loop.
- **Dataset columns verified BEFORE the collator** — rename in `prepare_data.py`;
  restructuring needed → stop and ask.
- **Hardware-sizing thumb (bf16):** ≤3B → 24 GB, 7–13B → 80 GB, 30B+ → multi-GPU
  or LoRA on 1× 80 GB, 70B+ → 8× 80 GB or LoRA. Full finetune won't fit and no
  LoRA requested → ask before switching.

---

## Workflow — 6 steps

Single pass, sequential; each step has a clear gate before the next begins.

### Step 1 — Inspect & qualify

**Goal:** decide whether to proceed. Probe model + dataset, apply accept/reject,
register applicable compat fixes, write the initial `config.yaml`.

Prerequisites: `MODEL_ID`, optional `DATASET_ID` / `local_dataset_path`,
optional `HF_TOKEN`, `OUTPUT_DIR` (default `./output/<model_short_name>`). Probes
run in a CPU-only `python:3.12-slim` Docker container (bind-mounted `.probe/`
scratch) so the host needs no virtualenv — Docker must exist first. Docker-presence
guard, container env, full probe invocation, and the model/dataset probe scripts
are in `references/workflow-intake-preflight.md`, `references/model-discovery.md`,
and `references/dataset-sources.md`.

Probe requirements:

- Model: load `AutoConfig`, read model-card tags, detect task from
  `architectures` + tags + card examples (fallback logging in `model-discovery.md`).
- Dataset: for recommended datasets, first present 3-5 choices from
  `dataset-recommendations.md`; for local data, bind-mount read-only and use
  `dataset-sources.md` format detection.
- Reject early if the model config fails, the task is out of scope, no recipe
  source exists, or the dataset cannot load / match the task schema.
- Evaluate `compat-workarounds.md` against the model/task; defer hardware-dependent
  rules to Step 2.

Write the initial `config.yaml` (`model_id`, `task`, `dataset_id` or
`local_dataset_path`, `research_sources: []` filled in Step 3,
`applicable_workarounds:` from Step 1, `notes: []` for reference fallbacks,
`push_to_hub: true` default — annotated template in
`references/workflow-intake-preflight.md`). Optionally `rm -rf "$OUTPUT_DIR/.probe"`
once the gate is met.

**Gate:** `config.yaml` exists with model, dataset, task, applicable_workarounds;
do not proceed if any field is missing.

---

### Step 2 — Hardware audit & NGC image

**Goal:** verify Docker + GPU + disk, pick the NGC PyTorch image live, finalize
hardware-dependent compat rules.

**2a. Audit (hard gate)** — three checks (commands in
`references/workflow-intake-preflight.md`):
1. GPU host runtime — `tao-setup-nvidia-gpu-host`'s
   `setup-nvidia-gpu-host.sh --backend docker --check-only`; on fail, ask approval
   then re-run with `--install --yes`.
2. Free-disk soft-warn — override via `MIN_DISK_GB` (default 100 GB); recommend
   ≥ 100 GB for NGC base (~20 GB) + HF cache + checkpoints + data.
3. Conditional credential presence (from the session environment, values never
   read) — `HF_TOKEN` only when gated or `push_to_hub` is on; `WANDB_*` only when
   WandB is on.

**Do not proceed to Step 4 on a hard-fail** — Step 4's `docker build` pulls a
20+ GB NGC base, and a missing `nvidia-container-toolkit` only surfaces later as
`could not select device driver "" with capabilities: [[gpu]]`. Record `gpu_count`,
`gpu_name`, `driver_major`, `vram_gb_per_gpu` in `config.yaml`.

**2b. Pick NGC image (live):** from the NVIDIA Deep Learning Frameworks support
matrix (<https://docs.nvidia.com/deeplearning/frameworks/support-matrix/index.html>),
PyTorch NGC container section, pick the highest-versioned image where
`Min driver ≤ detected driver_major` and container CUDA `≤` host CUDA Toolkit
(match closely so cuDNN / TensorRT line up). Do **not** reject an image for an
`aN`/`bN`/`rcN` PyTorch tag — NGC validates the full image; pick the newest
CUDA-aligned one and let `compat-workarounds.md` handle per-version issues. If the
matrix is unreachable, use the fallbacks in `references/hardware-container.md`;
default `nvcr.io/nvidia/pytorch:24.09-py3` (driver ≥ 545; SDPA+GQA bug — if
`num_key_value_heads < num_attention_heads`, set `attn_implementation: "eager"`).
Record `ngc_image` in `config.yaml`.

**2c. Re-evaluate hardware-dependent compat rules:** re-run the
`compat-workarounds.md` walk for entries whose `detect` needs `hw`; update
`applicable_workarounds:` in place.

**2d. Model-fit check:** estimate `param_bytes ≈ 2×param_count` (bf16); if
> 60% of `vram_gb_per_gpu × 1e9`, recommend LoRA in the user-facing summary.

**Gate:** `config.yaml` has `ngc_image`, `gpu_count`, `gpu_name`, `driver_major`,
`vram_gb_per_gpu`; hardware-dependent compat fixes recorded.

---

### Step 3 — Research the recipe

**Goal:** fetch the live recipe — training-data knowledge of
`transformers`/`trl`/`peft` is suspect, so Step 3 is non-negotiable. Walk
`references/research-priorities.md` in priority order (Priority 1 → 6); stop once
you have, for the detected task:

- `AutoModel` / processor class
- Train + eval transforms
- Collator
- `compute_metrics`
- Hyperparameter hints (LR, batch size, epochs, scheduler)

Record findings in `meta/recipe.md`, append source URLs to
`config.yaml: research_sources:`. A slot with no live finding falls back to the
matching scaffold (`cv-scripts.md` / `vlm-scripts.md`), logged as "fallback to
scaffold — no live source for <slot>" under `notes:`. Conflict-resolution rules
are in `references/research-priorities.md`.

**Gate:** every required slot filled, with a source URL or scaffold-fallback note.

---

### Step 4 — Generate project & smoke-test

**Goal:** write all scripts, build the image, prepare data, run a 1-step smoke on
real data (one `docker build`, two `docker run`s).

**4a. Generate project files** in `output_dir/`: `config.yaml`, `Dockerfile`,
`requirements.txt`, `prepare_data.py`, `train.py`, `run_eval.py`, `infer.py`,
optional `merge_lora.py`, optional `tests/`, `.gitignore`. Live Step 3 research is
authority; `cv-scripts.md` / `vlm-scripts.md` give scaffold shape only. Apply every
`applicable_workarounds` entry as a Dockerfile block, requirement pin, config
override, or runtime env var. Hard rules: `run_eval.py` keeps that exact filename
(avoids colliding with the HF `evaluate` package); every generated `.py` starts
with the NVIDIA Apache-2.0 copyright header and any emitter fails when it is
missing; `emit_unit_tests: true` generates and runs tests per
`references/testing.md`. Script bodies, Dockerfile shape, and the emitter contract
are in `references/workflow-generate-train.md`.

**4b. Build, prepare, smoke** — `docker build -t run-<short>:latest .`, then
`prepare_data` and the `--smoke --max_steps 1` run (`references/docker-runs.md`
§1-3). Smoke pass criteria (in `logs/smoke.log`):
- No exception
- Loss is finite (not `0.0`, not `NaN`)
- `grad_norm > 0` at step 1

If `emit_unit_tests: true`, also run `pytest tests/` in the container. Any failure → STOP.

**4c. Preflight summary** — before full training, print and verify: reference URL,
dataset columns, Hub target, monitoring target, NGC image, hardware, smoke loss/grad norm.

**Gate:** project files written, image built, smoke PASSED, preflight has no
blank fields.

---

### Step 5 — Train, evaluate, infer

**Goal:** baseline eval, full training, post-train eval, optional LoRA merge, 5
inference samples (all commands: `references/docker-runs.md` §4-8).

| Sub-step | docker-runs.md | Skip if |
|---|---|---|
| 5a. Baseline eval (zero-shot) | §4 | `skip_baseline: true` |
| 5b. Full training (detached) | §5 | — |
| 5c. LoRA merge | §6 | not VLM+LoRA |
| 5d. Post-train eval | §7 | — |
| 5e. Inference (5 samples) | §8 | — |

Multi-GPU: prepend `torchrun --nproc_per_node=$gpu_count` to `python train.py`.

While training streams, watch `docker logs -f hft_train`: loss should drop within
10-20 steps; flat loss (collator/label-masking bug), NaN (LR too high), and OOM
all stop the run — recovery in `references/core-rules.md`. If `emit_report: true`,
run `report.py` after Step 5e per `references/reporting.md`.

**Gate:** all of:
- `checkpoints/final/` (or `checkpoints/merged/` for LoRA) exists
- `reports/eval_results.json` has a numeric primary metric
- `reports/baseline_results.json` exists (unless skipped)
- `reports/inference_samples/` has 5 samples
- wandb URL shows descending loss

---

### Step 6 — Push & emit rerun skill

**Goal:** publish the run and make it reproducible without re-research.

Push per `references/hub-push.md` (weights, model card, eval/baseline JSONs,
`config.yaml`, `Dockerfile`, `requirements.txt`, inference samples, reports when
emitted) unless `push_to_hub: false` is explicit. Emit
`<output_dir>/skills/run-<short>/SKILL.md` from
`references/pipeline-skill-template.md` — substitute every placeholder, include
full YAML metadata + the NVIDIA copyright HTML comment, and make any emitter fail
if those are missing.

**Gate (Done criteria):** all of:
- Step 5 gate met
- HF Hub repo exists at the resolved URL with weights + card + `results/`
  (unless `push_to_hub: false`)
- `<output_dir>/skills/run-<short>/SKILL.md` exists, no `<placeholder>` left,
  with metadata + copyright HTML comment per `pipeline-skill-template.md`

Final message: wandb URL, HF Hub URL, baseline -> fine-tuned primary metric,
`reports/inference_samples/`, and the rerun skill path.

---

## Error playbook

On a known runtime error, consult the symptom → minimal-fix table in
`references/error-playbook.md` (NGC entrypoint, PyTorch/Transformers regressions,
numpy ABI, Albumentations bbox, PEFT/checkpointing, LoRA target breadth, CV
augmentation gaps, OOM at step 0) before redesigning anything. When a row there
fires twice across runs, lift it into `compat-workarounds.md` with a `detect` rule
— auto-applied in Step 1 before the error can fire.

---

## Communication style

- Terse. No filler, no restating the request; one-word answers when appropriate.
- Always include direct Hub and wandb URLs when referencing artifacts.
- On error: state what went wrong, why, what you changed — no menus.
- Never present "Option A/B/C" for a request with a clear answer. Act.

## Example pipelines

- [tao-rerun-convnext-cifar10](references/tao-rerun-convnext-cifar10.md)
- [tao-rerun-detr-cppe5](references/tao-rerun-detr-cppe5.md)
- [tao-rerun-segformer-foodseg103](references/tao-rerun-segformer-foodseg103.md)
- [tao-rerun-smolvlm-vqav2](references/tao-rerun-smolvlm-vqav2.md)

모든 파일

69개 파일

tao-finetune-huggingface-model 설치

스킬 파일을 다운로드하여 .claude/skills/ 디렉토리에 추출하세요.

ZIP 다운로드

저장소를 클론하고 스킬 파일을 프로젝트에 복사하세요.

git clone https://github.com/NVIDIA/skills/tree/main/skills/tao-finetune-huggingface-model # Copy SKILL.md to your .claude/skills/ directory

복사 복사
빠른 설정: 스킬 폴더를 .claude/skills/에 복사하세요. Claude는 자동으로 해당 스킬을 감지하고 사용합니다.
저장소 NVIDIA/skills

관련 스킬

web-search
업데이트 된 시간 2026년 6월 29일
webapp-testing
업데이트 된 시간 2026년 6월 29일
lark-base
업데이트 된 시간 2026년 7월 5일
agentmail
업데이트 된 시간 2026년 6월 29일
OR