tao-analyze-gaps-visual-changenet
NVIDIA/skills
NVIDIA TAO VCN Classify 실험에서 각 정답 레이블별로 가장 취약한 샘플을 식별합니다. 이를 위해 임계값 탐색, 취약도 점수 산정, 조명별 확장을 수행하는 Docker 컨테이너를 실행한 후, 다운스트림 보강 또는 재라벨링을 위해 상위 K개의 취약 샘플을 추출합니다.
...모든 것을 확장하십시오TAO VCN Classify 격차 분석 기술
당신은 NVIDIA TAO VCN Classify(Visual Component Net) 추론 결과 분석가입니다. 당신의 임무는 각 정답 라벨별로 결정 임계값으로부터 잘못된 방향으로의 부호 거리(signed distance)를 측정하여 가장 취약한 샘플을 식별한 후, 이를 후속 처리 단계인 데이터 증강 또는 재라벨링을 위해 추출하는 것입니다.
이 스킬은 의도적으로 가볍게 설계되었습니다. VCN의 분류 헤드는 단일 점수 이진 경계( siamese_score에 의한 PASS 대 NO_PASS)이므로, 이 분석은 조사적 성격이 아닌 계산적 성격을 띱니다. 전체 계산은 versions.yaml 에 선언된 tao_toolkit.data_services 이미지(실행 시 해결됨 — ‘설정’ 참조)에 대한 단일 직접 Docker 실행 호출로 수행됩니다. 컨테이너의 엔트리포인트는 ;를 받으며, 우리는 gap_analysis vcn_aoi key=value …를 전달합니다 . 각 재정의는 스크립트의 GapAnalysisConfig 스키마를 선택적으로 재정의하는 기본 Hydra key=value 형식입니다(기본값은 컨테이너에 내장되어 있습니다; docker run ... gap_analysis vcn_aoi --cfg=job을 사용하여 내부 구조를 확인하십시오). (컨테이너 내부에는 dataset 키워드가 없습니다. 이는 TAO 런처의 필러 접두사이며, 여기서는 생략됩니다.) 위임 분석, 다단계 이미지 감사 또는 구성 요소 유형 클러스터링은 필요하지 않습니다. VCN은 이러한 차원을 노출하지 않습니다. 컨테이너가 반환된 후, 소수의 대표적인 취약 샘플만 확인하여 격차를 검증하십시오.
CLI 표면은 데이터 서비스 컨테이너 빌드 간에 달라질 수 있습니다. gap_analysis vcn_aoi 호출이 인자 구문 분석에서 실패할 경우, docker run --rm "$DS_IMAGE" gap_analysis vcn_aoi --cfg=job 명령어를 사용하여 이미지당 한 번씩 실제 스키마를 점검하고, 재시도 전에 이름이 변경된 키(예: inference_csv 대 inference_results_dir, output_dir 대 results_dir)를 조정하십시오. 출력 파케트 파일명은 kpi_gaps.parquet입니다.
입력
- 실험 결과 디렉터리 — TAO VCN Classify 추론 결과인
inference/inference.csv가포함되어 있습니다. 필수 열:input_path,object_name,label,siamese_score. CSV 파일이 아닌 디렉터리 (예:inference/latest/)를 전달하십시오. 컨테이너는inference_results_dir/inference.csv를읽습니다. - 훈련 코드/구성 디렉터리 — VCN 훈련 YAML 파일이 포함되어 있습니다. 컨테이너는 여기서
dataset.classify.input_map(조명 조건 목록)과dataset.classify.image_ext를읽어 각 약한 샘플을 조명별 한 행으로 확장합니다. - 데이터셋 디렉터리 — 각 행(
kpi_media_path)의 상대적input_path앞에 이미지 루트가 추가됩니다. - 스키마 재정의 —
min_recall,top_k_per_label및 선택적으로 고정된임계값이Hydra 재정의 값으로 전달됩니다(기본값:min_recall=1.0,top_k_per_label=50,threshold=-1.0, 이는 스윕을 의미함).top_k_per_label은양의 정수여야 합니다. 이를 생략하면 컨테이너가 "임계값 미만 필터" 모드로 전환되며, 이 경우min_recall=1.0일때 PASS 오분류만 반환하고 NO_PASS 행은 반환하지 않습니다. ‘흔히 발생하는 오류’를 참조하십시오.
설정
임계값 스윕, 취약점 순위 지정 및 조명별 확장은 모두 versions.yaml에 선언된 tao_toolkit.data_services 이미지 내에서 실행됩니다. 실행 시작 시점에 구체적인 URI를 한 번 해결한 후, Docker, NVIDIA 컨테이너 툴킷 및 GPU가 설치되어 있는지 확인하고 이미지가 캐시되어 있는지 확인하십시오:
# versions.yaml에서 tao_toolkit.data_services → 구체적인 nvcr.io/... URI를 확인
DS_IMAGE=$(python3 -c "import yaml,os; print(yaml.safe_load(open(os.environ['TAO_SKILL_BANK_PATH']+'/versions.yaml'))['images']['tao_toolkit']['data_services'])")
echo "DS_IMAGE=$DS_IMAGE"
docker info > /dev/null && echo "OK: docker"
nvidia-smi > /dev/null && echo "OK: GPU"
docker image inspect "$DS_IMAGE" > /dev/null \
|| docker pull "$DS_IMAGE"
TAO_SKILL_BANK_PATH는 일반적으로 설치된 스킬 뱅크에 의해 설정됩니다. 이 변수가 설정되어 있지 않다면, 해결하기 전에 스킬 뱅크 저장소의 루트 디렉터리를 지정하십시오. GPU가 필수입니다. GPU가 없는 호스트에서는 초기 단계에서 중단하면 나중에 발생할 수 있는 혼란스러운 오류를 방지할 수 있습니다.
다음 세 가지 설정 규칙은 작업 수행에 핵심적이며 실수하기 쉬운 부분입니다:
- 경로 마운팅 — 컨테이너가 읽거나 쓰는 모든 호스트 경로(
inference.csv, 훈련용 YAML, 데이터셋 이미지 루트, 출력 디렉터리)는 바인드 마운트되어야 하며, 가장 간단한 방법은-v $WORKSPACE:$WORKSPACE -w $WORKSPACE를사용하여 양쪽에서 절대 경로가 동일하게 해석되도록 하는 것입니다. -
--user $(id -u):$(id -g)를전달하지 마십시오. 이는 컨테이너의transformers임포트 중에KeyError: 'getpwuid(): uid not found:유발하며, 대신' 오류를 chown이호스트 UID로 다시 출력됩니다. -e선택 사항이 아닌 필수 사항입니다. 현재 이미지는 이를 엄격히 요구하며, CLI 재정의 구문을 파싱하기 전에는 ValueError: The subtask vcn_aoi requires the following argument: -e/--experiment_spec_file 오류를발생시켜 종료됩니다.
전체 경로 마운팅 패턴,--user/chown의 근거 및 Alpine chown 명령어, multi--v 지침, 그리고 -e 요구 사항에 대한 자세한 내용은 references/container-setup.md를 참조하십시오.
방법
전체 스킬은 단일 docker run 호출과 그 뒤를 잇는 간단한 시각적 부분 점검으로 구성됩니다. 컨테이너는 내부적으로 1~4단계(임계값 스윕, 취약점 점수 산정, 상위 K개 선정, 조명별 확장)를 수행합니다. 사용자는 Read 도구를 사용하여 5단계(시각적 부분 점검)를 직접 처리합니다.
1~4단계 — 컨테이너 실행
$DOCKER gap_analysis vcn_aoi \
inference_results_dir=/inference/
항상
top_k_per_label을전달해야 합니다. 이 인수는 컨테이너를 기본값인 "임계값 미만의 샘플" 필터에서 적절한 레이블별 상위 K개 순위 지정 모드로 전환합니다.min_recall=1.0일때 임계값은 구조상 모든 NO_PASS 점수 이하로 설정되므로, 임계값 미만 필터는 오분류된 PASS 행만 반환하고 NO_PASS 행은 전혀 반환하지 않습니다. 따라서 증강 큐로는 쓸모가 없습니다.top_k_per_label을양수 정수(사양에서 지정하거나 Hydra 오버라이드로 설정)로 설정하면, 컨테이너는 모든 행에 대해 임계값 대비 부호 있는 취약도를 계산하고 정답 라벨별로 가장 취약한 K개 행을 추출하며, 이는 후속 단계에서 소비하는 라벨별 순위 지정된 출력입니다.
inference.csv를 읽고, 모든 고유한 siamese_score와 최소값 바로 아래의 값 하나를 샅샅이 훑은 후, NO_PASS 클래스 리콜이 min_recall 이상인 후보( 1e-12 허용 오차 적용)를 유지하고, 최상의 F1을 제공하는 임계값을 선택합니다(동점 시: 정밀도, 그 다음 임계값 순으로 결정). 각 행에 대해, 해당 임계값을 기준으로 부호가 있는 약점 값을 계산합니다(양수 = 오분류, 음수 = 정답, 크기 = 마진). weakness 값이 큰 순서대로 정렬한 후, 각 정답 레이블별로 top_k_per_label 개를 추출하고, 훈련용 YAML 파일의 dataset.classify.input_map 및 dataset.classify.image_ext를 사용하여 각 약한 행을 조명 조건별로 하나의 행으로 확장합니다.
리콜 목표를 충족하는 후보 임계값이 없는 경우, 컨테이너는 0이 아닌 값으로 종료되며, 모델이 실제로 달성할 수 있는 리콜 값을 설명하는 unreachable_kpi.txt 파일을 results_dir에 작성합니다. 이 경우, Docker 호출 후 분석을 중단하고, 모델이 어떤 작동 지점에서도 근본적으로 KPI에 도달할 수 없음을 설명하는 단일 섹션의 보고서를 작성하며, 재훈련 또는 라벨 재지정을 권장합니다. 시각적 무작위 검사는 건너뜁니다.
컨테이너는 results_dir에 다음 내용을 기록합니다:
| 아티팩트 | 내용 |
|---|---|
kpi_gaps.parquet |
라벨별 최하위 K개(Top-K) 약점, 조명별로 확장됨. 열: filepath, label, siamese_score, weakness. |
threshold.txt |
선택된 결정 임계값(단일 부동 소수점, 일반 텍스트). |
metrics.json |
선택된 임계값에서: 정밀도, 재현율, F1, 혼동 행렬 {tp, fp, tn, fn}, 그리고 레이블별 {총계, 평균_취약도, 중앙값_취약도, 최대_취약도, 오분류 수}. |
weak_samples_breakdown.txt |
라벨별 유지된 행 세부 정보: 총계, <%> 유지된 모든 행 중, 오분류된 행 수 (약점 > 0), 한계 행 수 (약점 ≤ 0). |
unreachable_kpi.txt |
리콜 목표에 도달할 수 없는 경우에만 작성됩니다. 이 파일이 존재한다는 것은 다음을 의미합니다: 5단계를 건너뛰고, 요약 보고서를 작성하며, 재훈련을 권장합니다. |
컨테이너의 stdout 요약(선택된 임계값, 유지된 행 수, 레이블별 내역)을 사용자의 stdout으로 출력하여, script-check 후크가 실행 결과물을 확인할 수 있도록 하십시오.
5단계 — 시각적 무작위 점검 (소규모, 고정)
unreachable_kpi.txt 파일이 존재하면 이 단계를 건너뜁니다. 그렇지 않은 경우 Read 도구를 사용하여 kpi_gaps.parquet 파일에서 가장 취약한 PASS 샘플 5개와 가장 취약한 NO_PASS 샘플 5개를 확인합니다 (FIRST-lighting 파일 경로를 사용하여 샘플당 한 행으로 중복 제거됨). 각 샘플을 ‘라벨 오류’ / ‘경계 사례’ / ‘데이터 품질 ’ / ‘체계적 오류’ 중 정확히 하나에 분류하고, 확인한 각 이미지(PIL이 있는 경우 128×128로 크기 조정, 없는 경우 그대로 복사)를 에 복사하십시오 . 이것이 필요한 유일한 이미지 검사입니다. 수십 장의 이미지를 검토하거나, 오류 모드 클러스터링을 실행하거나, 골든 이미지를 감사하지 마십시오(VCN에는 골든 이미지가 없습니다).
정확한 샘플 선택 순서, 조명별 중복 제거 규칙, 각 판정 범주의 전체 정의 및 이미지 복사 세부 사항은 references/visual-spot-check.md를 참조하십시오.
호출 방법
워크스페이스, 네 개의 경로, 두 개의 수치 조절값을 복사하여 편집하십시오. 이렇게 하면 엔드투엔드로 실행됩니다. 스크립트 검사 후크가 행 수를 확인할 수 있도록 stdout을 캡처하십시오.
WORKSPACE= # 컨테이너 내부에서 동일하게 마운트됨
EXP_DIR= # inference/inference.csv 및 train.yaml 포함; $WORKSPACE 내부에 있어야 함
DATASET_ROOT= # inference.csv의 input_path 항목에 대한 이미지 루트; $WORKSPACE 내에 있어야 함
MIN_RECALL=1.0 # 누락 0을 기본값으로 설정; KPI 기준이 완화되면 더 낮게 설정
TOP_K=50 # 레이블별 데이터 증강 할당량
OUT="$EXP_DIR/rca_results/$(date +%Y-%m-%d_%H%M%S)"
SPEC="$OUT/vcn_aoi_spec.yaml"
IMG=$(python3 -c "import yaml,os; print(yaml.safe_load(open(os.environ['TAO_SKILL_BANK_PATH']+'/versions.yaml'))['images']['tao_toolkit']['data_services'])")
mkdir -p "$OUT"
# 이번 실행에 대한 갭 분석 사양 작성
cat > "$SPEC" <<EOF
min_recall: $MIN_RECALL
top_k_per_label: $TOP_K
EOF
docker run --gpus all --rm --ipc=host \
-v "$WORKSPACE:$WORKSPACE" -w "$WORKSPACE" \
"$IMG" gap_analysis vcn_aoi \
-e "$SPEC" \
inference_results_dir="$EXP_DIR/inference/latest/" \
train_config="$EXP_DIR/train.yaml" \
kpi_media_path="$DATASET_ROOT" \
results_dir="$OUT"
# 컨테이너는 --user 옵션이 해제된 상태에서 루트 권한으로 쓰기 작업을 수행하므로, 필요한 경우 호스트 UID로 소유권을 다시 설정합니다.
docker run --rm -v "$WORKSPACE:/w" alpine chown -R "$(id -u):$(id -g)" "/w/$(realpath --relative-to="$WORKSPACE" "$OUT")"
# 스크립트 검사 훅이 실제 숫자를 확인할 수 있도록 정상성 출력
python3 - "$OUT" << 'PYEOF'
import json, os, sys
out = sys.argv[1]
unreachable = os.path.join(out, "unreachable_kpi.txt")
if os.path.isfile(unreachable):
print("KPI UNREACHABLE — see", unreachable)
sys.exit(0)
with open(os.path.join(out, "threshold.txt")) as f:
print("threshold:", f.read().strip())
with open(os.path.join(out, "metrics.json")) as f:
m = json.load(f)
print(f"precision={m['precision']:.4f} recall={m['recall']:.4f} f1={m['f1']:.4f}")
import pandas as pd
df = pd.read_parquet(os.path.join(out, "kpi_gaps.parquet"))
print(f"kpi_gaps.parquet: rows={len(df)}, cols={list(df.columns)}")
print(df['label'].value_counts())
PYEOF
출력 결과
모든 내용을 실험 결과 디렉터리 아래의 타임스탬프가 포함된 폴더에 기록합니다. 컨테이너의 출력 결과는 바로 해당 폴더로 전송되며, 시각적 무작위 검사는 rca_images/에 기록됩니다. 실행 시 패키징 훅은 RCA_Report.md 파일이 작성된 후 session/config 캡처 아티팩트를 추가할 수 있습니다.
/rca_results/YYYY-MM-DD_HHMMSS/
├── RCA_Report.md # 전체 격차 분석 보고서 (사용자가 작성)
├── kpi_gaps.parquet # 컨테이너: 레이블별 상위 K개 가장 취약한 항목, 라이트닝별로 확장
├── threshold.txt # 컨테이너: 선택된 결정 임계값 (단일 부동소수점 수)
├── metrics.json # 컨테이너: 혼동 행렬 + 라벨별 분포 통계
├── weak_samples_breakdown.txt # 컨테이너: 라벨별 개수/오분류 개수/한계값 개수
├── unreachable_kpi.txt # 컨테이너: 임계값 중 min_recall을 충족하는 것이 없을 때만 생성됨
├── rca_images/ # 사용자: 확인된 취약 샘플 10개의 썸네일
├── rca_config/ # 후크에 의해 자동 복사됨
└── session log/artifacts # 선택 사항, 런타임에 따라 달라지는 패키징 캡처
실행 시작 시, Bash에서 `date +%Y-%m-%d_%H%M%S `를 실행하여 실제 타임스탬프를 확인하십시오. 하드코딩하거나 추측하지 마십시오. 사용자가 사용자 지정 출력 경로를 지정한 경우, 대신 해당 경로를 사용하되 내부 구조는 동일하게 유지하십시오.
흔히 발생하는 실수
가장 심각한 오류는 min_recall=1.0일 때 top_k_per_label을 생략하는 경우입니다: 해당 리콜 수준에서는 선택된 임계값이 모든 NO_PASS 점수보다 낮거나 같으므로, top_k_per_label이 없으면 컨테이너는 "임계값 미만 샘플" 필터로 대체되어 오분류된 PASS 행만 반환하고 NO_PASS 행은 전혀 반환하지 않게 되어, 증강 큐가 깨집니다. 항상 사양에 명시적인 양의 top_k_per_label (기본값 50)을 포함하거나 Hydra 오버라이드로 지정하십시오.
다음 사항을 다루는 전체 체크리스트는 references/pitfalls.md를 참조하십시오: top_k_per_label 누락; --user 전달; Hydra 오버라이드만 사용하여 호출( `-e ` 없음); $WORKSPACE 외부의 사양 파일; 해결되지 않은 ??? 센티넬이 포함된 사양 파일; 이미지 미다운로드/잘못된 태그; 경로 마운트 불일치; unreachable_kpi.txt 파일 생성; inference.csv에 필수 열 누락; train YAML에 dataset.classify.input_map 또는 image_ext 누락; kpi_media_path가 input_path 접두사와 일치하지 않음; 컨테이너 내부에서 GPU가 감지되지 않음.
보고서 구조
RCA_Report.md 파일을 간결한(1000–1800단어) 계산적 격차 분석 보고서로 작성하십시오. 보고서의 심도는 서술이 아닌 정확한 수치와 명확한 조치 목록에서 비롯되어야 합니다. 전체 보고서 템플릿(7개 섹션: 결론, 임계값 선택, 취약점 분포, 상위 K개 가장 취약한 샘플, 시각적 무작위 점검, 레이블별 분석, 권장 조치 — 혼동 행렬 및 표 레이아웃 포함)은 references/output-template.md에 있습니다 . unreachable_kpi.txt 파일이 존재하는 경우, 3~6절을 해당 파일의 내용을 인용한 하나의 짧은 절로 대체하고, 7절은 ‘재훈련’ 또는 ‘라벨 재지정’이라는 단 하나의 권장 사항으로 요약하십시오.
실행 순서
versions.yaml(images.tao_toolkit.data_services)에서DS_IMAGE를설정한 후,docker info,nvidia-smi,docker image inspect "$DS_IMAGE"(없는 경우 다운로드)를 한 번씩 실행하여 환경을 확인합니다. 실패하는 항목이 있으면 명확한 메시지와 함께 중단합니다.date +%Y-%m-%d_%H%M%S 명령을실행하여 타임스탬프를 얻은 후,디렉터리를 생성합니다./rca_results/ / vcn_aoi_spec.yaml파일을 타임스탬프가 포함된 디렉터리에 작성하고,min_recall및top_k_per_label값을 입력합니다.-e경로가 컨테이너 내에서 올바르게 인식되도록 이 파일을$WORKSPACE아래에 보관합니다.docker run … "$DS_IMAGE" gap_analysis vcn_aoi -e vcn_aoi_spec.yaml inference_results_dir=… train_config=… kpi_media_path=… output_dir=…을실행합니다.컨테이너는kpi_gaps.parquet,threshold.txt,metrics.json,weak_samples_breakdown.txt파일을 results_dir에작성합니다. 스크립트 검사 훅이 실행 결과물을 확인할 수 있도록, 선택된 임계값과 유지된 행 수를 표준 출력(stdout)으로 출력하십시오.unreachable_kpi.txt파일이 존재하면 6단계를 건너뛰고 요약 보고서를 작성합니다. 그렇지 않으면 계속 진행합니다.kpi_gaps.parquet에서10개의 약한 샘플(가장 약한 PASS 5개 + 가장 약한 NO_PASS 5개)을 선택하고, Read를 사용하여 각 테스트 이미지를 확인한 후, 분류하고, 각각을rca_images/에 복사합니다.- 마지막으로
RCA_Report.md를작성합니다. 이 파일을 작성하면 패키징 훅이 트리거되어 세션 로그와 스킬 구성을 함께 복사합니다.
---
name: tao-analyze-gaps-visual-changenet
description: Identifies the weakest samples per ground-truth label in NVIDIA TAO VCN Classify experiments by running a Docker container that performs threshold sweep, weakness scoring, and per-lighting expansion, then surfaces top-K weak samples for downstream augmentation or relabeling.
license: Apache-2.0
---
# TAO VCN Classify Gap Analysis Skill
You are an analyst for NVIDIA TAO VCN Classify (Visual Component Net) inference results. Your job is to identify the **weakest samples per ground-truth label** by measuring signed distance from the decision threshold *in the wrong direction*, then surface them for downstream augmentation or relabeling.
This skill is intentionally lightweight. VCN's classify head is a single-score binary boundary (PASS vs NO_PASS by `siamese_score`), so the analysis is computational, not investigative. The whole computation lives behind one direct `docker run` invocation against the `tao_toolkit.data_services` image declared in `versions.yaml` (resolved at runtime — see Setup). The container's entrypoint takes `<category> <action> [hydra overrides...]`; we pass `gap_analysis vcn_aoi key=value …`. Each override is a bare Hydra `key=value` that selectively overrides the script's `GapAnalysisConfig` schema (defaults are baked into the container; introspect with `docker run ... gap_analysis vcn_aoi --cfg=job`). (There is no `dataset` keyword inside the container — that's the TAO launcher's pillar prefix and is dropped here.) You do **not** need delegated analysis, multi-phase image audits, or component-type clustering — VCN does not expose those dimensions. View only a small set of representative weak samples to qualify the gaps after the container returns.
CLI surface can shift between data-services container builds. If a `gap_analysis vcn_aoi` invocation fails on argument parsing, introspect the actual schema once per image with `docker run --rm "$DS_IMAGE" gap_analysis vcn_aoi --cfg=job` and reconcile any renamed keys (e.g. `inference_csv` vs `inference_results_dir`, `output_dir` vs `results_dir`) before retrying. Output parquet name is `kpi_gaps.parquet`.
---
## Inputs
1. **Experiment result directory** — contains `inference/inference.csv` from TAO VCN Classify inference. Required columns: `input_path`, `object_name`, `label`, `siamese_score`. Pass the **directory** (e.g. `inference/latest/`), not the CSV file — the container reads `inference_results_dir/inference.csv`.
2. **Training code/config directory** — contains the VCN train YAML. The container reads `dataset.classify.input_map` (lighting condition list) and `dataset.classify.image_ext` from it to expand each weak sample into one row per lighting.
3. **Dataset directory** — image root prepended to the relative `input_path` from each row (`kpi_media_path`).
4. **Schema overrides** — `min_recall`, `top_k_per_label`, and optionally a hard-pinned `threshold` are passed as Hydra overrides (defaults: `min_recall=1.0`, `top_k_per_label=50`, `threshold=-1.0` meaning sweep). **`top_k_per_label` must be a positive integer** — omitting it flips the container into "below-threshold filter" mode, which at `min_recall=1.0` returns only PASS misclassifications and zero NO_PASS rows. See Common pitfalls.
---
## Setup
The threshold sweep, weakness ranking, and per-lighting expansion all run inside the `tao_toolkit.data_services` image declared in `versions.yaml`. Resolve the concrete URI once at the top of the run, then confirm Docker, the NVIDIA container toolkit, and a GPU are present and ensure the image is cached:
```bash
# Resolve tao_toolkit.data_services → concrete nvcr.io/... URI from versions.yaml
DS_IMAGE=$(python3 -c "import yaml,os; print(yaml.safe_load(open(os.environ['TAO_SKILL_BANK_PATH']+'/versions.yaml'))['images']['tao_toolkit']['data_services'])")
echo "DS_IMAGE=$DS_IMAGE"
docker info > /dev/null && echo "OK: docker"
nvidia-smi > /dev/null && echo "OK: GPU"
docker image inspect "$DS_IMAGE" > /dev/null \
|| docker pull "$DS_IMAGE"
```
`TAO_SKILL_BANK_PATH` is usually exported by the installed skill bank. If it is unset, point it at the skill-bank repo root before resolving. A GPU is required; aborting early on a GPU-less host saves a confusing late error.
Three setup rules are load-bearing and easy to get wrong:
- **Path mounting** — every host path the container reads or writes (`inference.csv`, train YAML, dataset image root, output dir) must be bind-mounted, simplest with `-v $WORKSPACE:$WORKSPACE -w $WORKSPACE` so absolute paths resolve identically on both sides.
- **Do not pass `--user $(id -u):$(id -g)`** — it triggers `KeyError: 'getpwuid(): uid not found: <uid>'` during the container's `transformers` import; `chown` outputs back to the host UID afterwards instead.
- **`-e <spec>` is required, not optional** — current images hard-require it and exit with `ValueError: The subtask vcn_aoi requires the following argument: -e/--experiment_spec_file` before parsing CLI overrides.
See `references/container-setup.md` for the full path-mounting pattern, the `--user`/`chown` rationale and `alpine` chown command, multi-`-v` guidance, and the `-e <spec>` requirement detail.
---
## Method
The whole skill is a single `docker run` invocation followed by a small visual spot-check. The container does Steps 1–4 internally (threshold sweep, weakness scoring, top-K selection, per-lighting expansion). You handle Step 5 (visual spot-check) directly with the Read tool.
### Step 1–4 — Run the container
```bash
$DOCKER gap_analysis vcn_aoi \
inference_results_dir=<exp_dir>/inference/<label>/ \
train_config=<exp_dir>/train.yaml \
kpi_media_path=<dataset_root> \
results_dir=<rca_results_dir> \
top_k_per_label=50
```
> **Always pass `top_k_per_label`.** This is the argument that switches the container
> from the default "samples below threshold" filter into proper top-K-per-label
> ranking. At `min_recall=1.0` the threshold is by construction at-or-below every
> NO_PASS score, so the below-threshold filter returns ONLY misclassified PASS rows
> and zero NO_PASS rows — useless as an augmentation queue. With `top_k_per_label`
> set to a positive integer (either in the spec or as a Hydra override), the
> container computes signed weakness against the threshold for every row and
> surfaces the K weakest **per ground-truth label**, which is the per-label ranked
> output downstream steps consume.
Reads `inference.csv`, sweeps every unique `siamese_score` plus one value just below the minimum, keeps the candidates with NO_PASS-class recall ≥ `min_recall` (with `1e-12` tolerance), then picks the threshold with the best F1 (tie-break: precision, then threshold value). For every row, computes signed weakness from that threshold (positive = misclassified, negative = correct, magnitude = margin). Sorts by weakness descending and takes the top `top_k_per_label` per ground-truth label, then expands each weak row into one row per lighting condition using `dataset.classify.input_map` and `dataset.classify.image_ext` from the train YAML.
If **no** candidate threshold meets the recall target, the container exits non-zero and writes `unreachable_kpi.txt` into `results_dir` explaining which recall the model can actually achieve. In that case, stop the analysis after the docker call, write a one-section report explaining the model fundamentally cannot reach the KPI at any operating point, and recommend retraining or relabeling — skip the visual spot-check.
**Container writes into `results_dir`:**
| Artifact | Contents |
|----------|----------|
| `kpi_gaps.parquet` | Top-K weakest per label, expanded per lighting. Columns: `filepath`, `label`, `siamese_score`, `weakness`. |
| `threshold.txt` | Chosen decision threshold (single float, plain text). |
| `metrics.json` | At the chosen threshold: `precision`, `recall`, `f1`, confusion matrix `{tp, fp, tn, fn}`, plus per-label `{total, mean_weakness, median_weakness, max_weakness, n_misclassified}`. |
| `weak_samples_breakdown.txt` | Per-label kept-row breakdown: `<count>` total, `<%>` of all kept rows, `N` misclassified (weakness > 0), `N` marginal (weakness ≤ 0). |
| `unreachable_kpi.txt` | Only written when the recall target is unreachable. Presence of this file means: skip Step 5, write the abridged report, recommend retrain. |
Print the container's stdout summary (chosen threshold, kept-row counts, per-label breakdown) to your own stdout so the script-check hook can verify the run produced output.
### Step 5 — Visual spot check (small, fixed)
Skip this step if `unreachable_kpi.txt` exists. Otherwise use the Read tool to **view** the 5 weakest PASS samples and the 5 weakest NO_PASS samples from `kpi_gaps.parquet` (deduplicated to one row per sample, using the FIRST-lighting `filepath`), classify each as exactly one of **mislabeled** / **edge case** / **data quality** / **systematic**, and copy each viewed image (resized to 128×128 if PIL is available, otherwise just copy) into `<results_dir>/rca_images/`. This is the only image inspection required — do not view dozens of images, run failure mode clustering, or audit goldens (VCN has no golden images).
See `references/visual-spot-check.md` for the exact sample-selection sort, the per-lighting deduplication rule, the full definition of each verdict category, and the image-copy detail.
---
## Reference invocation
Paste-and-edit the workspace, the four paths, and the two numeric knobs; this runs end-to-end. Capture stdout so the script-check hook sees row counts.
```bash
WORKSPACE=<absolute path> # mounted identically inside the container
EXP_DIR=<experiment_result_dir> # contains inference/inference.csv and train.yaml; must be inside $WORKSPACE
DATASET_ROOT=<dataset_root> # image root for inference.csv input_path entries; must be inside $WORKSPACE
MIN_RECALL=1.0 # zero-miss default; lower if KPI relaxes
TOP_K=50 # per-label augmentation budget
OUT="$EXP_DIR/rca_results/$(date +%Y-%m-%d_%H%M%S)"
SPEC="$OUT/vcn_aoi_spec.yaml"
IMG=$(python3 -c "import yaml,os; print(yaml.safe_load(open(os.environ['TAO_SKILL_BANK_PATH']+'/versions.yaml'))['images']['tao_toolkit']['data_services'])")
mkdir -p "$OUT"
# Write the gap-analysis spec for this run
cat > "$SPEC" <<EOF
min_recall: $MIN_RECALL
top_k_per_label: $TOP_K
EOF
docker run --gpus all --rm --ipc=host \
-v "$WORKSPACE:$WORKSPACE" -w "$WORKSPACE" \
"$IMG" gap_analysis vcn_aoi \
-e "$SPEC" \
inference_results_dir="$EXP_DIR/inference/latest/" \
train_config="$EXP_DIR/train.yaml" \
kpi_media_path="$DATASET_ROOT" \
results_dir="$OUT"
# Container writes as root with --user dropped; chown back to host UID if needed.
docker run --rm -v "$WORKSPACE:/w" alpine chown -R "$(id -u):$(id -g)" "/w/$(realpath --relative-to="$WORKSPACE" "$OUT")"
# Sanity print so the script-check hook sees real numbers
python3 - "$OUT" << 'PYEOF'
import json, os, sys
out = sys.argv[1]
unreachable = os.path.join(out, "unreachable_kpi.txt")
if os.path.isfile(unreachable):
print("KPI UNREACHABLE — see", unreachable)
sys.exit(0)
with open(os.path.join(out, "threshold.txt")) as f:
print("threshold:", f.read().strip())
with open(os.path.join(out, "metrics.json")) as f:
m = json.load(f)
print(f"precision={m['precision']:.4f} recall={m['recall']:.4f} f1={m['f1']:.4f}")
import pandas as pd
df = pd.read_parquet(os.path.join(out, "kpi_gaps.parquet"))
print(f"kpi_gaps.parquet: rows={len(df)}, cols={list(df.columns)}")
print(df['label'].value_counts())
PYEOF
```
---
## Outputs
Write everything into a timestamped folder under the experiment result directory. The container's outputs go straight there; the visual spot-check writes `rca_images/`; any runtime packaging hook may add session/config capture artifacts after `RCA_Report.md` is written.
```
<experiment_result_dir>/rca_results/YYYY-MM-DD_HHMMSS/
├── RCA_Report.md # Full gap analysis report (you write this)
├── kpi_gaps.parquet # Container: top-K weakest per label, expanded per lighting
├── threshold.txt # Container: chosen decision threshold (single float)
├── metrics.json # Container: confusion matrix + per-label distribution stats
├── weak_samples_breakdown.txt # Container: per-label count/misclassified/marginal counts
├── unreachable_kpi.txt # Container: ONLY when no threshold meets min_recall
├── rca_images/ # You: thumbnails of the 10 viewed weak samples
├── rca_config/ # Auto-copied by hook
└── session log/artifacts # Optional, runtime-dependent packaging capture
```
At the start of the run, get the real timestamp by running `date +%Y-%m-%d_%H%M%S` in Bash. Do NOT hardcode or guess. If the user specifies a custom output path, use that instead but maintain the same internal structure.
---
## Common pitfalls
The single most consequential failure mode is **forgetting `top_k_per_label` when `min_recall=1.0`**: at that recall the chosen threshold sits at or below every NO_PASS score, so without `top_k_per_label` the container falls back to a "samples below threshold" filter that returns ONLY misclassified PASS rows and zero NO_PASS rows, breaking the augmentation queue. Always include an explicit positive `top_k_per_label` (default 50) in the spec or as a Hydra override.
See `references/pitfalls.md` for the complete checklist, covering: forgetting `top_k_per_label`; passing `--user`; calling with only Hydra overrides (no `-e <spec>`); spec file outside `$WORKSPACE`; spec file with unresolved `???` sentinels; image not pulled / wrong tag; path-mount mismatch; `unreachable_kpi.txt` written; `inference.csv` missing required columns; train YAML missing `dataset.classify.input_map` or `image_ext`; `kpi_media_path` not matching `input_path` prefixes; and no GPU detected from inside the container.
---
## Report Structure
Write `RCA_Report.md` as a tight (1000–1800 word) computational gap analysis — depth comes from accurate numbers and a clear action list, not narrative. The full report template (7 sections: Verdict, Threshold Selection, Weakness Distribution, Top-K Weakest Samples, Visual Spot Check, Per-Label Breakdown, Recommended Actions — with the confusion-matrix and table layouts) is in `references/output-template.md`. When `unreachable_kpi.txt` exists, replace sections 3–6 with a single short section quoting that file's contents and collapse section 7 to one recommendation: retrain or relabel.
---
## Execution Order
1. Resolve `DS_IMAGE` from `versions.yaml` (`images.tao_toolkit.data_services`), then run `docker info`, `nvidia-smi`, and `docker image inspect "$DS_IMAGE"` (pulling if missing) once to confirm the environment. Abort with a clear message if any fail.
2. Run `date +%Y-%m-%d_%H%M%S` to get the timestamp; create `<experiment_result_dir>/rca_results/<timestamp>/`.
3. Write `vcn_aoi_spec.yaml` into the timestamped dir with `min_recall` and `top_k_per_label` filled in. Keep it under `$WORKSPACE` so the `-e` path resolves inside the container.
4. Run `docker run … "$DS_IMAGE" gap_analysis vcn_aoi -e vcn_aoi_spec.yaml inference_results_dir=… train_config=… kpi_media_path=… output_dir=…`. The container writes `kpi_gaps.parquet`, `threshold.txt`, `metrics.json`, `weak_samples_breakdown.txt` into `results_dir`. Print the chosen threshold and kept-row counts to stdout so the script-check hook can verify the run produced output.
5. If `unreachable_kpi.txt` exists, skip Step 6 and write the abridged report. Otherwise continue.
6. Pick 10 weak samples (5 weakest PASS + 5 weakest NO_PASS) from `kpi_gaps.parquet`, view each test image with Read, classify, and copy each into `rca_images/`.
7. Write `RCA_Report.md` last — writing it triggers the packaging hook, which copies session logs and skill config alongside.
모든 파일
1개 파일tao-analyze-gaps-visual-changenet 설치
스킬 파일을 다운로드하여 .claude/skills/ 디렉터리에 압축을 풀어주세요.
ZIP 다운로드저장소를 클론하고 스킬 파일을 프로젝트에 복사하세요.
git clone https://github.com/NVIDIA/skills/tree/main/skills/tao-analyze-gaps-visual-changenet # Copy SKILL.md to your .claude/skills/ directory
복사





집
