選項
首頁首頁 Skill 數據科學與機器學習 tao-analyze-gaps-visual-changenet

tao-analyze-gaps-visual-changenet

NVIDIA/skills NVIDIA/skills

透過執行一個 Docker 容器,該容器會執行閾值掃描、弱點評分及按光照條件擴展等操作,據此在 NVIDIA TAO VCN Classify 實驗中,針對每個真實標籤識別出最弱的樣本,並將前 K 個弱樣本呈現出來,供後續的增強或重新標註使用。

...展開全部
2
更新時間 2026-09-28

TAO VCN 分類差距分析技能

身為 NVIDIA TAO VCN Classify(視覺元件網路)推論結果的分析師,您的任務是根據各真實標籤,透過測量樣本與決策閾值在錯誤方向上的有符距離,來識別最薄弱的樣本,並將其篩選出來以供後續進行資料增強或重新標記。

此技能刻意設計為輕量級。VCN 的分類頭(Classify head)是一個單分數二元邊界(透過siamese_score 判定為 PASS 或 NO_PASS),因此此分析屬運算性質,而非調查性質。 整個運算流程透過單次直接執行 Docker指令,針對versions.yaml中宣告的tao_toolkit.data_services映像檔進行(於執行時解析 — 請參閱「設定」)。 容器的入口點接受 [hydra 覆寫項目...];我們傳入gap_analysis vcn_aoi key=value …。 每個覆寫項皆為純粹的 Hydrakey=value格式,用於選擇性地覆寫腳本中的GapAnalysisConfig架構(預設值已內建於容器中;可透過 `docker run ... gap_analysis vcn_aoi --cfg=job` 進行檢視)。 (容器內部沒有「dataset」關鍵字——那是 TAO 啟動器的支柱前綴,此處已省略。) 您無需委派式分析、多階段映像審計或元件類型聚類 — VCN 並未公開這些維度。待容器返回後,僅需檢視一小組具代表性的弱樣本,即可判定差距。

CLI 介面可能因資料服務容器的建置版本而有所不同。若gap_analysis vcn_aoi執行時因參數解析失敗,請使用docker run --rm "$DS_IMAGE" gap_analysis vcn_aoi針對每個映像檔檢查一次實際的資料結構,--cfg=job指令對每個映像執行一次實際結構的檢視,並在重新嘗試前調整任何更名的鍵值(例如inference_csv與inference_results_dir、output_dir與results_dir)。輸出 Parquet 檔案名稱為kpi_gaps.parquet。

輸入

  1. 實驗結果目錄— 包含來自 TAO VCN Classify 推論的`inference/inference.csv` 檔案。必備欄位:input_path、object_name、label、siamese_score。 請傳入目錄(例如inference/latest/),而非 CSV 檔案 — 容器會從inference_results_dir/inference.csv 讀取資料。
  2. 訓練程式碼/設定目錄— 包含 VCN 訓練用的 YAML 檔案。容器會從中讀取dataset.classify.input_map(光照條件清單)和dataset.classify.image_ext,以將每個弱樣本擴展為每種光照條件各一行。
  3. 資料集目錄— 將圖像根目錄附加至每行(kpi_media_path)的相對input_path之前。
  4. 模式覆寫—min_recall、top_k_per_label,以及可選的硬性釘定閾值,將作為 Hydra 覆寫參數傳入(預設值:min_recall=1.0、top_k_per_label=50、threshold=-1.0表示掃描)。top_k_per_label必須為正整數— 若省略此參數,容器將切換至「低於閾值篩選」模式;當min_recall=1.0時,此模式僅會回傳 PASS 類別的誤分類結果,且不產生任何 NO_PASS 資料列。請參閱「常見陷阱」。

設定

閾值掃描、弱點排序以及按照明擴展,皆在versions.yaml 中宣告的tao_toolkit.data_services映像內執行。請在執行開始時先解析一次具體的 URI,接著確認 Docker、NVIDIA 容器工具包及 GPU 是否就緒,並確保該映像已快取:

# 從 versions.yaml 解析 tao_toolkit.data_services → 具體的 nvcr.io/... URI
DS_IMAGE=$(python3 -c "import yaml,os; print(yaml.safe_load(open(os.environ['TAO_SKILL_BANK_PATH']+'/versions.yaml'))['images']['tao_toolkit']['data_services'])")
echo "DS_IMAGE=$DS_IMAGE"

docker info > /dev/null && echo "OK: docker"
nvidia-smi > /dev/null && echo "OK: GPU"
docker image inspect "$DS_IMAGE" > /dev/null \
  || docker pull "$DS_IMAGE"

TAO_SKILL_BANK_PATH通常由已安裝的技能庫所匯出。若未設定,請在解析前將其指向技能庫儲存庫的根目錄。必須具備 GPU;若在無 GPU 的主機上及早中止,可避免後續出現令人困惑的錯誤。

有三項設定規則至關重要且容易出錯:

  • 路徑掛載— 容器讀取或寫入的每個主機路徑(inference.csv、訓練 YAML、資料集映像檔根目錄、輸出目錄)都必須進行綁定掛載,最簡單的方法是使用-v $WORKSPACE:$WORKSPACE -w $WORKSPACE,以確保絕對路徑在雙方都能解析為相同結果。
  • 請勿傳入--user $(id -u):$(id -g)— 這會在容器匯入transformers時觸發KeyError:'getpwuid(): uid not found:';此後,chown會將權限還原為主機的 UID。
  • -e是必需的,並非可選— 當前映像檔嚴格要求此參數,並會在解析 CLI 覆寫設定前,因ValueError: 子任務 vcn_aoi 需要以下參數:-e/--experiment_spec_file而退出。

請參閱references/container-setup.md以了解完整的路徑掛載模式、--user/chown的設計理念與Alpinechown 指令、多重 --v設定的指引,以及-e要求的詳細說明。

方法

整個技能僅包含一次Docker run調用,隨後進行簡短的視覺抽查。容器會在內部執行步驟 1–4(閾值掃描、弱點評分、前 K 名篩選、每種照明條件下的擴展分析)。您則需直接使用 Read 工具處理步驟 5(視覺抽查)。

步驟 1–4 — 執行容器

$DOCKER gap_analysis vcn_aoi \
    inference_results_dir=/inference/

請務必傳入top_k_per_label 參數。此參數會將容器 從預設的「低於閾值的樣本」過濾機制,切換為正確的「每標籤前 K 名」 排序機制。 當min_recall=1.0時,閾值在設計上會設定為等於或低於每個 NO_PASS 分數,因此「低於閾值」篩選機制僅會返回誤分類的 PASS 資料列 以及零個 NO_PASS 資料列——作為增強佇列毫無用處。 當top_k_per_label 設定為正整數(無論是在規格中或透過 Hydra 覆寫), 容器會針對每筆資料列計算其相對於閾值的帶符號弱點程度, 並篩選出每個真實標籤下最弱的 K筆資料,這便是下游步驟所使用的 按標籤排序的輸出結果。

讀取inference.csv 檔案,掃描每個唯一的siamese_score以及略低於最小值的數值,保留 NO_PASS 類別召回率 ≥min_recall(容差為1e-12)的候選值,然後選擇 F1 值最佳的閾值(平手時依精確度、閾值順序決定)。 針對每一行,根據該閾值計算帶符號的弱點(正值 = 誤分類,負值 = 正確,絕對值 = 邊界)。 依「弱點」值由大至小排序,並針對每個真實標籤選取前top_k_per_label個候選閾值,接著利用訓練 YAML 中的dataset.classify.input_map和dataset.classify.image_ext,將每筆弱點資料行展開為每種照明條件各一筆資料行。

若無任何候選閾值能達到召回率目標,容器將以非零狀態退出,並在results_dir 目錄中寫入unreachable_kpi.txt 檔案,說明模型實際能達到的召回率。 在此情況下,於呼叫 Docker 後停止分析,撰寫一份單節報告說明模型在任何運作點上皆根本無法達到 KPI,並建議重新訓練或重新標記 — 跳過視覺抽查步驟。

容器會寫入results_dir:

產出檔案 內容
kpi_gaps.parquet 各標籤下 Top-K 最弱項目,按照明條件細分。欄位:filepath、label、siamese_score、weakness。
threshold.txt 選定的決策閾值(單一浮點數,純文字)。
metrics.json 在選定的閾值下:精確率、召回率、F1 分數、混淆矩陣{真陽性, 假陽性, 真陰性, 假陰性},以及各標籤的{總數, 平均弱點值, 中位數弱點值, 最大弱點值, 誤分類數}。
weak_samples_breakdown.txt 按標籤劃分的保留行細項: 總數, <%> 在所有保留的資料列中,N 個誤分類(弱點 > 0),N 個邊緣案例(弱點 ≤ 0)。
unreachable_kpi.txt 僅在召回率目標無法達成時寫入。此檔案的存在表示:跳過步驟 5、寫入摘要報告、建議重新訓練。

將容器的標準輸出摘要(選定的閾值、保留列數、各標籤細項)輸出至您自己的標準輸出,以便腳本檢查掛鉤能驗證執行所產生的輸出。

步驟 5 — 視覺抽查(小規模、固定)

若存在unreachable_kpi.txt檔案,則跳過此步驟。否則,請使用 Read工具檢視 kpi_gaps.parquet中的 5 個最弱的「PASS」樣本與 5 個最弱的「NO_PASS」樣本(已根據 FIRST-lighting檔案路徑進行去重,每樣本僅保留一行), 將每個樣本精確歸類為「標籤錯誤」、「邊界案例」、「資料品質」或「系統性問題」其中一類,並將每個檢視過的影像(若 PIL 可用則調整為 128×128 尺寸,否則直接複製)複製到/rca_images/ 目錄中。 這是唯一需要的影像檢查步驟——請勿檢視數十張影像、執行故障模式聚類分析,或審核黃金標準影像(VCN 並無黃金標準影像)。

有關確切的樣本選取排序、各照明條件下的去重規則、各判定類別的完整定義,以及圖片複製的詳細說明,請參閱references/visual-spot-check.md。

參考執行方式

請複製並編輯工作區、四個路徑以及兩個數值參數;此操作將執行端到端測試。請擷取標準輸出(stdout),以便 script-check 鉤子能偵測到行數。

WORKSPACE=           # 在容器內以相同方式掛載
EXP_DIR=     # 包含 inference/inference.csv 和 train.yaml;必須位於 $WORKSPACE 內
DATASET_ROOT=         # 用於 inference.csv 中 input_path 條目的影像根目錄;必須位於 $WORKSPACE 內
MIN_RECALL=1.0                       # 預設為零漏檢;若 KPI 要求放寬則可調低
TOP_K=50                             # 每標籤的增強預算
OUT="$EXP_DIR/rca_results/$(date +%Y-%m-%d_%H%M%S)"
SPEC="$OUT/vcn_aoi_spec.yaml"
IMG=$(python3 -c "import yaml,os; print(yaml.safe_load(open(os.environ['TAO_SKILL_BANK_PATH']+'/versions.yaml'))['images']['tao_toolkit']['data_services'])")

mkdir -p "$OUT"

# 為本次執行撰寫差距分析規格
cat > "$SPEC" <<EOF
min_recall: $MIN_RECALL
top_k_per_label: $TOP_K
EOF

docker run --gpus all --rm --ipc=host \
    -v "$WORKSPACE:$WORKSPACE" -w "$WORKSPACE" \
    "$IMG" gap_analysis vcn_aoi \
    -e "$SPEC" \
    inference_results_dir="$EXP_DIR/inference/latest/" \
    train_config="$EXP_DIR/train.yaml" \
    kpi_media_path="$DATASET_ROOT" \
    results_dir="$OUT"

# 容器以 root 身分寫入,且已移除 --user 參數;如有需要,將檔案所有權改回主機 UID。
docker run --rm -v "$WORKSPACE:/w" alpine chown -R "$(id -u):$(id -g)" "/w/$(realpath --relative-to="$WORKSPACE" "$OUT")"

# 輸出驗證訊息,讓 script-check 鉤子能看到真實數值
python3 - "$OUT" << 'PYEOF'
import json, os, sys
out = sys.argv[1]
unreachable = os.path.join(out, "unreachable_kpi.txt")
if os.path.isfile(unreachable):
    print("KPI 無法存取 — 請參閱", unreachable)
    sys.exit(0)
with open(os.path.join(out, "threshold.txt")) as f:
    print("閾值:", f.read().strip())
with open(os.path.join(out, "metrics.json")) as f:
    m = json.load(f)
print(f"precision={m['precision']:.4f} recall={m['recall']:.4f} f1={m['f1']:.4f}")
import pandas as pd
df = pd.read_parquet(os.path.join(out, "kpi_gaps.parquet"))
print(f"kpi_gaps.parquet: 行數={len(df)}, 欄數={list(df.columns)}")
print(df['label'].value_counts())
PYEOF

輸出結果

將所有內容寫入實驗結果目錄下的帶有時間戳記的資料夾中。容器的輸出會直接存入該處;視覺抽查結果會寫入rca_images/;任何執行時封裝掛鉤都可能在寫入RCA_Report.md之後,將 session/config 擷取的副產品追加至該資料夾。

/rca_results/YYYY-MM-DD_HHMMSS/
├── RCA_Report.md              # 完整的差距分析報告(由您撰寫)
├── kpi_gaps.parquet           # 容器:每個標籤的前 K 個最弱樣本,按照明條件細分
├── threshold.txt              # 容器:選定的決策閾值(單一浮點數)
├── metrics.json               # 容器:混淆矩陣 + 各標籤的分布統計資料
├── weak_samples_breakdown.txt # 容器:各標籤的計數/誤分類/邊際計數
├── unreachable_kpi.txt        # 容器:僅當無任何閾值滿足 min_recall 時產生
├── rca_images/                # 由您提供:10 個已檢視的弱樣本縮圖
├── rca_config/                # 由 hook 自動複製
└── session log/artifacts      # 可選,依執行時環境而定的封裝擷取

執行開始時,請在 Bash 中執行 `date +%Y-%m-%d_%H%M%S` 以取得實際時間戳記。切勿硬編碼或憑空猜測。若使用者指定了自訂輸出路徑,請改用該路徑,但須維持相同的內部結構。

常見陷阱

當min_recall=1.0 時,最嚴重的失敗模式就是遺漏top_k_per_label: 在該召回率下,選定的閾值位於每個 NO_PASS 分數之下或等於該分數,因此若未設定 top_k_per_label,容器將回退至「低於閾值的樣本」篩選機制,該機制僅會返回誤分類的 PASS 列,而不會返回任何 NO_PASS 列,從而破壞增強佇列。 請務必在規格中或透過 Hydra 覆寫,明確包含正類別的top_k_per_label(預設值為 50)。

請參閱references/pitfalls.md以獲取完整的檢查清單,內容涵蓋:遺漏top_k_per_label;傳入--user 參數;僅使用 Hydra 覆寫參數進行呼叫(未使用-e );規格檔位於$WORKSPACE 之外;規格檔中含有未解析的???哨兵;未拉取映像檔/標籤錯誤; 路徑掛載不符;寫入了unreachable_kpi.txt;inference.csv缺少必備欄位;訓練 YAML 缺少dataset.classify.input_map或image_ext;kpi_media_path與input_path前綴不匹配;以及從容器內部未偵測到 GPU。

報告結構

撰寫RCA_Report.md文件,作為一份簡明扼要(1000–1800 字)的運算差距分析報告——報告的深度應來自精確的數據和清晰的行動清單,而非敘述性內容。 完整的報告範本(共 7 個章節:結論、閾值選擇、弱點分佈、Top-K 最弱樣本、視覺抽查、各標籤細項分析、建議行動——包含混淆矩陣與表格佈局)位於references/output-template.md 中。 若存在unreachable_kpi.txt檔案,請將第 3 至 6 節替換為單一簡短節,引用該檔案的內容,並將第 7 節簡化為一項建議:重新訓練或重新標記。

執行順序

  1. 從versions.yaml(images.tao_toolkit.data_services)解析DS_IMAGE,然後執行docker info、nvidia-smi 以及docker image inspect "$DS_IMAGE"(若不存在則先拉取),以確認環境設定。若任何步驟失敗,則以明確訊息終止執行。
  2. 執行`date +%Y-%m-%d_%H%M%S` 以取得時間戳記;建立目錄 `/rca_results/` 及 `/`。
  3. 將vcn_aoi_spec.yaml寫入帶有時間戳記的目錄中,並填入min_recall和top_k_per_label參數。將該檔案保留在$WORKSPACE目錄下,以便-e路徑能在容器內部解析。
  4. 執行docker run … "$DS_IMAGE" gap_analysis vcn_aoi -e vcn_aoi_spec.yaml inference_results_dir=… train_config=… kpi_media_path=… output_dir=…. 容器會將kpi_gaps.parquet、threshold.txt、metrics.json 及weak_samples_breakdown.txt寫入results_dir。將選定的閾值與保留的行數輸出至標準輸出(stdout),以便 script-check 鉤子能驗證執行過程是否產生了預期輸出。
  5. 若unreachable_kpi.txt檔案存在,則跳過步驟 6 並寫入摘要報告;否則繼續執行。
  6. 從kpi_gaps.parquet 中挑選 10 個弱樣本(5 個最弱的「通過」+ 5 個最弱的「未通過」),使用 Read 檢視每張測試圖片,進行分類,並將每張圖片複製到rca_images/ 目錄中。
  7. 最後寫入RCA_Report.md— 寫入此檔案會觸發封裝鉤子,該鉤子會一併複製會話日誌和技能配置。
在 GitHub 上查看
---
name: tao-analyze-gaps-visual-changenet
description: Identifies the weakest samples per ground-truth label in NVIDIA TAO VCN Classify experiments by running a Docker container that performs threshold sweep, weakness scoring, and per-lighting expansion, then surfaces top-K weak samples for downstream augmentation or relabeling.
license: Apache-2.0
---

# TAO VCN Classify Gap Analysis Skill

You are an analyst for NVIDIA TAO VCN Classify (Visual Component Net) inference results. Your job is to identify the **weakest samples per ground-truth label** by measuring signed distance from the decision threshold *in the wrong direction*, then surface them for downstream augmentation or relabeling.

This skill is intentionally lightweight. VCN's classify head is a single-score binary boundary (PASS vs NO_PASS by `siamese_score`), so the analysis is computational, not investigative. The whole computation lives behind one direct `docker run` invocation against the `tao_toolkit.data_services` image declared in `versions.yaml` (resolved at runtime — see Setup). The container's entrypoint takes `<category> <action> [hydra overrides...]`; we pass `gap_analysis vcn_aoi key=value …`. Each override is a bare Hydra `key=value` that selectively overrides the script's `GapAnalysisConfig` schema (defaults are baked into the container; introspect with `docker run ... gap_analysis vcn_aoi --cfg=job`). (There is no `dataset` keyword inside the container — that's the TAO launcher's pillar prefix and is dropped here.) You do **not** need delegated analysis, multi-phase image audits, or component-type clustering — VCN does not expose those dimensions. View only a small set of representative weak samples to qualify the gaps after the container returns.

CLI surface can shift between data-services container builds. If a `gap_analysis vcn_aoi` invocation fails on argument parsing, introspect the actual schema once per image with `docker run --rm "$DS_IMAGE" gap_analysis vcn_aoi --cfg=job` and reconcile any renamed keys (e.g. `inference_csv` vs `inference_results_dir`, `output_dir` vs `results_dir`) before retrying. Output parquet name is `kpi_gaps.parquet`.

---

## Inputs

1. **Experiment result directory** — contains `inference/inference.csv` from TAO VCN Classify inference. Required columns: `input_path`, `object_name`, `label`, `siamese_score`. Pass the **directory** (e.g. `inference/latest/`), not the CSV file — the container reads `inference_results_dir/inference.csv`.
2. **Training code/config directory** — contains the VCN train YAML. The container reads `dataset.classify.input_map` (lighting condition list) and `dataset.classify.image_ext` from it to expand each weak sample into one row per lighting.
3. **Dataset directory** — image root prepended to the relative `input_path` from each row (`kpi_media_path`).
4. **Schema overrides** — `min_recall`, `top_k_per_label`, and optionally a hard-pinned `threshold` are passed as Hydra overrides (defaults: `min_recall=1.0`, `top_k_per_label=50`, `threshold=-1.0` meaning sweep). **`top_k_per_label` must be a positive integer** — omitting it flips the container into "below-threshold filter" mode, which at `min_recall=1.0` returns only PASS misclassifications and zero NO_PASS rows. See Common pitfalls.

---

## Setup

The threshold sweep, weakness ranking, and per-lighting expansion all run inside the `tao_toolkit.data_services` image declared in `versions.yaml`. Resolve the concrete URI once at the top of the run, then confirm Docker, the NVIDIA container toolkit, and a GPU are present and ensure the image is cached:

```bash
# Resolve tao_toolkit.data_services → concrete nvcr.io/... URI from versions.yaml
DS_IMAGE=$(python3 -c "import yaml,os; print(yaml.safe_load(open(os.environ['TAO_SKILL_BANK_PATH']+'/versions.yaml'))['images']['tao_toolkit']['data_services'])")
echo "DS_IMAGE=$DS_IMAGE"

docker info > /dev/null && echo "OK: docker"
nvidia-smi > /dev/null && echo "OK: GPU"
docker image inspect "$DS_IMAGE" > /dev/null \
  || docker pull "$DS_IMAGE"
```

`TAO_SKILL_BANK_PATH` is usually exported by the installed skill bank. If it is unset, point it at the skill-bank repo root before resolving. A GPU is required; aborting early on a GPU-less host saves a confusing late error.

Three setup rules are load-bearing and easy to get wrong:

- **Path mounting** — every host path the container reads or writes (`inference.csv`, train YAML, dataset image root, output dir) must be bind-mounted, simplest with `-v $WORKSPACE:$WORKSPACE -w $WORKSPACE` so absolute paths resolve identically on both sides.
- **Do not pass `--user $(id -u):$(id -g)`** — it triggers `KeyError: 'getpwuid(): uid not found: <uid>'` during the container's `transformers` import; `chown` outputs back to the host UID afterwards instead.
- **`-e <spec>` is required, not optional** — current images hard-require it and exit with `ValueError: The subtask vcn_aoi requires the following argument: -e/--experiment_spec_file` before parsing CLI overrides.

See `references/container-setup.md` for the full path-mounting pattern, the `--user`/`chown` rationale and `alpine` chown command, multi-`-v` guidance, and the `-e <spec>` requirement detail.

---

## Method

The whole skill is a single `docker run` invocation followed by a small visual spot-check. The container does Steps 1–4 internally (threshold sweep, weakness scoring, top-K selection, per-lighting expansion). You handle Step 5 (visual spot-check) directly with the Read tool.

### Step 1–4 — Run the container

```bash
$DOCKER gap_analysis vcn_aoi \
    inference_results_dir=<exp_dir>/inference/<label>/ \
    train_config=<exp_dir>/train.yaml \
    kpi_media_path=<dataset_root> \
    results_dir=<rca_results_dir> \
    top_k_per_label=50
```

> **Always pass `top_k_per_label`.** This is the argument that switches the container
> from the default "samples below threshold" filter into proper top-K-per-label
> ranking. At `min_recall=1.0` the threshold is by construction at-or-below every
> NO_PASS score, so the below-threshold filter returns ONLY misclassified PASS rows
> and zero NO_PASS rows — useless as an augmentation queue. With `top_k_per_label`
> set to a positive integer (either in the spec or as a Hydra override), the
> container computes signed weakness against the threshold for every row and
> surfaces the K weakest **per ground-truth label**, which is the per-label ranked
> output downstream steps consume.

Reads `inference.csv`, sweeps every unique `siamese_score` plus one value just below the minimum, keeps the candidates with NO_PASS-class recall ≥ `min_recall` (with `1e-12` tolerance), then picks the threshold with the best F1 (tie-break: precision, then threshold value). For every row, computes signed weakness from that threshold (positive = misclassified, negative = correct, magnitude = margin). Sorts by weakness descending and takes the top `top_k_per_label` per ground-truth label, then expands each weak row into one row per lighting condition using `dataset.classify.input_map` and `dataset.classify.image_ext` from the train YAML.

If **no** candidate threshold meets the recall target, the container exits non-zero and writes `unreachable_kpi.txt` into `results_dir` explaining which recall the model can actually achieve. In that case, stop the analysis after the docker call, write a one-section report explaining the model fundamentally cannot reach the KPI at any operating point, and recommend retraining or relabeling — skip the visual spot-check.

**Container writes into `results_dir`:**

| Artifact | Contents |
|----------|----------|
| `kpi_gaps.parquet` | Top-K weakest per label, expanded per lighting. Columns: `filepath`, `label`, `siamese_score`, `weakness`. |
| `threshold.txt` | Chosen decision threshold (single float, plain text). |
| `metrics.json` | At the chosen threshold: `precision`, `recall`, `f1`, confusion matrix `{tp, fp, tn, fn}`, plus per-label `{total, mean_weakness, median_weakness, max_weakness, n_misclassified}`. |
| `weak_samples_breakdown.txt` | Per-label kept-row breakdown: `<count>` total, `<%>` of all kept rows, `N` misclassified (weakness > 0), `N` marginal (weakness ≤ 0). |
| `unreachable_kpi.txt` | Only written when the recall target is unreachable. Presence of this file means: skip Step 5, write the abridged report, recommend retrain. |

Print the container's stdout summary (chosen threshold, kept-row counts, per-label breakdown) to your own stdout so the script-check hook can verify the run produced output.

### Step 5 — Visual spot check (small, fixed)

Skip this step if `unreachable_kpi.txt` exists. Otherwise use the Read tool to **view** the 5 weakest PASS samples and the 5 weakest NO_PASS samples from `kpi_gaps.parquet` (deduplicated to one row per sample, using the FIRST-lighting `filepath`), classify each as exactly one of **mislabeled** / **edge case** / **data quality** / **systematic**, and copy each viewed image (resized to 128×128 if PIL is available, otherwise just copy) into `<results_dir>/rca_images/`. This is the only image inspection required — do not view dozens of images, run failure mode clustering, or audit goldens (VCN has no golden images).

See `references/visual-spot-check.md` for the exact sample-selection sort, the per-lighting deduplication rule, the full definition of each verdict category, and the image-copy detail.

---

## Reference invocation

Paste-and-edit the workspace, the four paths, and the two numeric knobs; this runs end-to-end. Capture stdout so the script-check hook sees row counts.

```bash
WORKSPACE=<absolute path>            # mounted identically inside the container
EXP_DIR=<experiment_result_dir>      # contains inference/inference.csv and train.yaml; must be inside $WORKSPACE
DATASET_ROOT=<dataset_root>          # image root for inference.csv input_path entries; must be inside $WORKSPACE
MIN_RECALL=1.0                       # zero-miss default; lower if KPI relaxes
TOP_K=50                             # per-label augmentation budget
OUT="$EXP_DIR/rca_results/$(date +%Y-%m-%d_%H%M%S)"
SPEC="$OUT/vcn_aoi_spec.yaml"
IMG=$(python3 -c "import yaml,os; print(yaml.safe_load(open(os.environ['TAO_SKILL_BANK_PATH']+'/versions.yaml'))['images']['tao_toolkit']['data_services'])")

mkdir -p "$OUT"

# Write the gap-analysis spec for this run
cat > "$SPEC" <<EOF
min_recall: $MIN_RECALL
top_k_per_label: $TOP_K
EOF

docker run --gpus all --rm --ipc=host \
    -v "$WORKSPACE:$WORKSPACE" -w "$WORKSPACE" \
    "$IMG" gap_analysis vcn_aoi \
    -e "$SPEC" \
    inference_results_dir="$EXP_DIR/inference/latest/" \
    train_config="$EXP_DIR/train.yaml" \
    kpi_media_path="$DATASET_ROOT" \
    results_dir="$OUT"

# Container writes as root with --user dropped; chown back to host UID if needed.
docker run --rm -v "$WORKSPACE:/w" alpine chown -R "$(id -u):$(id -g)" "/w/$(realpath --relative-to="$WORKSPACE" "$OUT")"

# Sanity print so the script-check hook sees real numbers
python3 - "$OUT" << 'PYEOF'
import json, os, sys
out = sys.argv[1]
unreachable = os.path.join(out, "unreachable_kpi.txt")
if os.path.isfile(unreachable):
    print("KPI UNREACHABLE — see", unreachable)
    sys.exit(0)
with open(os.path.join(out, "threshold.txt")) as f:
    print("threshold:", f.read().strip())
with open(os.path.join(out, "metrics.json")) as f:
    m = json.load(f)
print(f"precision={m['precision']:.4f} recall={m['recall']:.4f} f1={m['f1']:.4f}")
import pandas as pd
df = pd.read_parquet(os.path.join(out, "kpi_gaps.parquet"))
print(f"kpi_gaps.parquet: rows={len(df)}, cols={list(df.columns)}")
print(df['label'].value_counts())
PYEOF
```

---

## Outputs

Write everything into a timestamped folder under the experiment result directory. The container's outputs go straight there; the visual spot-check writes `rca_images/`; any runtime packaging hook may add session/config capture artifacts after `RCA_Report.md` is written.

```
<experiment_result_dir>/rca_results/YYYY-MM-DD_HHMMSS/
├── RCA_Report.md              # Full gap analysis report (you write this)
├── kpi_gaps.parquet           # Container: top-K weakest per label, expanded per lighting
├── threshold.txt              # Container: chosen decision threshold (single float)
├── metrics.json               # Container: confusion matrix + per-label distribution stats
├── weak_samples_breakdown.txt # Container: per-label count/misclassified/marginal counts
├── unreachable_kpi.txt        # Container: ONLY when no threshold meets min_recall
├── rca_images/                # You: thumbnails of the 10 viewed weak samples
├── rca_config/                # Auto-copied by hook
└── session log/artifacts      # Optional, runtime-dependent packaging capture
```

At the start of the run, get the real timestamp by running `date +%Y-%m-%d_%H%M%S` in Bash. Do NOT hardcode or guess. If the user specifies a custom output path, use that instead but maintain the same internal structure.

---

## Common pitfalls

The single most consequential failure mode is **forgetting `top_k_per_label` when `min_recall=1.0`**: at that recall the chosen threshold sits at or below every NO_PASS score, so without `top_k_per_label` the container falls back to a "samples below threshold" filter that returns ONLY misclassified PASS rows and zero NO_PASS rows, breaking the augmentation queue. Always include an explicit positive `top_k_per_label` (default 50) in the spec or as a Hydra override.

See `references/pitfalls.md` for the complete checklist, covering: forgetting `top_k_per_label`; passing `--user`; calling with only Hydra overrides (no `-e <spec>`); spec file outside `$WORKSPACE`; spec file with unresolved `???` sentinels; image not pulled / wrong tag; path-mount mismatch; `unreachable_kpi.txt` written; `inference.csv` missing required columns; train YAML missing `dataset.classify.input_map` or `image_ext`; `kpi_media_path` not matching `input_path` prefixes; and no GPU detected from inside the container.

---

## Report Structure

Write `RCA_Report.md` as a tight (1000–1800 word) computational gap analysis — depth comes from accurate numbers and a clear action list, not narrative. The full report template (7 sections: Verdict, Threshold Selection, Weakness Distribution, Top-K Weakest Samples, Visual Spot Check, Per-Label Breakdown, Recommended Actions — with the confusion-matrix and table layouts) is in `references/output-template.md`. When `unreachable_kpi.txt` exists, replace sections 3–6 with a single short section quoting that file's contents and collapse section 7 to one recommendation: retrain or relabel.

---

## Execution Order

1. Resolve `DS_IMAGE` from `versions.yaml` (`images.tao_toolkit.data_services`), then run `docker info`, `nvidia-smi`, and `docker image inspect "$DS_IMAGE"` (pulling if missing) once to confirm the environment. Abort with a clear message if any fail.
2. Run `date +%Y-%m-%d_%H%M%S` to get the timestamp; create `<experiment_result_dir>/rca_results/<timestamp>/`.
3. Write `vcn_aoi_spec.yaml` into the timestamped dir with `min_recall` and `top_k_per_label` filled in. Keep it under `$WORKSPACE` so the `-e` path resolves inside the container.
4. Run `docker run … "$DS_IMAGE" gap_analysis vcn_aoi -e vcn_aoi_spec.yaml inference_results_dir=… train_config=… kpi_media_path=… output_dir=…`. The container writes `kpi_gaps.parquet`, `threshold.txt`, `metrics.json`, `weak_samples_breakdown.txt` into `results_dir`. Print the chosen threshold and kept-row counts to stdout so the script-check hook can verify the run produced output.
5. If `unreachable_kpi.txt` exists, skip Step 6 and write the abridged report. Otherwise continue.
6. Pick 10 weak samples (5 weakest PASS + 5 weakest NO_PASS) from `kpi_gaps.parquet`, view each test image with Read, classify, and copy each into `rca_images/`.
7. Write `RCA_Report.md` last — writing it triggers the packaging hook, which copies session logs and skill config alongside.

所有檔案

1 個檔案

安裝 tao-analyze-gaps-visual-changenet

請下載並將技能檔案解壓縮至您的 .claude/skills/ 目錄中。

下載 ZIP

複製儲存庫並將技能檔案複製到您的專案中。

git clone https://github.com/NVIDIA/skills/tree/main/skills/tao-analyze-gaps-visual-changenet # Copy SKILL.md to your .claude/skills/ directory

複製 複製
快速設定: 將技能資料夾複製到 .claude/skills/ Claude 會自動偵測並使用該技能
儲存庫 NVIDIA/skills

相關技能

web-search
更新時間 2026-06-29
webapp-testing
更新時間 2026-06-29
lark-base
更新時間 2026-07-05
agentmail
更新時間 2026-06-29
OR