tao-analyze-gaps-visual-changenet
NVIDIA/skills
NVIDIA TAO VCN Classifyの実験において、グラウンドトゥルースラベルごとに最も脆弱なサンプルを特定します。これには、閾値スイープ、脆弱性スコアリング、および照明ごとの拡張を行うDockerコンテナを実行し、下流のデータ拡張や再ラベル付けのために上位K個の脆弱なサンプルを抽出します。
...すべて拡張しますTAO VCN Classify ギャップ分析スキル
あなたは、NVIDIA TAO VCN Classify(Visual Component Net)の推論結果を担当するアナリストです。あなたの仕事は、各 ground-truth ラベルごとに、決定閾値から誤った方向への符号付き距離を測定して最も精度の低いサンプルを特定し、それらを抽出・提示して、後続のデータ拡張や再ラベル付けに活用できるようにすることです。
このスキルは意図的に軽量に設計されています。VCNのclassifyヘッドは単一スコアの二値境界(siamese_scoreによるPASS対NO_PASS)であるため、この分析は調査的なものではなく、計算的なものです。 計算処理全体は、versions.yamlで宣言されたtao_toolkit.data_servicesイメージに対する1回の直接的なDocker実行呼び出しによって行われます(実行時に解決されます — 「セットアップ」を参照)。 コンテナのエントリポイントは ` を受け付け、`gap_analysis vcn_aoi key=value ...` を渡します。 各オーバーライドは、スクリプトのGapAnalysisConfigスキーマを選択的に上書きする、単純な Hydra のkey=value形式です(デフォルト値はコンテナに組み込まれています。docker run ... gap_analysis vcn_aoi --cfg=job で内部を確認してください)。 (コンテナ内には「dataset」キーワードは存在しません。これは TAO ランチャーのピラープレフィックスであり、ここでは省略されています。) 委任分析、多段階のイメージ監査、コンポーネント型クラスタリングは必要ありません。VCN はこれらの次元を公開していません。コンテナが返却された後、代表的な少数の脆弱なサンプルのみを閲覧して、ギャップを特定してください。
CLI インターフェースは、データサービス・コンテナのビルド間で異なる場合があります。gap_analysis vcn_aoiの呼び出しが引数の解析で失敗した場合は、docker run --rm "$DS_IMAGE" gap_analysis vcn_aoiを実行して、イメージごとに 1 回、実際のスキーマを確認してください。--cfg=jobを実行して各イメージごとに実際のスキーマを検証し、再試行前に名前が変更されたキー(例:inference_csv対inference_results_dir、output_dir対results_dir)を調整してください。出力される Parquet ファイル名はkpi_gaps.parquet です。
入力
- 実験結果ディレクトリ— TAO VCN Classifyの推論
結果であるinference/inference.csvが含まれています。必須の列:input_path、object_name、label、siamese_score。 CSVファイルではなく、ディレクトリ(例:inference/latest/)を指定してください。コンテナはinference_results_dir/inference.csv を読み込みます。 - トレーニングコード/設定ディレクトリ— VCN トレーニング用 YAML が含まれています。コンテナは、そこから
dataset.classify.input_map(照明条件リスト)およびdataset.classify.image_extを読み取り、各弱サンプルを照明ごとに 1 行ずつ展開します。 - データセットディレクトリ— 各行(
kpi_media_path)の相対的なinput_pathの先頭に、画像のルートパスが追加されます。 - スキーマのオーバーライド—
min_recall、top_k_per_label、およびオプションで固定された閾値が、Hydraのオーバーライドとして渡されます(デフォルト:min_recall=1.0、top_k_per_label=50、threshold=-1.0(スイープを意味))。top_k_per_label は正の整数でなければなりません。これを省略すると、コンテナは「閾値未満フィルタ」モードに切り替わり、min_recall=1.0の場合、PASS の誤分類のみが返され、NO_PASS 行は 0 行となります。「よくある落とし穴」を参照してください。
セットアップ
しきい値スイープ、弱点ランキング、および照明ごとの拡張は、すべてversions.yaml で宣言されているtao_toolkit.data_servicesイメージ内で実行されます。実行の最初に具体的な URI を一度解決し、Docker、NVIDIA コンテナツールキット、および GPU が存在することを確認し、イメージがキャッシュされていることを確認してください:
# versions.yaml から tao_toolkit.data_services → 具体的な nvcr.io/... URI を解決
DS_IMAGE=$(python3 -c "import yaml,os; print(yaml.safe_load(open(os.environ['TAO_SKILL_BANK_PATH']+'/versions.yaml'))['images']['tao_toolkit']['data_services'])")
echo "DS_IMAGE=$DS_IMAGE"
docker info > /dev/null && echo "OK: docker"
nvidia-smi > /dev/null && echo "OK: GPU"
docker image inspect "$DS_IMAGE" > /dev/null \
|| docker pull "$DS_IMAGE"
TAO_SKILL_BANK_PATH は通常、インストール済みのスキルバンクによってエクスポートされます。設定されていない場合は、解決を行う前にスキルバンクのリポジトリのルートディレクトリを指定してください。GPU が必要です。GPU がないホストでは早期に処理を中止することで、後になって発生する混乱を招くエラーを防ぐことができます。
設定において、負荷がかかりやすく、ミスを起こしやすい3つのルールがあります:
- パスマウント— コンテナが読み書きするすべてのホストパス(
inference.csv、トレーニング用 YAML、データセットイメージのルート、出力ディレクトリ)はバインドマウントする必要があります。最も簡単な方法は、-v $WORKSPACE:$WORKSPACE -w $WORKSPACEを使用し、両側で絶対パスが同一に解決されるようにすることです。 -
--user $(id -u):$(id -g)を指定しないでください— これにより、コンテナのtransformersのインポート中にKeyError: 'getpwuid(): uid not found:が発生します。代わりに、' chownを実行すると、ホストの UID が出力されます。 -e必須であり、オプションではありません。現在のイメージではこれが厳密に要求されており、CLIによる上書きを解析する前に「は ValueError: サブタスク vcn_aoi には次の引数が必要です: -e/--experiment_spec_file」というエラーで終了します。
完全なパスマウントのパターン、--user/chownの根拠およびAlpine のchown コマンド、multi--vに関するガイダンス、および-e要件の詳細については、references/container-setup.md を参照してください。
方法
このスキル全体は、単一の `docker run` 呼び出しと、それに続く簡単な目視によるスポットチェックで構成されています。コンテナは内部でステップ 1~4(閾値スイープ、脆弱性スコアリング、トップ K 選定、照明ごとの展開)を実行します。ステップ 5(目視によるスポットチェック)は、Read ツールを使用して直接処理します。
ステップ1~4 — コンテナの実行
$DOCKER gap_analysis vcn_aoi \
inference_results_dir=/inference/
必ず
top_k_per_label を指定してください。この引数を指定することで、コンテナは デフォルトの「しきい値以下のサンプル」フィルタから、適切なラベルごとのトップK ランキングに切り替わります。min_recall=1.0の場合、しきい値は設計上、すべての NO_PASS スコア以下となるため、しきい値未満のフィルタは誤分類された PASS 行のみを 返し、NO_PASS 行はゼロとなります。これは拡張キューとしては役に立ちません。top_k_per_labelを正の整数に設定すると(仕様書で指定するか、Hydraによるオーバーライドのいずれか)、 コンテナは各行について閾値に対する符号付き弱さを計算し、 ground-truthラベルごとに最も弱いK行を抽出します。これが、下流のステップで消費される ラベルごとのランク付け済み出力となります。
inference.csvを読み込み、すべての固有のsiamese_scoreに加え、最小値のすぐ下にある値を1つずつ走査し、NO_PASSクラスのリコールがmin_recall以上(許容誤差1e-12)の候補を保持し、F1スコアが最も高い閾値を選択します(同値の場合は、精度、次に閾値の順で決定します)。 各行について、その閾値からの弱さの符号付き値を計算する(正=誤分類、負=正答、絶対値=マージン)。 weaknessの降順でソートし、各ground-truthラベルごとにtop_k_per_labelの候補を抽出します。その後、トレーニング用YAMLファイル内のdataset.classify.input_map およびdataset.classify.image_extを使用して、各weakness行を照明条件ごとに1行ずつ展開します。
候補となる閾値のいずれもリコール目標を満たさない場合、コンテナはゼロ以外の値を返して終了し、results_dirにunreachable_kpi.txt を書き込んで、モデルが実際に達成可能なリコールを説明します。 その場合は、Dockerの呼び出し後に分析を中止し、モデルがいかなる動作点においても根本的にKPIに到達できないことを説明する1セクションのレポートを作成し、再学習または再ラベリングを推奨します。視覚的なスポットチェックはスキップします。
コンテナはresults_dir に以下を書き込みます:
| アーティファクト | 内容 |
|---|---|
kpi_gaps.parquet |
ラベルごとの最弱トップK件を、照明ごとに展開したもの。列:filepath、label、siamese_score、weakness。 |
threshold.txt |
選択された決定閾値(単一の浮動小数点数、プレーンテキスト)。 |
metrics.json |
選択された閾値における、精度、再現率、F1スコア、混同行列{tp, fp, tn, fn}、およびラベルごとの{total, mean_weakness, median_weakness, max_weakness, n_misclassified}。 |
weak_samples_breakdown.txt |
ラベルごとの保持行の内訳: 合計、 <%> 保持された全行のうち、誤分類された行数(弱さ > 0)、境界線上の行数(弱さ ≤ 0)。 |
unreachable_kpi.txt |
リコール目標が達成不可能な場合にのみ書き込まれます。このファイルが存在する場合の意味:ステップ5をスキップし、要約レポートを作成し、再トレーニングを推奨します。 |
コンテナの標準出力サマリー(選択された閾値、保持行数、ラベルごとの内訳)を自身の標準出力に書き出し、script-checkフックが実行結果の出力を検証できるようにする。
ステップ 5 — 目視によるスポットチェック(小規模かつ固定)
unreachable_kpi.txtが存在する場合は、このステップをスキップします。そうでない場合は、Read ツールを使用してkpi_gaps.parquetから最も弱い PASS サンプル 5 件と最も弱い NO_PASS サンプル 5件を表示し(FIRST-lightingファイルパスを使用して、サンプルごとに 1 行に重複排除済み)、 それぞれを「誤ラベル付け」「エッジケース」「データ品質」「系統的」のいずれか1つに正確に分類し、閲覧した各画像(PILが利用可能な場合は128×128にリサイズ、そうでない場合はそのままコピー)をにコピーしてください。 必要な画像確認はこれだけです。数十枚の画像を閲覧したり、故障モードのクラスタリングを実行したり、ゴールデン画像を監査したりしないでください(VCNにはゴールデン画像がありません)。
正確なサンプル選択の順序、照明ごとの重複排除ルール、各判定カテゴリの完全な定義、および画像コピーの詳細については、references/visual-spot-check.md を参照してください。
参照の呼び出し
ワークスペース、4つのパス、および2つの数値パラメータをコピーして編集してください。これによりエンドツーエンドで実行されます。スクリプトチェックフックが行数を認識できるよう、標準出力をキャプチャしてください。
WORKSPACE= # コンテナ内で同一にマウントされる
EXP_DIR= # inference/inference.csv および train.yaml を含む;$WORKSPACE 内に存在しなければならない
DATASET_ROOT= # inference.csvのinput_pathエントリ用の画像ルート;$WORKSPACE内にある必要があります
MIN_RECALL=1.0 # デフォルトはゼロミス;KPIの要件が緩い場合はこの値を小さくしてください
TOP_K=50 # ラベルごとのデータ拡張予算
OUT="$EXP_DIR/rca_results/$(date +%Y-%m-%d_%H%M%S)"
SPEC="$OUT/vcn_aoi_spec.yaml"
IMG=$(python3 -c "import yaml,os; print(yaml.safe_load(open(os.environ['TAO_SKILL_BANK_PATH']+'/versions.yaml'))['images']['tao_toolkit']['data_services'])")
mkdir -p "$OUT"
# 今回の実行用のギャップ分析仕様を記述
cat > "$SPEC" <<EOF
min_recall: $MIN_RECALL
top_k_per_label: $TOP_K
EOF
docker run --gpus all --rm --ipc=host \
-v "$WORKSPACE:$WORKSPACE" -w "$WORKSPACE" \
"$IMG" gap_analysis vcn_aoi \
-e "$SPEC" \
inference_results_dir="$EXP_DIR/inference/latest/" \
train_config="$EXP_DIR/train.yaml" \
kpi_media_path="$DATASET_ROOT" \
results_dir="$OUT"
# コンテナは --user オプションを無効にして root 権限で書き込みを行うため、必要に応じてホストの UID に所有権を戻す。
docker run --rm -v "$WORKSPACE:/w" alpine chown -R "$(id -u):$(id -g)" "/w/$(realpath --relative-to="$WORKSPACE" "$OUT")"
# スクリプトチェックフックが実際の数値を認識できるよう、確認用の出力を表示
python3 - "$OUT" << 'PYEOF'
import json, os, sys
out = sys.argv[1]
unreachable = os.path.join(out, "unreachable_kpi.txt")
if os.path.isfile(unreachable):
print("KPI UNREACHABLE — see", unreachable)
sys.exit(0)
with open(os.path.join(out, "threshold.txt")) as f:
print("threshold:", f.read().strip())
with open(os.path.join(out, "metrics.json")) as f:
m = json.load(f)
print(f"precision={m['precision']:.4f} recall={m['recall']:.4f} f1={m['f1']:.4f}")
import pandas as pd
df = pd.read_parquet(os.path.join(out, "kpi_gaps.parquet"))
print(f"kpi_gaps.parquet: rows={len(df)}, cols={list(df.columns)}")
print(df['label'].value_counts())
PYEOF
出力
すべての実験結果を、実験結果ディレクトリ内のタイムスタンプ付きフォルダに書き出します。コンテナの出力は直接そこに書き込まれます。視覚的なスポットチェックの結果はrca_images/ に書き込まれます。実行時のパッケージングフックは、RCA_Report.mdの書き込み後に session/config キャプチャアーティファクトを追加する場合があります。
/rca_results/YYYY-MM-DD_HHMMSS/
├── RCA_Report.md # 完全なギャップ分析レポート(作成者による)
├── kpi_gaps.parquet # コンテナ:ラベルごとの上位K件の最も弱いサンプル、ライティングごとに展開
├── threshold.txt # コンテナ:選択された決定閾値(単一の浮動小数点数)
├── metrics.json # コンテナ:混同行列+ラベルごとの分布統計
├── weak_samples_breakdown.txt # コンテナ:ラベルごとのカウント/誤分類数/マージナルカウント
├── unreachable_kpi.txt # コンテナ:min_recallを満たすしきい値が存在しない場合のみ
├── rca_images/ # ユーザーによる作成:閲覧された10個の弱いサンプルのサムネイル
├── rca_config/ # フックによって自動コピーされる
└── session log/artifacts # オプション、実行環境に依存するパッケージング情報の取得
実行開始時に、Bashで`date +%Y-%m-%d_%H%M%S` を実行して実際のタイムスタンプを取得してください。ハードコーディングや推測は絶対に行わないでください。ユーザーがカスタム出力パスを指定した場合は、そのパスを使用してください。ただし、内部構造は同じまま維持してください。
よくある落とし穴
min_recall=1.0の場合にtop_k_per_label を忘れることが、最も重大な失敗モードです: そのリコール値では、選択された閾値がすべての NO_PASS スコア以下となるため、top_k_per_label を指定しないと、コンテナは「閾値未満のサンプル」というフィルタにフォールバックし、誤分類された PASS 行のみを返し、NO_PASS 行を一切返さなくなるため、オーグメンテーションキューが破綻します。 仕様書またはHydraのオーバーライドとして、常に明示的な正例用top_k_per_label(デフォルトは50)を含めてください。
完全なチェックリストについては、references/pitfalls.md を参照してください。これには、top_k_per_label の指定漏れ、--user オプションの指定、Hydra オーバーライドのみでの呼び出し(-e指定なし)、$WORKSPACE 以外の場所に置かれた仕様ファイル、未解決の???センチネルを含む仕様ファイル、イメージの取得失敗/タグの不一致、 パスマウントの不一致;unreachable_kpi.txt が書き込まれている;inference.csvに必要な列が欠落している;train YAMLに dataset.classify.input_mapまたはimage_ext が欠落している;kpi_media_path が input_pathのプレフィックスと一致しない;およびコンテナ内部から GPU が検出されない。
レポートの構成
RCA_Report.md を、簡潔な(1000~1800語)計算上のギャップ分析として作成してください。深みは、物語的な記述ではなく、正確な数値と明確なアクションリストから生まれます。 完全なレポートテンプレート(7つのセクション:結論、閾値の選択、弱点の分布、上位K個の最も弱いサンプル、視覚的スポットチェック、ラベルごとの内訳、推奨されるアクション — 混同行列および表のレイアウトを含む)は、references/output-template.md にあります。unreachable_kpi.txtが存在する場合は、セクション 3~6 を、そのファイルの内容を引用した単一の短いセクションに置き換え、セクション 7 を「再学習」または「再ラベル付け」という 1 つの推奨事項に集約してください。
実行順序
versions.yaml(images.tao_toolkit.data_services)からDS_IMAGE を解決し、docker info、nvidia-smi、およびdocker image inspect "$DS_IMAGE"(存在しない場合はプルする)を一度実行して環境を確認します。いずれかが失敗した場合は、明確なメッセージを表示して中止します。- `
date +%Y-%m-%d_%H%M%S` を実行してタイムスタンプを取得し、`を作成します。`、`/rca_results/`、` ` vcn_aoi_spec.yaml を、min_recallおよびtop_k_per_labelを記入した上で、タイムスタンプ付きのディレクトリに書き出します。コンテナ内で-eパスが解決されるように、$WORKSPACE配下に配置してください。docker run … "$DS_IMAGE" gap_analysis vcn_aoi -e vcn_aoi_spec.yaml inference_results_dir=… train_config=… kpi_media_path=… output_dir=…を実行します。コンテナは、kpi_gaps.parquet、threshold.txt、metrics.json、weak_samples_breakdown.txt をresults_dirに書き込みます。選択された閾値と保持された行数を標準出力に出力し、script-check フックが実行結果の出力を検証できるようにします。unreachable_kpi.txtが存在する場合は、ステップ 6 をスキップして要約レポートを出力します。そうでない場合は、処理を続行します。kpi_gaps.parquetから 10 個の弱いサンプル(最も弱い PASS 5 個 + 最も弱い NO_PASS 5 個)を選び、Read で各テスト画像を表示して分類し、それぞれをrca_images/にコピーします。- 最後に
RCA_Report.mdを書き出します。これを書き出すとパッケージングフックがトリガーされ、セッションログとスキル設定が併せてコピーされます。
---
name: tao-analyze-gaps-visual-changenet
description: Identifies the weakest samples per ground-truth label in NVIDIA TAO VCN Classify experiments by running a Docker container that performs threshold sweep, weakness scoring, and per-lighting expansion, then surfaces top-K weak samples for downstream augmentation or relabeling.
license: Apache-2.0
---
# TAO VCN Classify Gap Analysis Skill
You are an analyst for NVIDIA TAO VCN Classify (Visual Component Net) inference results. Your job is to identify the **weakest samples per ground-truth label** by measuring signed distance from the decision threshold *in the wrong direction*, then surface them for downstream augmentation or relabeling.
This skill is intentionally lightweight. VCN's classify head is a single-score binary boundary (PASS vs NO_PASS by `siamese_score`), so the analysis is computational, not investigative. The whole computation lives behind one direct `docker run` invocation against the `tao_toolkit.data_services` image declared in `versions.yaml` (resolved at runtime — see Setup). The container's entrypoint takes `<category> <action> [hydra overrides...]`; we pass `gap_analysis vcn_aoi key=value …`. Each override is a bare Hydra `key=value` that selectively overrides the script's `GapAnalysisConfig` schema (defaults are baked into the container; introspect with `docker run ... gap_analysis vcn_aoi --cfg=job`). (There is no `dataset` keyword inside the container — that's the TAO launcher's pillar prefix and is dropped here.) You do **not** need delegated analysis, multi-phase image audits, or component-type clustering — VCN does not expose those dimensions. View only a small set of representative weak samples to qualify the gaps after the container returns.
CLI surface can shift between data-services container builds. If a `gap_analysis vcn_aoi` invocation fails on argument parsing, introspect the actual schema once per image with `docker run --rm "$DS_IMAGE" gap_analysis vcn_aoi --cfg=job` and reconcile any renamed keys (e.g. `inference_csv` vs `inference_results_dir`, `output_dir` vs `results_dir`) before retrying. Output parquet name is `kpi_gaps.parquet`.
---
## Inputs
1. **Experiment result directory** — contains `inference/inference.csv` from TAO VCN Classify inference. Required columns: `input_path`, `object_name`, `label`, `siamese_score`. Pass the **directory** (e.g. `inference/latest/`), not the CSV file — the container reads `inference_results_dir/inference.csv`.
2. **Training code/config directory** — contains the VCN train YAML. The container reads `dataset.classify.input_map` (lighting condition list) and `dataset.classify.image_ext` from it to expand each weak sample into one row per lighting.
3. **Dataset directory** — image root prepended to the relative `input_path` from each row (`kpi_media_path`).
4. **Schema overrides** — `min_recall`, `top_k_per_label`, and optionally a hard-pinned `threshold` are passed as Hydra overrides (defaults: `min_recall=1.0`, `top_k_per_label=50`, `threshold=-1.0` meaning sweep). **`top_k_per_label` must be a positive integer** — omitting it flips the container into "below-threshold filter" mode, which at `min_recall=1.0` returns only PASS misclassifications and zero NO_PASS rows. See Common pitfalls.
---
## Setup
The threshold sweep, weakness ranking, and per-lighting expansion all run inside the `tao_toolkit.data_services` image declared in `versions.yaml`. Resolve the concrete URI once at the top of the run, then confirm Docker, the NVIDIA container toolkit, and a GPU are present and ensure the image is cached:
```bash
# Resolve tao_toolkit.data_services → concrete nvcr.io/... URI from versions.yaml
DS_IMAGE=$(python3 -c "import yaml,os; print(yaml.safe_load(open(os.environ['TAO_SKILL_BANK_PATH']+'/versions.yaml'))['images']['tao_toolkit']['data_services'])")
echo "DS_IMAGE=$DS_IMAGE"
docker info > /dev/null && echo "OK: docker"
nvidia-smi > /dev/null && echo "OK: GPU"
docker image inspect "$DS_IMAGE" > /dev/null \
|| docker pull "$DS_IMAGE"
```
`TAO_SKILL_BANK_PATH` is usually exported by the installed skill bank. If it is unset, point it at the skill-bank repo root before resolving. A GPU is required; aborting early on a GPU-less host saves a confusing late error.
Three setup rules are load-bearing and easy to get wrong:
- **Path mounting** — every host path the container reads or writes (`inference.csv`, train YAML, dataset image root, output dir) must be bind-mounted, simplest with `-v $WORKSPACE:$WORKSPACE -w $WORKSPACE` so absolute paths resolve identically on both sides.
- **Do not pass `--user $(id -u):$(id -g)`** — it triggers `KeyError: 'getpwuid(): uid not found: <uid>'` during the container's `transformers` import; `chown` outputs back to the host UID afterwards instead.
- **`-e <spec>` is required, not optional** — current images hard-require it and exit with `ValueError: The subtask vcn_aoi requires the following argument: -e/--experiment_spec_file` before parsing CLI overrides.
See `references/container-setup.md` for the full path-mounting pattern, the `--user`/`chown` rationale and `alpine` chown command, multi-`-v` guidance, and the `-e <spec>` requirement detail.
---
## Method
The whole skill is a single `docker run` invocation followed by a small visual spot-check. The container does Steps 1–4 internally (threshold sweep, weakness scoring, top-K selection, per-lighting expansion). You handle Step 5 (visual spot-check) directly with the Read tool.
### Step 1–4 — Run the container
```bash
$DOCKER gap_analysis vcn_aoi \
inference_results_dir=<exp_dir>/inference/<label>/ \
train_config=<exp_dir>/train.yaml \
kpi_media_path=<dataset_root> \
results_dir=<rca_results_dir> \
top_k_per_label=50
```
> **Always pass `top_k_per_label`.** This is the argument that switches the container
> from the default "samples below threshold" filter into proper top-K-per-label
> ranking. At `min_recall=1.0` the threshold is by construction at-or-below every
> NO_PASS score, so the below-threshold filter returns ONLY misclassified PASS rows
> and zero NO_PASS rows — useless as an augmentation queue. With `top_k_per_label`
> set to a positive integer (either in the spec or as a Hydra override), the
> container computes signed weakness against the threshold for every row and
> surfaces the K weakest **per ground-truth label**, which is the per-label ranked
> output downstream steps consume.
Reads `inference.csv`, sweeps every unique `siamese_score` plus one value just below the minimum, keeps the candidates with NO_PASS-class recall ≥ `min_recall` (with `1e-12` tolerance), then picks the threshold with the best F1 (tie-break: precision, then threshold value). For every row, computes signed weakness from that threshold (positive = misclassified, negative = correct, magnitude = margin). Sorts by weakness descending and takes the top `top_k_per_label` per ground-truth label, then expands each weak row into one row per lighting condition using `dataset.classify.input_map` and `dataset.classify.image_ext` from the train YAML.
If **no** candidate threshold meets the recall target, the container exits non-zero and writes `unreachable_kpi.txt` into `results_dir` explaining which recall the model can actually achieve. In that case, stop the analysis after the docker call, write a one-section report explaining the model fundamentally cannot reach the KPI at any operating point, and recommend retraining or relabeling — skip the visual spot-check.
**Container writes into `results_dir`:**
| Artifact | Contents |
|----------|----------|
| `kpi_gaps.parquet` | Top-K weakest per label, expanded per lighting. Columns: `filepath`, `label`, `siamese_score`, `weakness`. |
| `threshold.txt` | Chosen decision threshold (single float, plain text). |
| `metrics.json` | At the chosen threshold: `precision`, `recall`, `f1`, confusion matrix `{tp, fp, tn, fn}`, plus per-label `{total, mean_weakness, median_weakness, max_weakness, n_misclassified}`. |
| `weak_samples_breakdown.txt` | Per-label kept-row breakdown: `<count>` total, `<%>` of all kept rows, `N` misclassified (weakness > 0), `N` marginal (weakness ≤ 0). |
| `unreachable_kpi.txt` | Only written when the recall target is unreachable. Presence of this file means: skip Step 5, write the abridged report, recommend retrain. |
Print the container's stdout summary (chosen threshold, kept-row counts, per-label breakdown) to your own stdout so the script-check hook can verify the run produced output.
### Step 5 — Visual spot check (small, fixed)
Skip this step if `unreachable_kpi.txt` exists. Otherwise use the Read tool to **view** the 5 weakest PASS samples and the 5 weakest NO_PASS samples from `kpi_gaps.parquet` (deduplicated to one row per sample, using the FIRST-lighting `filepath`), classify each as exactly one of **mislabeled** / **edge case** / **data quality** / **systematic**, and copy each viewed image (resized to 128×128 if PIL is available, otherwise just copy) into `<results_dir>/rca_images/`. This is the only image inspection required — do not view dozens of images, run failure mode clustering, or audit goldens (VCN has no golden images).
See `references/visual-spot-check.md` for the exact sample-selection sort, the per-lighting deduplication rule, the full definition of each verdict category, and the image-copy detail.
---
## Reference invocation
Paste-and-edit the workspace, the four paths, and the two numeric knobs; this runs end-to-end. Capture stdout so the script-check hook sees row counts.
```bash
WORKSPACE=<absolute path> # mounted identically inside the container
EXP_DIR=<experiment_result_dir> # contains inference/inference.csv and train.yaml; must be inside $WORKSPACE
DATASET_ROOT=<dataset_root> # image root for inference.csv input_path entries; must be inside $WORKSPACE
MIN_RECALL=1.0 # zero-miss default; lower if KPI relaxes
TOP_K=50 # per-label augmentation budget
OUT="$EXP_DIR/rca_results/$(date +%Y-%m-%d_%H%M%S)"
SPEC="$OUT/vcn_aoi_spec.yaml"
IMG=$(python3 -c "import yaml,os; print(yaml.safe_load(open(os.environ['TAO_SKILL_BANK_PATH']+'/versions.yaml'))['images']['tao_toolkit']['data_services'])")
mkdir -p "$OUT"
# Write the gap-analysis spec for this run
cat > "$SPEC" <<EOF
min_recall: $MIN_RECALL
top_k_per_label: $TOP_K
EOF
docker run --gpus all --rm --ipc=host \
-v "$WORKSPACE:$WORKSPACE" -w "$WORKSPACE" \
"$IMG" gap_analysis vcn_aoi \
-e "$SPEC" \
inference_results_dir="$EXP_DIR/inference/latest/" \
train_config="$EXP_DIR/train.yaml" \
kpi_media_path="$DATASET_ROOT" \
results_dir="$OUT"
# Container writes as root with --user dropped; chown back to host UID if needed.
docker run --rm -v "$WORKSPACE:/w" alpine chown -R "$(id -u):$(id -g)" "/w/$(realpath --relative-to="$WORKSPACE" "$OUT")"
# Sanity print so the script-check hook sees real numbers
python3 - "$OUT" << 'PYEOF'
import json, os, sys
out = sys.argv[1]
unreachable = os.path.join(out, "unreachable_kpi.txt")
if os.path.isfile(unreachable):
print("KPI UNREACHABLE — see", unreachable)
sys.exit(0)
with open(os.path.join(out, "threshold.txt")) as f:
print("threshold:", f.read().strip())
with open(os.path.join(out, "metrics.json")) as f:
m = json.load(f)
print(f"precision={m['precision']:.4f} recall={m['recall']:.4f} f1={m['f1']:.4f}")
import pandas as pd
df = pd.read_parquet(os.path.join(out, "kpi_gaps.parquet"))
print(f"kpi_gaps.parquet: rows={len(df)}, cols={list(df.columns)}")
print(df['label'].value_counts())
PYEOF
```
---
## Outputs
Write everything into a timestamped folder under the experiment result directory. The container's outputs go straight there; the visual spot-check writes `rca_images/`; any runtime packaging hook may add session/config capture artifacts after `RCA_Report.md` is written.
```
<experiment_result_dir>/rca_results/YYYY-MM-DD_HHMMSS/
├── RCA_Report.md # Full gap analysis report (you write this)
├── kpi_gaps.parquet # Container: top-K weakest per label, expanded per lighting
├── threshold.txt # Container: chosen decision threshold (single float)
├── metrics.json # Container: confusion matrix + per-label distribution stats
├── weak_samples_breakdown.txt # Container: per-label count/misclassified/marginal counts
├── unreachable_kpi.txt # Container: ONLY when no threshold meets min_recall
├── rca_images/ # You: thumbnails of the 10 viewed weak samples
├── rca_config/ # Auto-copied by hook
└── session log/artifacts # Optional, runtime-dependent packaging capture
```
At the start of the run, get the real timestamp by running `date +%Y-%m-%d_%H%M%S` in Bash. Do NOT hardcode or guess. If the user specifies a custom output path, use that instead but maintain the same internal structure.
---
## Common pitfalls
The single most consequential failure mode is **forgetting `top_k_per_label` when `min_recall=1.0`**: at that recall the chosen threshold sits at or below every NO_PASS score, so without `top_k_per_label` the container falls back to a "samples below threshold" filter that returns ONLY misclassified PASS rows and zero NO_PASS rows, breaking the augmentation queue. Always include an explicit positive `top_k_per_label` (default 50) in the spec or as a Hydra override.
See `references/pitfalls.md` for the complete checklist, covering: forgetting `top_k_per_label`; passing `--user`; calling with only Hydra overrides (no `-e <spec>`); spec file outside `$WORKSPACE`; spec file with unresolved `???` sentinels; image not pulled / wrong tag; path-mount mismatch; `unreachable_kpi.txt` written; `inference.csv` missing required columns; train YAML missing `dataset.classify.input_map` or `image_ext`; `kpi_media_path` not matching `input_path` prefixes; and no GPU detected from inside the container.
---
## Report Structure
Write `RCA_Report.md` as a tight (1000–1800 word) computational gap analysis — depth comes from accurate numbers and a clear action list, not narrative. The full report template (7 sections: Verdict, Threshold Selection, Weakness Distribution, Top-K Weakest Samples, Visual Spot Check, Per-Label Breakdown, Recommended Actions — with the confusion-matrix and table layouts) is in `references/output-template.md`. When `unreachable_kpi.txt` exists, replace sections 3–6 with a single short section quoting that file's contents and collapse section 7 to one recommendation: retrain or relabel.
---
## Execution Order
1. Resolve `DS_IMAGE` from `versions.yaml` (`images.tao_toolkit.data_services`), then run `docker info`, `nvidia-smi`, and `docker image inspect "$DS_IMAGE"` (pulling if missing) once to confirm the environment. Abort with a clear message if any fail.
2. Run `date +%Y-%m-%d_%H%M%S` to get the timestamp; create `<experiment_result_dir>/rca_results/<timestamp>/`.
3. Write `vcn_aoi_spec.yaml` into the timestamped dir with `min_recall` and `top_k_per_label` filled in. Keep it under `$WORKSPACE` so the `-e` path resolves inside the container.
4. Run `docker run … "$DS_IMAGE" gap_analysis vcn_aoi -e vcn_aoi_spec.yaml inference_results_dir=… train_config=… kpi_media_path=… output_dir=…`. The container writes `kpi_gaps.parquet`, `threshold.txt`, `metrics.json`, `weak_samples_breakdown.txt` into `results_dir`. Print the chosen threshold and kept-row counts to stdout so the script-check hook can verify the run produced output.
5. If `unreachable_kpi.txt` exists, skip Step 6 and write the abridged report. Otherwise continue.
6. Pick 10 weak samples (5 weakest PASS + 5 weakest NO_PASS) from `kpi_gaps.parquet`, view each test image with Read, classify, and copy each into `rca_images/`.
7. Write `RCA_Report.md` last — writing it triggers the packaging hook, which copies session logs and skill config alongside.
すべてのファイル
1件のファイルtao-analyze-gaps-visual-changenetをインストール
スキルファイルをダウンロードし、.claude/skills/ ディレクトリに解凍してください。
ZIPをダウンロードリポジトリをクローンし、スキルファイルをプロジェクトにコピーしてください。
git clone https://github.com/NVIDIA/skills/tree/main/skills/tao-analyze-gaps-visual-changenet # Copy SKILL.md to your .claude/skills/ directory
コピー





家
