选项
首页首页 Skill 数据科学与机器学习 tao-analyze-gaps-visual-changenet

tao-analyze-gaps-visual-changenet

NVIDIA/skills NVIDIA/skills

通过运行一个 Docker 容器,该容器执行阈值扫描、弱样本评分和按光照条件扩展操作,从而在 NVIDIA TAO VCN Classify 实验中,针对每个真实标签识别出最弱的样本,并筛选出前 K 个弱样本,供下游增强或重新标注使用。

...展开全部
2
更新时间 2026-09-28

TAO VCN 分类差距分析技能

您是 NVIDIA TAO VCN Classify(视觉组件网络)推理结果的分析师。您的工作是根据每个真实标签,通过测量样本与决策阈值之间在错误方向上的带号距离,识别出最弱的样本,然后将其筛选出来,以便进行后续的数据增强或重新标注。

该技能设计上刻意保持轻量级。VCN 的分类头是一个单分数二元边界(通过siamese_score 判断 PASS 还是 NO_PASS),因此分析过程侧重于计算,而非调查。 整个计算过程通过一次直接调用Docker 运行实现,该调用针对versions.yaml中声明的tao_toolkit.data_services镜像(运行时解析——参见“设置”部分)。 容器的入口点接受 [hydra 覆盖项...];我们传入gap_analysis vcn_aoi key=value …。 每个覆盖项都是一个简单的 Hydrakey=value对,用于有选择地覆盖脚本的GapAnalysisConfig架构(默认值已内置于容器中;可通过docker run ... gap_analysis vcn_aoi --cfg=job 进行检查)。 (容器内部不存在“dataset”关键字——这是 TAO 启动器的支柱前缀,此处已被省略。) 您无需委托分析、多阶段镜像审计或组件类型聚类——VCN 不会暴露这些维度。只需查看一小部分具有代表性的弱样本,即可在容器返回后确认差距。

CLI 界面可能在不同数据服务容器的构建版本之间切换。如果调用gap_analysis vcn_aoi时因参数解析失败,请使用docker run --rm "$DS_IMAGE" gap_analysis vcn_aoi命令,针对每个镜像检查一次实际的架构,--cfg=job命令对每个镜像执行一次实际架构检查,并在重试前调整任何重命名的键(例如inference_csv与inference_results_dir、output_dir与results_dir)。Parquet 输出文件名为kpi_gaps.parquet。

输入

  1. 实验结果目录— 包含来自 TAO VCN Classify 推理的`inference/inference.csv` 文件。必备字段:input_path、object_name、label、siamese_score。 请传入目录(例如inference/latest/),而非 CSV 文件——容器会从inference_results_dir/inference.csv 中读取数据。
  2. 训练代码/配置目录— 包含 VCN 训练的 YAML 文件。容器会从中读取dataset.classify.input_map(光照条件列表)和dataset.classify.image_ext,将每个弱样本按光照条件扩展为一行。
  3. 数据集目录——将图像根目录附加到每行(kpi_media_path)的相对input_path前面。
  4. 模式覆盖项—min_recall、top_k_per_label 以及可选的硬编码阈值将作为 Hydra 覆盖参数传递(默认值:min_recall=1.0 ,top_k_per_label=50,threshold=-1.0表示进行阈值扫描)。top_k_per_label必须为正整数— 省略该参数会将容器切换为“低于阈值过滤”模式,当min_recall=1.0时,该模式仅返回 PASS 误分类结果,且不返回任何 NO_PASS 行。参见“常见陷阱”。

设置

阈值扫描、弱点排序和按照明扩展均在versions.yaml 中声明的tao_toolkit.data_services镜像内运行。在运行开始时解析一次具体的 URI,然后确认 Docker、NVIDIA 容器工具包和 GPU 是否存在,并确保镜像已缓存:

# 从 versions.yaml 中解析 tao_toolkit.data_services → 具体的 nvcr.io/... URI
DS_IMAGE=$(python3 -c "import yaml,os; print(yaml.safe_load(open(os.environ['TAO_SKILL_BANK_PATH']+'/versions.yaml'))['images']['tao_toolkit']['data_services'])")
echo "DS_IMAGE=$DS_IMAGE"

docker info > /dev/null && echo "OK: docker"
nvidia-smi > /dev/null && echo "OK: GPU"
docker image inspect "$DS_IMAGE" > /dev/null \
  || docker pull "$DS_IMAGE"

TAO_SKILL_BANK_PATH通常由已安装的技能库导出。如果该变量未设置,请在解析之前将其指向技能库仓库的根目录。必须配备 GPU;在无 GPU 的主机上尽早中止操作,可避免后期出现令人困惑的错误。

以下三条配置规则至关重要,且容易出错:

  • 路径挂载—— 容器读写的所有主机路径(inference.csv、训练 YAML、数据集镜像根目录、输出目录)都必须进行绑定挂载,最简单的方法是使用-v $WORKSPACE:$WORKSPACE -w $WORKSPACE,这样绝对路径在双方的解析结果将完全一致。
  • 请勿传递`--user $(id -u):$(id -g)` —— 这会在容器导入transformers时触发`KeyError: 'getpwuid(): uid not found:'`错误;随后`chown` 命令会将所有权归还给主机 UID。
  • -e是必需的,而非可选的— 当前镜像严格要求该参数,并在解析 CLI 覆盖设置之前会因ValueError: 子任务 vcn_aoi 需要以下参数:-e/--experiment_spec_file而退出。

请参阅references/container-setup.md以了解完整的路径挂载模式、--user/chown的设计原理及Alpine 系统的chown 命令、multi--v相关指南,以及-e要求的详细说明。

方法

整个技能仅包含一次docker run调用,随后进行一次简短的视觉抽查。容器会在内部执行步骤 1–4(阈值扫描、弱点评分、Top-K 筛选、按光照条件扩展)。您需直接使用 Read 工具处理步骤 5(视觉抽查)。

步骤 1–4 — 运行容器

$DOCKER gap_analysis vcn_aoi \
    inference_results_dir=/inference/

请务必传入top_k_per_label 参数。该参数可将容器 从默认的“低于阈值的样本”过滤模式切换为正确的按标签 Top-K 排序模式。 当min_recall=1.0时,阈值按设计位于每个 NO_PASS 得分之下或等于该得分,因此“低于阈值”过滤器仅返回误分类的 PASS 行 且不返回任何 NO_PASS 行——作为增强队列毫无用处。 当top_k_per_label 设置为正整数(无论是在规范中还是作为 Hydra 覆盖参数)时, 容器会针对每一行计算相对于阈值的有符号弱点,并 筛选出每个 ground-truth 标签下最弱的 K 行,这便是下游步骤所使用的 按标签排序的输出。

读取inference.csv 文件,遍历每个唯一的siamese_score 值以及略低于最小值的下一个数值,保留 NO_PASS 类召回率 ≥min_recall(容差为1e-12)的候选项,然后选择 F1 值最佳的阈值(平局时按精度优先,其次为阈值数值的顺序)。 对于每一行,根据该阈值计算带符号的弱点(正值 = 误分类,负值 = 正确,绝对值 = 容差)。 按弱点值降序排序,并为每个真实标签选取前top_k_per_label个候选阈值,随后利用训练 YAML 文件中的dataset.classify.input_map和dataset.classify.image_ext,将每行弱点数据扩展为按光照条件分列的行。

如果没有候选阈值能满足召回率目标,则容器以非零状态退出,并在results_dir中写入unreachable_kpi.txt 文件,说明模型实际能达到的召回率。 在这种情况下,在调用 Docker 后停止分析,编写一份单节报告说明模型在任何工作点上都根本无法达到 KPI,并建议重新训练或重新标注——跳过视觉抽查。

容器将以下内容写入results_dir:

构建产物 内容
kpi_gaps.parquet 每个标签下最弱的 Top-K 结果,按光照条件展开。列:filepath、label、siamese_score、weakness。
threshold.txt 选定的决策阈值(单个浮点数,纯文本)。
metrics.json 在选定的阈值下:精确率、召回率、F1 值、混淆矩阵{真阳性, 假阳性, 真阴性, 假阴性},以及按标签划分的{总数, 平均弱点值, 中位数弱点值, 最大弱点值, 误分类数}。
weak_samples_breakdown.txt 按标签划分的保留行明细: 总数, <%> 在所有保留行中,误分类数(弱点 > 0)为 N,临界行数(弱点 ≤ 0)为 N。
unreachable_kpi.txt 仅在召回率目标无法达成时生成。该文件的存在意味着:跳过步骤 5,生成简要报告,并建议重新训练。

将容器的标准输出摘要(选定的阈值、保留行数、按标签细分的统计数据)打印到您的标准输出中,以便脚本检查钩子能够验证运行生成的输出。

步骤 5 — 可视化抽查(小规模、固定)

若存在unreachable_kpi.txt 文件,则跳过此步骤。否则,使用 Read 工具查看 kpi_gaps.parquet文件中 5 个最弱的“PASS”样本和 5 个最弱的“NO_PASS”样本(已去重为每个样本一行,使用 FIRST-lighting文件路径), 将每个样本精确归类为“标签错误”、“边界案例”、“数据质量 ”或“系统性”其中之一,并将每个查看过的图像(如有 PIL 则调整为 128×128 尺寸,否则直接复制)复制到/rca_images/ 中。 这是唯一需要的图像检查工作——请勿查看数十张图像、运行故障模式聚类分析,或审核黄金样本(VCN 没有黄金样本)。

有关确切的样本选择排序、按光照条件的去重规则、各判定类别的完整定义以及图像复制详情,请参阅references/visual-spot-check.md。

参考调用

将工作区、四个路径和两个数值参数粘贴并编辑;这将执行端到端测试。捕获标准输出,以便脚本检查钩子能获取行数。

WORKSPACE=           # 在容器内以相同方式挂载
EXP_DIR=     # 包含 inference/inference.csv 和 train.yaml;必须位于 $WORKSPACE 内
DATASET_ROOT=         # 用于 inference.csv 中 input_path 条目的图像根目录;必须位于 $WORKSPACE 内
MIN_RECALL=1.0                       # 默认值为零遗漏;若 KPI 要求放宽,则可调低
TOP_K=50                             # 每个标签的增生预算
OUT="$EXP_DIR/rca_results/$(date +%Y-%m-%d_%H%M%S)"
SPEC="$OUT/vcn_aoi_spec.yaml"
IMG=$(python3 -c "import yaml,os; print(yaml.safe_load(open(os.environ['TAO_SKILL_BANK_PATH']+'/versions.yaml'))['images']['tao_toolkit']['data_services'])")

mkdir -p "$OUT"

# 写入本次运行的差距分析规范
cat > "$SPEC" <<EOF
min_recall: $MIN_RECALL
top_k_per_label: $TOP_K
EOF

docker run --gpus all --rm --ipc=host \
    -v "$WORKSPACE:$WORKSPACE" -w "$WORKSPACE" \
    "$IMG" gap_analysis vcn_aoi \
    -e "$SPEC" \
    inference_results_dir="$EXP_DIR/inference/latest/" \
    train_config="$EXP_DIR/train.yaml" \
    kpi_media_path="$DATASET_ROOT" \
    results_dir="$OUT"

# 容器以 root 身份写入,且 --user 选项已被移除;如有需要,请将所有权变回主机 UID。
docker run --rm -v "$WORKSPACE:/w" alpine chown -R "$(id -u):$(id -g)" "/w/$(realpath --relative-to="$WORKSPACE" "$OUT")"

# 输出验证信息,以便脚本检查钩子能看到真实数值
python3 - "$OUT" << 'PYEOF'
import json, os, sys
out = sys.argv[1]
unreachable = os.path.join(out, "unreachable_kpi.txt")
if os.path.isfile(unreachable):
    print("KPI 无法访问 — 详见", unreachable)
    sys.exit(0)
with open(os.path.join(out, "threshold.txt")) as f:
    print("阈值:", f.read().strip())
with open(os.path.join(out, "metrics.json")) as f:
    m = json.load(f)
print(f"precision={m['precision']:.4f} recall={m['recall']:.4f} f1={m['f1']:.4f}")
import pandas as pd
df = pd.read_parquet(os.path.join(out, "kpi_gaps.parquet"))
print(f"kpi_gaps.parquet: rows={len(df)}, cols={list(df.columns)}")
print(df['label'].value_counts())
PYEOF

输出

将所有内容写入实验结果目录下的带时间戳的文件夹中。容器的输出直接写入该文件夹;可视化抽查结果写入rca_images/;任何运行时打包钩子都可能在写入RCA_Report.md之后,将会话/配置捕获的生成文件添加到该文件夹中。

/rca_results/YYYY-MM-DD_HHMMSS/
├── RCA_Report.md              # 完整的差距分析报告(由您编写)
├── kpi_gaps.parquet           # 容器:每个标签的前 K 个最弱样本,按照明条件展开
├── threshold.txt              # 容器:选定的决策阈值(单个浮点数)
├── metrics.json               # 容器:混淆矩阵 + 按标签划分的分布统计数据
├── weak_samples_breakdown.txt # 容器:按标签划分的样本数/误分类数/边际计数
├── unreachable_kpi.txt        # 容器:仅当无阈值满足 min_recall 时生成
├── rca_images/                # 由您提供:10 个已查看的弱样本的缩略图
├── rca_config/                # 由钩子自动复制
└── session log/artifacts      # 可选,取决于运行时环境的打包捕获

运行开始时,请在 Bash 中执行`date +%Y-%m-%d_%H%M%S` 命令获取真实时间戳。切勿硬编码或凭空猜测。如果用户指定了自定义输出路径,请使用该路径,但需保持相同的内部结构。

常见陷阱

最严重的故障模式莫过于在 min_recall=1.0时忘记设置top_k_per_label: 在该召回率下,所选阈值位于每个 NO_PASS 分数的水平或以下,因此如果没有top_k_per_label,容器将回退到“低于阈值的样本”过滤器,该过滤器仅返回误分类的 PASS 行,而 NO_PASS 行为零,从而破坏了增强队列。 请务必在规范中或通过 Hydra 覆盖显式指定正样本的top_k_per_label(默认值为 50)。

请参阅references/pitfalls.md获取完整的检查清单,涵盖以下内容:遗漏top_k_per_label;传递--user 参数;仅使用 Hydra 覆盖项调用(未使用-e );spec 文件位于$WORKSPACE 之外;spec 文件中存在未解析的???哨兵;未拉取镜像/标签错误; 路径挂载不匹配;写入了unreachable_kpi.txt;inference.csv缺少必需的列;训练 YAML文件缺少dataset.classify.input_map或image_ext;kpi_media_path与input_path前缀不匹配;以及从容器内部未检测到 GPU。

报告结构

编写RCA_Report.md文件,将其作为一份简洁(1000–1800 词)的计算差距分析报告——报告的深度应源于准确的数值和清晰的行动清单,而非叙述性内容。 完整的报告模板(包含 7 个部分:结论、阈值选择、弱点分布、前 K 个最弱样本、视觉抽查、按标签细分、建议措施——附混淆矩阵和表格布局)位于references/output-template.md 中。 当存在unreachable_kpi.txt 文件时,请用一个简短章节替换第 3–6 节,该章节应引用该文件的内容,并将第 7 节简化为一条建议:重新训练或重新标注。

执行顺序

  1. 从versions.yaml(images.tao_toolkit.data_services)中解析DS_IMAGE,然后运行docker info、nvidia-smi 以及docker image inspect "$DS_IMAGE"(若不存在则拉取),执行一次以确认环境。若任何步骤失败,则显示明确信息并终止。
  2. 运行 `date +%Y-%m-%d_%H%M%S` 获取时间戳;创建目录 `/rca_results/` 及 `/`。
  3. 将vcn_aoi_spec.yaml写入带时间戳的目录中,并填写min_recall和top_k_per_label参数。将其保存在$WORKSPACE目录下,以便-e路径能在容器内部解析。
  4. 运行docker run … "$DS_IMAGE" gap_analysis vcn_aoi -e vcn_aoi_spec.yaml inference_results_dir=… train_config=… kpi_media_path=… output_dir=…. 容器会将kpi_gaps.parquet、threshold.txt、metrics.json 和weak_samples_breakdown.txt写入results_dir。将选定的阈值和保留行数打印到标准输出(stdout),以便脚本检查钩子能够验证运行是否产生了预期输出。
  5. 如果存在unreachable_kpi.txt 文件,则跳过步骤 6 并生成简要报告。否则继续执行。
  6. 从kpi_gaps.parquet 中挑选 10 个弱样本(5 个最弱的“PASS”样本 + 5 个最弱的“NO_PASS”样本),通过 Read 查看每张测试图像并进行分类,然后将每张图像复制到rca_images/ 目录下。
  7. 最后生成RCA_Report.md—— 生成该文件将触发打包钩子,该钩子会一并复制会话日志和技能配置。
在 GitHub 上查看
---
name: tao-analyze-gaps-visual-changenet
description: Identifies the weakest samples per ground-truth label in NVIDIA TAO VCN Classify experiments by running a Docker container that performs threshold sweep, weakness scoring, and per-lighting expansion, then surfaces top-K weak samples for downstream augmentation or relabeling.
license: Apache-2.0
---

# TAO VCN Classify Gap Analysis Skill

You are an analyst for NVIDIA TAO VCN Classify (Visual Component Net) inference results. Your job is to identify the **weakest samples per ground-truth label** by measuring signed distance from the decision threshold *in the wrong direction*, then surface them for downstream augmentation or relabeling.

This skill is intentionally lightweight. VCN's classify head is a single-score binary boundary (PASS vs NO_PASS by `siamese_score`), so the analysis is computational, not investigative. The whole computation lives behind one direct `docker run` invocation against the `tao_toolkit.data_services` image declared in `versions.yaml` (resolved at runtime — see Setup). The container's entrypoint takes `<category> <action> [hydra overrides...]`; we pass `gap_analysis vcn_aoi key=value …`. Each override is a bare Hydra `key=value` that selectively overrides the script's `GapAnalysisConfig` schema (defaults are baked into the container; introspect with `docker run ... gap_analysis vcn_aoi --cfg=job`). (There is no `dataset` keyword inside the container — that's the TAO launcher's pillar prefix and is dropped here.) You do **not** need delegated analysis, multi-phase image audits, or component-type clustering — VCN does not expose those dimensions. View only a small set of representative weak samples to qualify the gaps after the container returns.

CLI surface can shift between data-services container builds. If a `gap_analysis vcn_aoi` invocation fails on argument parsing, introspect the actual schema once per image with `docker run --rm "$DS_IMAGE" gap_analysis vcn_aoi --cfg=job` and reconcile any renamed keys (e.g. `inference_csv` vs `inference_results_dir`, `output_dir` vs `results_dir`) before retrying. Output parquet name is `kpi_gaps.parquet`.

---

## Inputs

1. **Experiment result directory** — contains `inference/inference.csv` from TAO VCN Classify inference. Required columns: `input_path`, `object_name`, `label`, `siamese_score`. Pass the **directory** (e.g. `inference/latest/`), not the CSV file — the container reads `inference_results_dir/inference.csv`.
2. **Training code/config directory** — contains the VCN train YAML. The container reads `dataset.classify.input_map` (lighting condition list) and `dataset.classify.image_ext` from it to expand each weak sample into one row per lighting.
3. **Dataset directory** — image root prepended to the relative `input_path` from each row (`kpi_media_path`).
4. **Schema overrides** — `min_recall`, `top_k_per_label`, and optionally a hard-pinned `threshold` are passed as Hydra overrides (defaults: `min_recall=1.0`, `top_k_per_label=50`, `threshold=-1.0` meaning sweep). **`top_k_per_label` must be a positive integer** — omitting it flips the container into "below-threshold filter" mode, which at `min_recall=1.0` returns only PASS misclassifications and zero NO_PASS rows. See Common pitfalls.

---

## Setup

The threshold sweep, weakness ranking, and per-lighting expansion all run inside the `tao_toolkit.data_services` image declared in `versions.yaml`. Resolve the concrete URI once at the top of the run, then confirm Docker, the NVIDIA container toolkit, and a GPU are present and ensure the image is cached:

```bash
# Resolve tao_toolkit.data_services → concrete nvcr.io/... URI from versions.yaml
DS_IMAGE=$(python3 -c "import yaml,os; print(yaml.safe_load(open(os.environ['TAO_SKILL_BANK_PATH']+'/versions.yaml'))['images']['tao_toolkit']['data_services'])")
echo "DS_IMAGE=$DS_IMAGE"

docker info > /dev/null && echo "OK: docker"
nvidia-smi > /dev/null && echo "OK: GPU"
docker image inspect "$DS_IMAGE" > /dev/null \
  || docker pull "$DS_IMAGE"
```

`TAO_SKILL_BANK_PATH` is usually exported by the installed skill bank. If it is unset, point it at the skill-bank repo root before resolving. A GPU is required; aborting early on a GPU-less host saves a confusing late error.

Three setup rules are load-bearing and easy to get wrong:

- **Path mounting** — every host path the container reads or writes (`inference.csv`, train YAML, dataset image root, output dir) must be bind-mounted, simplest with `-v $WORKSPACE:$WORKSPACE -w $WORKSPACE` so absolute paths resolve identically on both sides.
- **Do not pass `--user $(id -u):$(id -g)`** — it triggers `KeyError: 'getpwuid(): uid not found: <uid>'` during the container's `transformers` import; `chown` outputs back to the host UID afterwards instead.
- **`-e <spec>` is required, not optional** — current images hard-require it and exit with `ValueError: The subtask vcn_aoi requires the following argument: -e/--experiment_spec_file` before parsing CLI overrides.

See `references/container-setup.md` for the full path-mounting pattern, the `--user`/`chown` rationale and `alpine` chown command, multi-`-v` guidance, and the `-e <spec>` requirement detail.

---

## Method

The whole skill is a single `docker run` invocation followed by a small visual spot-check. The container does Steps 1–4 internally (threshold sweep, weakness scoring, top-K selection, per-lighting expansion). You handle Step 5 (visual spot-check) directly with the Read tool.

### Step 1–4 — Run the container

```bash
$DOCKER gap_analysis vcn_aoi \
    inference_results_dir=<exp_dir>/inference/<label>/ \
    train_config=<exp_dir>/train.yaml \
    kpi_media_path=<dataset_root> \
    results_dir=<rca_results_dir> \
    top_k_per_label=50
```

> **Always pass `top_k_per_label`.** This is the argument that switches the container
> from the default "samples below threshold" filter into proper top-K-per-label
> ranking. At `min_recall=1.0` the threshold is by construction at-or-below every
> NO_PASS score, so the below-threshold filter returns ONLY misclassified PASS rows
> and zero NO_PASS rows — useless as an augmentation queue. With `top_k_per_label`
> set to a positive integer (either in the spec or as a Hydra override), the
> container computes signed weakness against the threshold for every row and
> surfaces the K weakest **per ground-truth label**, which is the per-label ranked
> output downstream steps consume.

Reads `inference.csv`, sweeps every unique `siamese_score` plus one value just below the minimum, keeps the candidates with NO_PASS-class recall ≥ `min_recall` (with `1e-12` tolerance), then picks the threshold with the best F1 (tie-break: precision, then threshold value). For every row, computes signed weakness from that threshold (positive = misclassified, negative = correct, magnitude = margin). Sorts by weakness descending and takes the top `top_k_per_label` per ground-truth label, then expands each weak row into one row per lighting condition using `dataset.classify.input_map` and `dataset.classify.image_ext` from the train YAML.

If **no** candidate threshold meets the recall target, the container exits non-zero and writes `unreachable_kpi.txt` into `results_dir` explaining which recall the model can actually achieve. In that case, stop the analysis after the docker call, write a one-section report explaining the model fundamentally cannot reach the KPI at any operating point, and recommend retraining or relabeling — skip the visual spot-check.

**Container writes into `results_dir`:**

| Artifact | Contents |
|----------|----------|
| `kpi_gaps.parquet` | Top-K weakest per label, expanded per lighting. Columns: `filepath`, `label`, `siamese_score`, `weakness`. |
| `threshold.txt` | Chosen decision threshold (single float, plain text). |
| `metrics.json` | At the chosen threshold: `precision`, `recall`, `f1`, confusion matrix `{tp, fp, tn, fn}`, plus per-label `{total, mean_weakness, median_weakness, max_weakness, n_misclassified}`. |
| `weak_samples_breakdown.txt` | Per-label kept-row breakdown: `<count>` total, `<%>` of all kept rows, `N` misclassified (weakness > 0), `N` marginal (weakness ≤ 0). |
| `unreachable_kpi.txt` | Only written when the recall target is unreachable. Presence of this file means: skip Step 5, write the abridged report, recommend retrain. |

Print the container's stdout summary (chosen threshold, kept-row counts, per-label breakdown) to your own stdout so the script-check hook can verify the run produced output.

### Step 5 — Visual spot check (small, fixed)

Skip this step if `unreachable_kpi.txt` exists. Otherwise use the Read tool to **view** the 5 weakest PASS samples and the 5 weakest NO_PASS samples from `kpi_gaps.parquet` (deduplicated to one row per sample, using the FIRST-lighting `filepath`), classify each as exactly one of **mislabeled** / **edge case** / **data quality** / **systematic**, and copy each viewed image (resized to 128×128 if PIL is available, otherwise just copy) into `<results_dir>/rca_images/`. This is the only image inspection required — do not view dozens of images, run failure mode clustering, or audit goldens (VCN has no golden images).

See `references/visual-spot-check.md` for the exact sample-selection sort, the per-lighting deduplication rule, the full definition of each verdict category, and the image-copy detail.

---

## Reference invocation

Paste-and-edit the workspace, the four paths, and the two numeric knobs; this runs end-to-end. Capture stdout so the script-check hook sees row counts.

```bash
WORKSPACE=<absolute path>            # mounted identically inside the container
EXP_DIR=<experiment_result_dir>      # contains inference/inference.csv and train.yaml; must be inside $WORKSPACE
DATASET_ROOT=<dataset_root>          # image root for inference.csv input_path entries; must be inside $WORKSPACE
MIN_RECALL=1.0                       # zero-miss default; lower if KPI relaxes
TOP_K=50                             # per-label augmentation budget
OUT="$EXP_DIR/rca_results/$(date +%Y-%m-%d_%H%M%S)"
SPEC="$OUT/vcn_aoi_spec.yaml"
IMG=$(python3 -c "import yaml,os; print(yaml.safe_load(open(os.environ['TAO_SKILL_BANK_PATH']+'/versions.yaml'))['images']['tao_toolkit']['data_services'])")

mkdir -p "$OUT"

# Write the gap-analysis spec for this run
cat > "$SPEC" <<EOF
min_recall: $MIN_RECALL
top_k_per_label: $TOP_K
EOF

docker run --gpus all --rm --ipc=host \
    -v "$WORKSPACE:$WORKSPACE" -w "$WORKSPACE" \
    "$IMG" gap_analysis vcn_aoi \
    -e "$SPEC" \
    inference_results_dir="$EXP_DIR/inference/latest/" \
    train_config="$EXP_DIR/train.yaml" \
    kpi_media_path="$DATASET_ROOT" \
    results_dir="$OUT"

# Container writes as root with --user dropped; chown back to host UID if needed.
docker run --rm -v "$WORKSPACE:/w" alpine chown -R "$(id -u):$(id -g)" "/w/$(realpath --relative-to="$WORKSPACE" "$OUT")"

# Sanity print so the script-check hook sees real numbers
python3 - "$OUT" << 'PYEOF'
import json, os, sys
out = sys.argv[1]
unreachable = os.path.join(out, "unreachable_kpi.txt")
if os.path.isfile(unreachable):
    print("KPI UNREACHABLE — see", unreachable)
    sys.exit(0)
with open(os.path.join(out, "threshold.txt")) as f:
    print("threshold:", f.read().strip())
with open(os.path.join(out, "metrics.json")) as f:
    m = json.load(f)
print(f"precision={m['precision']:.4f} recall={m['recall']:.4f} f1={m['f1']:.4f}")
import pandas as pd
df = pd.read_parquet(os.path.join(out, "kpi_gaps.parquet"))
print(f"kpi_gaps.parquet: rows={len(df)}, cols={list(df.columns)}")
print(df['label'].value_counts())
PYEOF
```

---

## Outputs

Write everything into a timestamped folder under the experiment result directory. The container's outputs go straight there; the visual spot-check writes `rca_images/`; any runtime packaging hook may add session/config capture artifacts after `RCA_Report.md` is written.

```
<experiment_result_dir>/rca_results/YYYY-MM-DD_HHMMSS/
├── RCA_Report.md              # Full gap analysis report (you write this)
├── kpi_gaps.parquet           # Container: top-K weakest per label, expanded per lighting
├── threshold.txt              # Container: chosen decision threshold (single float)
├── metrics.json               # Container: confusion matrix + per-label distribution stats
├── weak_samples_breakdown.txt # Container: per-label count/misclassified/marginal counts
├── unreachable_kpi.txt        # Container: ONLY when no threshold meets min_recall
├── rca_images/                # You: thumbnails of the 10 viewed weak samples
├── rca_config/                # Auto-copied by hook
└── session log/artifacts      # Optional, runtime-dependent packaging capture
```

At the start of the run, get the real timestamp by running `date +%Y-%m-%d_%H%M%S` in Bash. Do NOT hardcode or guess. If the user specifies a custom output path, use that instead but maintain the same internal structure.

---

## Common pitfalls

The single most consequential failure mode is **forgetting `top_k_per_label` when `min_recall=1.0`**: at that recall the chosen threshold sits at or below every NO_PASS score, so without `top_k_per_label` the container falls back to a "samples below threshold" filter that returns ONLY misclassified PASS rows and zero NO_PASS rows, breaking the augmentation queue. Always include an explicit positive `top_k_per_label` (default 50) in the spec or as a Hydra override.

See `references/pitfalls.md` for the complete checklist, covering: forgetting `top_k_per_label`; passing `--user`; calling with only Hydra overrides (no `-e <spec>`); spec file outside `$WORKSPACE`; spec file with unresolved `???` sentinels; image not pulled / wrong tag; path-mount mismatch; `unreachable_kpi.txt` written; `inference.csv` missing required columns; train YAML missing `dataset.classify.input_map` or `image_ext`; `kpi_media_path` not matching `input_path` prefixes; and no GPU detected from inside the container.

---

## Report Structure

Write `RCA_Report.md` as a tight (1000–1800 word) computational gap analysis — depth comes from accurate numbers and a clear action list, not narrative. The full report template (7 sections: Verdict, Threshold Selection, Weakness Distribution, Top-K Weakest Samples, Visual Spot Check, Per-Label Breakdown, Recommended Actions — with the confusion-matrix and table layouts) is in `references/output-template.md`. When `unreachable_kpi.txt` exists, replace sections 3–6 with a single short section quoting that file's contents and collapse section 7 to one recommendation: retrain or relabel.

---

## Execution Order

1. Resolve `DS_IMAGE` from `versions.yaml` (`images.tao_toolkit.data_services`), then run `docker info`, `nvidia-smi`, and `docker image inspect "$DS_IMAGE"` (pulling if missing) once to confirm the environment. Abort with a clear message if any fail.
2. Run `date +%Y-%m-%d_%H%M%S` to get the timestamp; create `<experiment_result_dir>/rca_results/<timestamp>/`.
3. Write `vcn_aoi_spec.yaml` into the timestamped dir with `min_recall` and `top_k_per_label` filled in. Keep it under `$WORKSPACE` so the `-e` path resolves inside the container.
4. Run `docker run … "$DS_IMAGE" gap_analysis vcn_aoi -e vcn_aoi_spec.yaml inference_results_dir=… train_config=… kpi_media_path=… output_dir=…`. The container writes `kpi_gaps.parquet`, `threshold.txt`, `metrics.json`, `weak_samples_breakdown.txt` into `results_dir`. Print the chosen threshold and kept-row counts to stdout so the script-check hook can verify the run produced output.
5. If `unreachable_kpi.txt` exists, skip Step 6 and write the abridged report. Otherwise continue.
6. Pick 10 weak samples (5 weakest PASS + 5 weakest NO_PASS) from `kpi_gaps.parquet`, view each test image with Read, classify, and copy each into `rca_images/`.
7. Write `RCA_Report.md` last — writing it triggers the packaging hook, which copies session logs and skill config alongside.

所有文件

1 个文件

安装 tao-analyze-gaps-visual-changenet

下载技能文件并将其解压到 .claude/skills/ 目录下。

下载ZIP

克隆仓库并复制技能文件到您的项目中。

git clone https://github.com/NVIDIA/skills/tree/main/skills/tao-analyze-gaps-visual-changenet # Copy SKILL.md to your .claude/skills/ directory

复制 复制
快速设置: 将技能文件夹复制到 .claude/skills/ Claude 将自动检测并使用该技能
仓库 NVIDIA/skills

相关技能

web-search
更新时间 2026-06-29
webapp-testing
更新时间 2026-06-29
lark-base
更新时间 2026-07-05
agentmail
更新时间 2026-06-29
OR