选项
首页首页 Skill 数据库管理 tao-mine-aoi-images

tao-mine-aoi-images

NVIDIA/skills NVIDIA/skills

在VCN AOI工作流中,该方法将目标图像和源图像的Parquet文件进行嵌入,随后从最近邻源图像中提取数据以进行增强。

...展开全部
26
更新时间 2026-09-24

DEFT 挖掘与嵌入技能

您是 VCN AOI 领域中 DEFT“先嵌入后挖掘”工作流的操作员。 您的任务是:接收一组弱目标图像(间隙分析或布线输出)的 Parquet 文件以及一个源图像池,然后生成一组经过去重处理的、与目标图像相似的源图像 Parquet 文件——这些文件已准备好用于下一轮训练。

该工作流是固定且确定性的:先对目标图像进行嵌入,再对源图像池进行嵌入,最后挖掘最近邻。每个步骤输出的 Parquet 文件即为下一步的输入。 这里没有迭代搜索、没有聚类处理、也没有人工干预的筛选——深度来源于选择正确的编码器和正确的TopN,而非多阶段的调查。

整个技能只是对三个直接调用`docker run`的薄包装,这些调用针对在 `versions.yaml` 中声明的 `tao_toolkit.data_services` 镜像(在运行时解析——参见“设置”部分)。 容器的入口点接受 -e [hydra 覆盖项...]——传递embedding 参数 image_embeddings -e …用于嵌入,以及tmm 最近邻 -e …用于挖掘。-e标志指向一个 YAML 文件,该文件为子任务的架构提供默认值;其后的内容均为纯粹的 Hydra 覆盖项(key=value),用于根据每次运行情况有选择地覆盖 spec 字段。 (容器内部没有dataset关键字——这是 TAO 启动器的 pillar 前缀,此处已被省略。)如果镜像未被缓存,请拉取一次:docker pull "$DS_IMAGE"(根据“设置”部分解析$DS_IMAGE后)。

数据服务发布之间,模式键名可能会更改(RCA 技能中,inference_csv变为inference_results_dir,output_dir变为results_dir)。 如有疑问,请针对每个镜像检查一次实际模式:docker run --rm "$DS_IMAGE" embedding image_embeddings --cfg=job以及... tmm nearest_neighbors --cfg=job。

输入

  1. 目标 Parquet文件——差距分析的输出结果,通常是来自tao-route-visual-changenet-samples 的 mining_gaps.parquet(或者如果跳过了路由步骤,则来自tao-analyze-gaps-visual-changenet的gaps.parquet)。 必填字段:filepath。如果同时存在label 字段,则在挖掘过程中可进行基于标签的过滤;否则,挖掘任务将默认跳过该过滤操作。
  2. 源图像池——用于挖掘的候选图像 Parquet 文件,其中包含filepath列。如果用户只有 CSV 文件,请在第 2 步之前将其转换为具有相同列的Parquet 文件。若要进行基于标签的过滤,图像池还必须包含label列。
  3. 嵌入式规范文件— 包含model、model_path、batch_size 以及(仅当model_path为 TAO.pth/.ckpt 时)model_config_path 的 YAML 文件。 该配置文件在步骤 1 和步骤 2 中复用;input_parquet/output_parquet由 Hydra 在每次运行时作为覆盖参数提供。两个嵌入步骤必须由同一份配置文件驱动——来自不同编码器的嵌入向量无法进行比较,编码器不匹配是导致“挖掘出的图像看似无关”报告的最常见原因。
  4. 挖掘规范文件——一个包含topn、knn_metric、filter_by_label 以及(极少更改的)source_embed_column_name/target_embed_column_name 的 YAML文件。source_parquet/target_parquet/output_parquet是运行时的 Hydra 覆盖值。 SigLIP 和 CLIP 嵌入向量应使用knn_metric: cosine。当filter_by_label: true时,若任一嵌入向量的 Parquet 文件缺少标签列,容器会记录一条警告并跳过过滤继续执行。

环境配置

在运行开始时,从versions.yaml中解析一次具体的tao_toolkit.data_servicesURI,然后在执行其他操作之前,确认已安装 Docker、NVIDIA 容器工具包以及 GPU。 编码器前向传播和 cuML/cuDF k-NN 搜索均需要 GPU;若无 CUDA,这两个步骤都会失败。

# 从 versions.yaml 中解析 tao_toolkit.data_services → 具体的 nvcr.io/... URI
DS_IMAGE=$(python3 -c "import yaml,os; print(yaml.safe_load(open(os.environ['TAO_SKILL_BANK_PATH']+'/versions.yaml'))['images']['tao_toolkit']['data_services'])")
echo "DS_IMAGE=$DS_IMAGE"

docker info > /dev/null && echo "OK: docker"
nvidia-smi > /dev/null && echo "OK: GPU"
docker image inspect "$DS_IMAGE" > /dev/null \
  || docker pull "$DS_IMAGE"

容器读取或写入的每个主机路径都必须通过绑定挂载。最可靠的方法是将工作区根目录挂载为容器内外路径完全一致,然后为这三个调用复用一个$DOCKER别名:

WORKSPACE=
DOCKER="docker run --gpus all --rm --ipc=host -v $WORKSPACE:$WORKSPACE -w $WORKSPACE $DS_IMAGE"

请勿传递--user $(id -u):$(id -g)—— 这会在任何工作开始前,于Transformers导入过程中触发getpwuid() KeyError错误。容器以 root 身份运行;随后 chown 命令会将所有权归还给主机的 UID。

每次迭代只需编写两次 spec 文件,并将它们放置在$WORKSPACE下,以便-e参数在挂载的两端都能解析;每次运行的值不写入 spec 文件,而是作为 Hydra 覆盖参数传递。 如果源数据池是 CSV 文件,请预先将其转换为 Parquet 格式(保留文件路径,如有标签则一并保留)。默认的embedding_spec.yaml使用model: SigLIP,model_path: google/siglip-base-patch16-224,batch_size: 64; 默认的mining_spec.yaml使用topn: 5,knn_metric: cosine,filter_by_label: "false"(加引号——模式将其视为字符串)。

请参阅references/setup.md以获取完整的环境说明、TAO_SKILL_BANK_PATH的处理方法、路径挂载的原理、getpwuid 的chown 变通方案、CSV 转 Parquet 的代码片段,以及规范文件的原文编写块。

方法

按顺序执行三个命令。每个命令输出的 Parquet 文件即为下一个命令的输入。请以纯 Bash 方式运行这些命令;Setup 中定义的$DOCKER别名会自动处理容器、GPU 及挂载。 每次调用都遵循相同的格式:-e用于应用内置默认值,随后通过少量 Hydra 覆盖设置来指定运行时所需的特定路径。

步骤 1 — 嵌入目标镜像

$DOCKER embedding image_embeddings \
    -e \
    input_parquet= \
    output_parquet=

读取差距分析/路由输出,并生成一个 Parquet 文件,其中包含文件路径、嵌入信息以及任何从输入原样保留的额外元数据列(例如label、siamese_score、weakness)。 将输出模式(pd.read_parquet(...).columns)打印到标准输出,以便脚本检查钩子能够确认嵌入向量列是否存在。

若需在不修改规范文件的情况下为单次运行覆盖model/model_path/batch_size,请将其作为 Hydra 覆盖项追加(例如model_path=...)。

步骤 2 — 嵌入源数据池

$DOCKER embedding image_embeddings \
    -e \
    input_parquet= \
    output_parquet=

命令格式与步骤 1 相同,但应用于源数据池。请使用与步骤 1完全相同的 embedding_spec.yaml文件,且在此处不要对model/model_path/batch_size进行不同设置——若两个步骤中的编码器配置不一致,将产生无法比较的嵌入向量。

步骤 3 — 挖掘最近邻

$DOCKER tmm nearest_neighbors \
    -e \
    source_parquet= \
    target_parquet= \
    output_parquet=

针对每个目标嵌入,根据选定的度量标准查找前 n 个最接近的源嵌入,在各目标之间去除重复项,并生成一个单列(文件路径)Parquet 文件,其中包含经过挖掘的唯一源路径。 该容器还会在输出 Parquet 文件旁生成一个mining_summary.txt文件,其中包含:查询计数、邻居计数、已移除的重复项,以及(当启用标签过滤时)保留与丢弃的配对计数。 在扫描时,可通过内联 Hydra 覆盖调整topn、knn_metric 或filter_by_label(例如topn=10)——无需重写规范。

当filter_by_label=true时,若某个嵌入式 Parquet 文件缺少标签列,容器会记录一条警告并跳过过滤直接继续。如果挖掘出的输出比预期大或包含跨标签对,请先检查 Docker 日志中是否有该警告,再判断任务是否执行正确。

请参阅references/reference-invocation.md,其中提供了最简化的“复制粘贴并编辑”端到端操作指南(解析$DS_IMAGE、编写两个规范、执行全部三个步骤、更改输出文件所有权并打印行数),可作为单个流式 Bash 代码块运行。

输出与报告

将所有内容写入实验 / 迭代目录下的带时间戳的文件夹中。通过在 Bash 中运行date +%Y-%m-%d_%H%M%S获取真实时间戳——切勿硬编码或随意猜测。如果用户指定了自定义输出路径,请直接使用该路径,但需保持相同的内部布局。 打包钩子会在写入Mining_Report.md时自动添加mining_config/和claude_session.jsonl。

挖掘生成的 Parquet 文件是下游训练所使用的成果。两个嵌入式 Parquet 文件虽为中间产物,但值得保留——它们可在针对同一源数据池的多次挖掘运行中重复使用,并且当报告显示“看似无关”时,这是进行编码器级调试的唯一参考依据。

请参阅references/outputs-and-reporting.md以了解完整的输出目录结构以及原封不动的Mining_Report.md模板(结论、输入、编码器一致性、挖掘运行、按标签细分、输出合理性、建议措施;字数控制在 600–1200 词之间)。

常见陷阱

最常见的失败原因是两个嵌入步骤中的编码器不匹配——这是导致垃圾挖掘输出最主要的原因;两个步骤必须使用相同的embedding_spec.yaml 文件。 其他常见陷阱包括:传递--user 参数(导致getpwuid KeyError)、跳过嵌入步骤、因缺少标签列而导致filter_by_label=true 被静默忽略、spec 文件位于$WORKSPACE 之外、未解析的???哨兵未解析、TAO检查点缺少model_config_path、直接导入CSV源数据池、主机/容器路径不匹配、无GPU、镜像标签未拉取或使用 :latest,以及topn × N_targets ≫ 源数据大小(属预期情况——请报告实际挖掘到的计数)。

请参阅references/troubleshooting.md以获取包含确切错误、原因及解决方案的完整陷阱列表。

执行顺序

  1. 从versions.yaml(images.tao_toolkit.data_services)中解析DS_IMAGE,然后运行docker info、nvidia-smi 以及docker image inspect "$DS_IMAGE"(若缺失则拉取)一次以确认环境。若任何步骤失败,则显示明确信息并中止。
  2. 运行 `date +%Y-%m-%d_%H%M%S` 获取时间戳;创建目录 `/mining_results/` 及文件 `/`。
  3. 将embedding_spec.yaml和mining_spec.yaml写入带时间戳的目录中,并填写编码器选择和数据挖掘参数。将这些文件保存在$WORKSPACE目录下,以便-e路径能在容器内部解析。
  4. 如果源数据池是 CSV 文件,请先将其转换为 Parquet 格式(保留文件路径和标签)。
  5. 通过 `docker run … embedding image_embeddings -e embedding_spec.yaml input_parquet=… output_parquet=…` 执行步骤 1(嵌入目标数据)。将生成的 Parquet 文件的行数和列数打印到标准输出。
  6. 使用与步骤 1完全相同的 embedding_spec.yaml运行步骤 2(嵌入源数据集)。将输出 Parquet 的行数和列数打印到标准输出。
  7. 通过 `docker run … tmm nearest_neighbors -e mining_spec.yaml source_parquet=… target_parquet=… output_parquet=…` 运行步骤 3(挖掘最近邻)。确认`mining_summary.txt`已写入`mined.parquet` 文件旁。
  8. 若目标嵌入 Parquet 文件和挖掘输出文件均包含标签,则通过文件路径将两者进行关联,从而计算各标签的分布情况(第 5 节)。
  9. 最后写入Mining_Report.md— 写入该文件会触发打包钩子,该钩子会同时复制会话日志和技能配置。
在 GitHub 上查看
---
name: tao-mine-aoi-images
description: Embeds target and source image parquets, then mines nearest-neighbour source images for augmentation in VCN AOI workflows.
license: Apache-2.0
---

# DEFT Mining and Embedding Skill

You are the operator of the DEFT embed-then-mine workflow for VCN AOI. Your job is to take a parquet of weak target images (the gap-analysis or routing output) and a source pool, then produce a deduplicated parquet of mined source images that look similar to the targets — ready to feed into the next training round.

The workflow is fixed and deterministic: **embed the targets, embed the source pool, then mine nearest neighbours.** Each step's output parquet is the next step's input. There is no iterative search, no clustering pass, no human-in-the-loop selection — depth comes from picking the right encoder and the right `topn`, not from a multi-phase investigation.

The whole skill is a thin wrapper around three direct `docker run` invocations against the `tao_toolkit.data_services` image declared in `versions.yaml` (resolved at runtime — see Setup). The container's entrypoint takes `<category> <action> -e <spec.yaml> [hydra overrides...]` — pass `embedding image_embeddings -e <embedding_spec.yaml> …` for embedding and `tmm nearest_neighbors -e <mining_spec.yaml> …` for mining. The `-e` flag points at a YAML that supplies default values for the subtask's schema; anything afterward is a bare Hydra override (`key=value`) that selectively overrides spec fields per run. (There is no `dataset` keyword inside the container — that's the TAO launcher's pillar prefix and is dropped here.) Pull the image once if it isn't cached: `docker pull "$DS_IMAGE"` (after resolving `$DS_IMAGE` per Setup).

Schema keys can rename between data-services releases (the RCA skill saw `inference_csv` → `inference_results_dir`, `output_dir` → `results_dir`). When in doubt, introspect the actual schema once per image: `docker run --rm "$DS_IMAGE" embedding image_embeddings --cfg=job` and `... tmm nearest_neighbors --cfg=job`.

---

## Inputs

1. **Target parquet** — the gap-analysis output, typically `mining_gaps.parquet` from `tao-route-visual-changenet-samples` (or `gaps.parquet` from `tao-analyze-gaps-visual-changenet` if routing was skipped). Required column: `filepath`. If `label` is also present, label-aware filtering during mining is available; otherwise the mining task silently no-ops the filter.
2. **Source pool** — a parquet of candidate images to mine against, with a `filepath` column. If the user only has a CSV, convert it to a parquet **with the same columns** before Step 2. For label-aware filtering, the pool must also carry a `label` column.
3. **Embedding spec file** — a YAML containing `model`, `model_path`, `batch_size`, and (only when `model_path` is a TAO `.pth`/`.ckpt`) `model_config_path`. Reused across Steps 1 and 2; `input_parquet`/`output_parquet` are supplied per run as Hydra overrides. The **same** spec MUST drive both embedding steps — embeddings from different encoders are not comparable, and mismatched encoders are the most common cause of "the mined images look unrelated" reports.
4. **Mining spec file** — a YAML containing `topn`, `knn_metric`, `filter_by_label`, and (rarely changed) `source_embed_column_name`/`target_embed_column_name`. `source_parquet`/`target_parquet`/`output_parquet` are Hydra overrides at run time. SigLIP and CLIP embeddings should use `knn_metric: cosine`. When `filter_by_label: true` but either embedding parquet lacks a `label` column, the container logs a warning and proceeds **without** filtering.

---

## Setup

Resolve the concrete `tao_toolkit.data_services` URI from `versions.yaml` once at the top of the run, then confirm Docker, the NVIDIA container toolkit, and a GPU are present before doing anything else. A GPU is required for both the encoder forward pass and the cuML/cuDF k-NN search; both steps fail without CUDA.

```bash
# Resolve tao_toolkit.data_services → concrete nvcr.io/... URI from versions.yaml
DS_IMAGE=$(python3 -c "import yaml,os; print(yaml.safe_load(open(os.environ['TAO_SKILL_BANK_PATH']+'/versions.yaml'))['images']['tao_toolkit']['data_services'])")
echo "DS_IMAGE=$DS_IMAGE"

docker info > /dev/null && echo "OK: docker"
nvidia-smi > /dev/null && echo "OK: GPU"
docker image inspect "$DS_IMAGE" > /dev/null \
  || docker pull "$DS_IMAGE"
```

Every host path the container reads or writes must be bind-mounted. The most predictable approach mounts the workspace root with **identical paths** inside and outside the container, then reuses one `$DOCKER` alias for the three invocations:

```bash
WORKSPACE=<absolute path that contains all parquets, outputs, and the source-pool images>
DOCKER="docker run --gpus all --rm --ipc=host -v $WORKSPACE:$WORKSPACE -w $WORKSPACE $DS_IMAGE"
```

Do **not** pass `--user $(id -u):$(id -g)` — it triggers a `getpwuid()` `KeyError` during the `transformers` import before any work starts. The container runs as root; chown outputs back to the host UID afterward.

Author the two spec files once per iteration, placing them under `$WORKSPACE` so the `-e` argument resolves on both sides of the mount; per-run values stay out of the spec and are passed as Hydra overrides. If the source pool is a CSV, convert it to parquet up front (preserving `filepath`, and `label` if present). The default `embedding_spec.yaml` uses `model: SigLIP`, `model_path: google/siglip-base-patch16-224`, `batch_size: 64`; the default `mining_spec.yaml` uses `topn: 5`, `knn_metric: cosine`, `filter_by_label: "false"` (quoted — the schema reads it as a string).

See `references/setup.md` for the full environment notes, `TAO_SKILL_BANK_PATH` handling, the path-mounting rationale, the `getpwuid` chown workaround, the CSV-to-parquet snippet, and the verbatim spec-file authoring blocks.

---

## Method

Three commands, in order. Each command's output parquet is the next command's input. Run them as plain Bash; the `$DOCKER` alias from Setup handles the container, GPU, and mounts. Every invocation follows the same shape: `-e <spec>` for the baked-in defaults, then a handful of Hydra overrides for the run-specific paths.

### Step 1 — Embed the target images

```bash
$DOCKER embedding image_embeddings \
    -e <embedding_spec.yaml> \
    input_parquet=<target_parquet> \
    output_parquet=<target_embeddings_parquet>
```

Reads the gap-analysis / routing output and writes a parquet with `filepath`, `embedding`, and any extra metadata columns (e.g. `label`, `siamese_score`, `weakness`) carried forward verbatim from the input. Print the output schema (`pd.read_parquet(...).columns`) to stdout so the script-check hook can confirm the embedding column exists.

If you need to override `model` / `model_path` / `batch_size` for one run without editing the spec, append them as Hydra overrides (e.g. `model_path=...`).

### Step 2 — Embed the source pool

```bash
$DOCKER embedding image_embeddings \
    -e <embedding_spec.yaml> \
    input_parquet=<source_pool_parquet> \
    output_parquet=<source_embeddings_parquet>
```

Same command shape as Step 1, applied to the source pool. Use the **identical** `embedding_spec.yaml` as Step 1, and do not override `model` / `model_path` / `batch_size` differently here — mismatched encoder configs across the two steps produce non-comparable embeddings.

### Step 3 — Mine nearest neighbours

```bash
$DOCKER tmm nearest_neighbors \
    -e <mining_spec.yaml> \
    source_parquet=<source_embeddings_parquet> \
    target_parquet=<target_embeddings_parquet> \
    output_parquet=<mined_parquet>
```

For each target embedding, finds the `topn` closest source embeddings under the chosen metric, deduplicates across targets, and writes a single-column (`filepath`) parquet of unique mined source paths. The container also drops a `mining_summary.txt` next to the output parquet with: query count, neighbour count, duplicates removed, and (when label filtering is on) kept-vs-dropped pair counts. Tweak `topn`, `knn_metric`, or `filter_by_label` via inline Hydra override when sweeping (e.g. `topn=10`) — no need to rewrite the spec.

When `filter_by_label=true` but one of the embedding parquets is missing the `label` column, the container logs a warning and proceeds without filtering. If the mined output looks larger than expected or contains cross-label pairs, scan the docker log for that warning before assuming the task did the right thing.

See `references/reference-invocation.md` for the minimal paste-and-edit end-to-end recipe (resolves `$DS_IMAGE`, writes both specs, runs all three steps, chowns outputs, and prints row counts) to run as a single streamed Bash block.

---

## Outputs and report

Write everything into a timestamped folder under the experiment / iteration directory. Get the real timestamp by running `date +%Y-%m-%d_%H%M%S` in Bash — do NOT hardcode or guess. If the user specifies a custom output path, use it directly but maintain the same internal layout. The packaging hook adds `mining_config/` and `claude_session.jsonl` automatically when `Mining_Report.md` is written.

The mined parquet is the artifact downstream training consumes. The two embedding parquets are intermediate but worth retaining — reusable across multiple mining runs against the same source pool, and the only place to look when a "looks unrelated" report needs encoder-level debugging.

See `references/outputs-and-reporting.md` for the full output-directory layout and the verbatim `Mining_Report.md` template (Verdict, Inputs, Encoder Consistency, Mining Run, Per-Label Breakdown, Output Sanity, Recommended Actions; keep it 600–1200 words).

---

## Common pitfalls

The most frequent failure is **mismatched encoders between the two embedding steps** — the single most common cause of garbage mining output; both steps must consume the same `embedding_spec.yaml`. Other recurring traps: passing `--user` (the `getpwuid` `KeyError`), skipping an embedding step, a missing `label` column silently no-oping `filter_by_label=true`, spec files outside `$WORKSPACE`, unresolved `???` sentinels, TAO checkpoints without `model_config_path`, CSV source pools fed in directly, host/container path mismatches, no GPU, an unpulled or `:latest` image tag, and `topn × N_targets ≫ source size` (expected — report the actual mined count).

See `references/troubleshooting.md` for the full pitfall list with the exact errors, causes, and fixes.

---

## Execution Order

1. Resolve `DS_IMAGE` from `versions.yaml` (`images.tao_toolkit.data_services`), then run `docker info`, `nvidia-smi`, and `docker image inspect "$DS_IMAGE"` (pulling if missing) once to confirm the environment. Abort with a clear message if any fail.
2. Run `date +%Y-%m-%d_%H%M%S` to get the timestamp; create `<output_dir>/mining_results/<timestamp>/`.
3. Write `embedding_spec.yaml` and `mining_spec.yaml` into the timestamped dir, filling in the encoder choice and mining knobs. Keep these under `$WORKSPACE` so the `-e` path resolves inside the container.
4. If the source pool is a CSV, convert to parquet first (preserve `filepath` and `label`).
5. Run Step 1 (embed targets) via `docker run … embedding image_embeddings -e embedding_spec.yaml input_parquet=… output_parquet=…`. Print the output parquet's row count and columns to stdout.
6. Run Step 2 (embed source pool) with the **identical** `embedding_spec.yaml` as Step 1. Print output row count and columns.
7. Run Step 3 (mine nearest neighbours) via `docker run … tmm nearest_neighbors -e mining_spec.yaml source_parquet=… target_parquet=… output_parquet=…`. Confirm `mining_summary.txt` was written next to `mined.parquet`.
8. Compute the per-label breakdown (Section 5) by joining the target embeddings parquet with the mined output on filepath, if both carry `label`.
9. Write `Mining_Report.md` last — writing it triggers the packaging hook, which copies session logs and skill config alongside.

安装 tao-mine-aoi-images

下载技能文件并将其解压到 .claude/skills/ 目录中。

下载ZIP

克隆仓库并复制技能文件到您的项目中。

git clone https://github.com/NVIDIA/skills/tree/main/skills/tao-mine-aoi-images # Copy SKILL.md to your .claude/skills/ directory

复制 复制
快速设置: 将技能文件夹复制到 .claude/skills/ Claude 将自动检测并使用该技能
仓库 NVIDIA/skills

相关技能

microservices-patterns
更新时间 2026-06-29
jpa-patterns
更新时间 2026-06-30
fabric-lakehouse
更新时间 2026-06-30
prisma-expert
更新时间 2026-06-29
OR