tao-mine-aoi-images
NVIDIA/skills
將目標與來源影像的 Parquet 檔案嵌入系統,隨後在 VCN AOI 工作流程中,從最近鄰來源影像中提取資料以進行增強。
...展開全部DEFT 挖掘與嵌入技能
您是 VCN AOI「先嵌入、再挖掘」工作流程的操作員。 您的任務是接收一組弱目標影像(間隙分析或路由輸出)的 Parquet 檔案,以及一個來源池,進而產生一組已去重且與目標影像相似的「挖掘後來源影像」Parquet 檔案——以便投入下一輪訓練。
此工作流程是固定且確定性的:先對目標圖像進行嵌入,接著對來源圖像池進行嵌入,最後挖掘最近鄰。每個步驟產出的 Parquet 檔案即為下一個步驟的輸入。 此流程不包含迭代搜尋、不進行聚類處理,亦無人工介入的選取環節——深度源自於選用正確的編碼器與適當的Top-n 閾值,而非來自多階段的調查。
整個技能僅是對 `versions.yaml` 中宣告的 `tao_toolkit.data_services` 映像檔進行三次直接Docker 執行呼叫的輕量級封裝(於執行時解析 — 詳見「設定」)。 容器的入口點接受— 將嵌入式資料 image_embeddings -e用於嵌入,並將tmm 最近鄰演算法 -e用於挖掘。-e旗標指向一個 YAML 檔案,該檔案為子任務的架構提供預設值;其後的所有內容皆為純粹的 Hydra 覆寫設定(key=value),可針對每次執行選擇性地覆寫規格欄位。 (容器內沒有「dataset」關鍵字——那是 TAO 啟動器的支柱前綴,此處已省略。)若映像檔未被快取,請拉取一次:docker pull "$DS_IMAGE"(在根據「Setup」解析$DS_IMAGE之後)。
模式鍵名可能會在不同版本的 data-services 之間更名(例如 RCA 技能中的inference_csv變更為inference_results_dir ,output_dir變更為results_dir)。 如有疑問,請針對每個映像檔檢查實際模式一次:docker run --rm "$DS_IMAGE" embedding image_embeddings --cfg=job以及... tmm nearest_neighbors --cfg=job。
輸入
- 目標 Parquet 檔案— 差距分析的輸出結果,通常為來自
tao-route-visual-changenet-samples 的mining_gaps.parquet(或若跳過路由步驟,則為來自tao-analyze-gaps-visual-changenet的gaps.parquet)。 必填欄位:filepath。若同時存在label欄位,則可在挖掘過程中進行標籤感知過濾;否則挖掘任務會靜默地將過濾器設為無操作。 - 來源資料庫— 用於挖掘的候選影像 Parquet 檔案,其中包含
filepath欄位。若使用者僅有 CSV 檔案,請在步驟 2 之前將其轉換為具有相同欄位的Parquet 檔案。若要進行標籤感知篩選,資料庫還必須包含label欄位。 - 嵌入式規格檔— 一個包含
model、model_path、batch_size以及(僅當model_path為 TAO 的.pth/.ckpt時)model_config_path的 YAML 檔案。 此規格檔在步驟 1 和步驟 2 中重複使用;input_parquet/output_parquet會於每次執行時由 Hydra 覆寫提供。同一份規格檔必須同時驅動兩個嵌入步驟 — 來自不同編碼器的嵌入向量無法相互比較,而編碼器不匹配是導致「挖掘出的圖片看似無關」報告的最常見原因。 - 挖掘規格檔— 一份包含
topn、knn_metric、filter_by_label以及(極少變更的)source_embed_column_name/target_embed_column_name的 YAML 檔案。source_parquet/target_parquet/output_parquet為執行時由 Hydra 覆寫的參數。 SigLIP 和 CLIP 嵌入向量應使用knn_metric: cosine。當filter_by_label: true,但任一嵌入向量 Parquet 檔案缺少標籤欄位時,容器會記錄警告並在不進行篩選的情況下繼續執行。
設定
在執行流程的起始階段,請先從versions.yaml檔案中解析出具體的tao_toolkit.data_servicesURI,接著在執行其他操作之前,請確認已安裝 Docker、NVIDIA 容器工具包以及 GPU。 編碼器的前向傳播以及 cuML/cuDF 的 k-NN 搜尋均需 GPU 支援;若無 CUDA,這兩項步驟皆會失敗。
# 從 versions.yaml 解析 tao_toolkit.data_services → 具體的 nvcr.io/... URI
DS_IMAGE=$(python3 -c "import yaml,os; print(yaml.safe_load(open(os.environ['TAO_SKILL_BANK_PATH']+'/versions.yaml'))['images']['tao_toolkit']['data_services'])")
echo "DS_IMAGE=$DS_IMAGE"
docker info > /dev/null && echo "OK: docker"
nvidia-smi > /dev/null && echo "OK: GPU"
docker image inspect "$DS_IMAGE" > /dev/null \
|| docker pull "$DS_IMAGE"
容器讀取或寫入的每個主機路徑都必須透過綁定掛載。最可預測的做法是將工作區根目錄掛載為容器內外路徑完全一致,然後針對這三個呼叫重用一個$DOCKER別名:
WORKSPACE=
DOCKER="docker run --gpus all --rm --ipc=host -v $WORKSPACE:$WORKSPACE -w $WORKSPACE $DS_IMAGE"
請勿傳入--user $(id -u):$(id -g)— 這會在任何工作開始前,於導入Transformer時觸發getpwuid()的KeyError 錯誤。容器以 root 身分執行;執行 chown 後,權限會回歸至主機的 UID。
每輪迭代僅需編寫兩份 spec 檔案一次,並將其放置於$WORKSPACE目錄下,以便-e參數能在掛載的兩側解析;每次執行所需的值不包含在 spec 中,而是作為 Hydra 的覆寫參數傳入。 若來源資料池為 CSV 檔案,請預先將其轉換為 Parquet 格式(保留檔案路徑,以及若有標籤則一併保留)。預設的embedding_spec.yaml使用model: SigLIP、model_path: google/siglip-base-patch16-224、batch_size: 64; 預設的mining_spec.yaml採用topn: 5,knn_metric: cosine,filter_by_label: "false"(需加引號——因為資料結構會將其視為字串)。
請參閱references/setup.md以獲取完整的環境說明、TAO_SKILL_BANK_PATH的處理方式、路徑掛載的理由、getpwuid 的chown 解決方案、CSV 轉 Parquet 的程式碼片段,以及規格檔案編寫的逐字範例區塊。
方法
依序執行三條指令。每條指令產出的 Parquet 檔案即為下一條指令的輸入。請以純 Bash 方式執行這些指令;Setup 中的$DOCKER別名會自動處理容器、GPU 及掛載設定。 每次執行皆遵循相同格式:先使用-e套用內建預設值,接著透過少數 Hydra 覆寫設定來指定執行時所需的特定路徑。
步驟 1 — 嵌入目標映像
$DOCKER embedding image_embeddings \
-e \
input_parquet= \
output_parquet=
讀取差距分析/路由輸出,並寫入一個 Parquet 檔案,其中包含檔案路徑、嵌入資料,以及從輸入原樣傳遞過來的任何額外元資料欄位(例如label、siamese_score、weakness)。 將輸出資料結構 (pd.read_parquet(...).columns) 輸出至標準輸出 (stdout),以便 script-check 掛鉤能確認嵌入向量欄位確實存在。
若需在不修改規格檔的情況下,針對單次執行覆寫model/model_path/batch_size,請將其作為 Hydra 覆寫參數附加(例如model_path=...)。
步驟 2 — 嵌入來源資料池
$DOCKER embedding image_embeddings \
-e \
input_parquet= \
output_parquet=
與步驟 1 相同的命令結構,應用於來源資料池。請使用與步驟 1完全相同的 embedding_spec.yaml,且在此處請勿對model/model_path/batch_size進行不同設定 — 若兩個步驟間的編碼器設定不一致,將產生無法相互比對的嵌入向量。
步驟 3 — 挖掘最近鄰
$DOCKER tmm nearest_neighbors \
-e \
source_parquet= \
target_parquet= \
output_parquet=
針對每個目標嵌入向量,根據選定的度量標準找出前 n 個最接近的來源嵌入向量,在各目標之間進行去重處理,並將挖掘出的唯一來源路徑寫入單欄(檔案路徑)Parquet 檔案中。 該容器還會於輸出 Parquet 檔案旁產生一個mining_summary.txt檔案,內容包含:查詢次數、鄰接節點數、已移除的重複項,以及(當標籤過濾開啟時)保留與捨棄的配對數量。 在掃描時,可透過內嵌的 Hydra 覆寫功能調整topn、knn_metric 或filter_by_label(例如topn=10)——無需重寫規格文件。
當filter_by_label=true,但其中一個嵌入式 Parquet 檔案缺少標籤欄位時,容器會記錄警告訊息並在不進行篩選的情況下繼續執行。若挖掘出的輸出量大於預期,或包含跨標籤配對,請先檢查 Docker 日誌中的該警告訊息,再判定任務是否執行正確。
請參閱references/reference-invocation.md,其中提供最簡化的「複製貼上並編輯」端到端操作指南(解決$DS_IMAGE、寫入兩份規格檔、執行全部三個步驟、變更輸出檔案所有權,並列印行數),可作為單一串流 Bash 區塊執行。
輸出與報告
將所有內容寫入實驗 / 迭代目錄下的帶有時間戳記的資料夾中。請在 Bash 中執行`date +%Y-%m-%d_%H%M%S` 以取得真實的時間戳記 — 切勿硬編碼或憑空猜測。若使用者指定了自訂輸出路徑,請直接使用該路徑,但須維持相同的內部結構。 當寫入Mining_Report.md時,封裝鉤子會自動新增mining_config/目錄及claude_session.jsonl 檔案。
挖掘產出的 Parquet 檔案是下游訓練所使用的產出成果。那兩個嵌入式 Parquet 檔案雖屬中間產出,但仍值得保留——它們可在針對同一來源資料池的多個挖掘執行中重複使用,且當報告顯示「看似無關」時,這是進行編碼器層級除錯的唯一參考來源。
請參閱references/outputs-and-reporting.md以了解完整的輸出目錄結構,以及Mining_Report.md模板的完整內容(結論、輸入、編碼器一致性、挖掘執行、各標籤細項分析、輸出合理性、建議行動;篇幅應控制在 600–1200 字之間)。
常見陷阱
最常見的失敗原因在於兩個嵌入步驟之間的編碼器不匹配——這是產生垃圾挖掘輸出最主要的單一原因;兩個步驟都必須使用相同的embedding_spec.yaml。 其他常見陷阱:傳入--user 參數(導致getpwuid KeyError)、跳過嵌入步驟、標籤欄位缺失導致filter_by_label=true 靜默無效、spec 檔案位於$WORKSPACE 以外、未解析的???哨兵未解析、缺乏model_config_path 的 TAO 檢查點、直接導入 CSV 來源資料池、主機/容器路徑不符、無 GPU、未拉取或使用:latest映像標籤,以及topn × N_targets ≫ 來源大小(此為預期行為——請回報實際挖掘的計數)。
請參閱references/troubleshooting.md以查看完整的陷阱清單,其中包含確切的錯誤、原因及解決方法。
執行順序
- 從
versions.yaml(images.tao_toolkit.data_services)解析DS_IMAGE,然後執行docker info、nvidia-smi以及docker image inspect "$DS_IMAGE"(若不存在則進行拉取)一次,以確認環境。若任何步驟失敗,請以清晰訊息終止執行。 - 執行
`date +%Y-%m-%d_%H%M%S` 以取得時間戳記;建立 `` 目錄。/mining_results/` 及 ` / - 將
`embedding_spec.yaml`和`mining_spec.yaml`寫入帶有時戳的目錄中,並填入編碼器選項及挖掘參數。將這些檔案存放於$WORKSPACE目錄下,以便在容器內能解析`-e` 路徑。 - 若來源資料集為 CSV 格式,請先轉換為 Parquet 格式(保留
檔案路徑與標籤)。 - 透過 `
docker run … embedding image_embeddings -e embedding_spec.yaml input_parquet=… output_parquet=…`執行步驟 1(嵌入目標資料)。將輸出 Parquet 檔案的列數與欄數輸出至標準輸出(stdout)。 - 使用與步驟 1完全相同的
embedding_spec.yaml執行步驟 2(嵌入來源資料池)。將輸出 Parquet 檔案的列數和欄數輸出至標準輸出。 - 透過 `
docker run … tmm nearest_neighbors -e mining_spec.yaml source_parquet=… target_parquet=… output_parquet=…`執行步驟 3(挖掘最近鄰)。確認`mining_summary.txt`已寫入至`mined.parquet` 旁。 - 若目標嵌入資料的 Parquet 檔案與挖掘產出檔案均包含
標籤,則透過檔案路徑將兩者進行聯結,以計算各標籤的細項分布(第 5 節)。 - 最後寫入
Mining_Report.md— 寫入此檔案會觸發封裝鉤子,該鉤子會一併複製工作階段日誌與技能設定。
---
name: tao-mine-aoi-images
description: Embeds target and source image parquets, then mines nearest-neighbour source images for augmentation in VCN AOI workflows.
license: Apache-2.0
---
# DEFT Mining and Embedding Skill
You are the operator of the DEFT embed-then-mine workflow for VCN AOI. Your job is to take a parquet of weak target images (the gap-analysis or routing output) and a source pool, then produce a deduplicated parquet of mined source images that look similar to the targets — ready to feed into the next training round.
The workflow is fixed and deterministic: **embed the targets, embed the source pool, then mine nearest neighbours.** Each step's output parquet is the next step's input. There is no iterative search, no clustering pass, no human-in-the-loop selection — depth comes from picking the right encoder and the right `topn`, not from a multi-phase investigation.
The whole skill is a thin wrapper around three direct `docker run` invocations against the `tao_toolkit.data_services` image declared in `versions.yaml` (resolved at runtime — see Setup). The container's entrypoint takes `<category> <action> -e <spec.yaml> [hydra overrides...]` — pass `embedding image_embeddings -e <embedding_spec.yaml> …` for embedding and `tmm nearest_neighbors -e <mining_spec.yaml> …` for mining. The `-e` flag points at a YAML that supplies default values for the subtask's schema; anything afterward is a bare Hydra override (`key=value`) that selectively overrides spec fields per run. (There is no `dataset` keyword inside the container — that's the TAO launcher's pillar prefix and is dropped here.) Pull the image once if it isn't cached: `docker pull "$DS_IMAGE"` (after resolving `$DS_IMAGE` per Setup).
Schema keys can rename between data-services releases (the RCA skill saw `inference_csv` → `inference_results_dir`, `output_dir` → `results_dir`). When in doubt, introspect the actual schema once per image: `docker run --rm "$DS_IMAGE" embedding image_embeddings --cfg=job` and `... tmm nearest_neighbors --cfg=job`.
---
## Inputs
1. **Target parquet** — the gap-analysis output, typically `mining_gaps.parquet` from `tao-route-visual-changenet-samples` (or `gaps.parquet` from `tao-analyze-gaps-visual-changenet` if routing was skipped). Required column: `filepath`. If `label` is also present, label-aware filtering during mining is available; otherwise the mining task silently no-ops the filter.
2. **Source pool** — a parquet of candidate images to mine against, with a `filepath` column. If the user only has a CSV, convert it to a parquet **with the same columns** before Step 2. For label-aware filtering, the pool must also carry a `label` column.
3. **Embedding spec file** — a YAML containing `model`, `model_path`, `batch_size`, and (only when `model_path` is a TAO `.pth`/`.ckpt`) `model_config_path`. Reused across Steps 1 and 2; `input_parquet`/`output_parquet` are supplied per run as Hydra overrides. The **same** spec MUST drive both embedding steps — embeddings from different encoders are not comparable, and mismatched encoders are the most common cause of "the mined images look unrelated" reports.
4. **Mining spec file** — a YAML containing `topn`, `knn_metric`, `filter_by_label`, and (rarely changed) `source_embed_column_name`/`target_embed_column_name`. `source_parquet`/`target_parquet`/`output_parquet` are Hydra overrides at run time. SigLIP and CLIP embeddings should use `knn_metric: cosine`. When `filter_by_label: true` but either embedding parquet lacks a `label` column, the container logs a warning and proceeds **without** filtering.
---
## Setup
Resolve the concrete `tao_toolkit.data_services` URI from `versions.yaml` once at the top of the run, then confirm Docker, the NVIDIA container toolkit, and a GPU are present before doing anything else. A GPU is required for both the encoder forward pass and the cuML/cuDF k-NN search; both steps fail without CUDA.
```bash
# Resolve tao_toolkit.data_services → concrete nvcr.io/... URI from versions.yaml
DS_IMAGE=$(python3 -c "import yaml,os; print(yaml.safe_load(open(os.environ['TAO_SKILL_BANK_PATH']+'/versions.yaml'))['images']['tao_toolkit']['data_services'])")
echo "DS_IMAGE=$DS_IMAGE"
docker info > /dev/null && echo "OK: docker"
nvidia-smi > /dev/null && echo "OK: GPU"
docker image inspect "$DS_IMAGE" > /dev/null \
|| docker pull "$DS_IMAGE"
```
Every host path the container reads or writes must be bind-mounted. The most predictable approach mounts the workspace root with **identical paths** inside and outside the container, then reuses one `$DOCKER` alias for the three invocations:
```bash
WORKSPACE=<absolute path that contains all parquets, outputs, and the source-pool images>
DOCKER="docker run --gpus all --rm --ipc=host -v $WORKSPACE:$WORKSPACE -w $WORKSPACE $DS_IMAGE"
```
Do **not** pass `--user $(id -u):$(id -g)` — it triggers a `getpwuid()` `KeyError` during the `transformers` import before any work starts. The container runs as root; chown outputs back to the host UID afterward.
Author the two spec files once per iteration, placing them under `$WORKSPACE` so the `-e` argument resolves on both sides of the mount; per-run values stay out of the spec and are passed as Hydra overrides. If the source pool is a CSV, convert it to parquet up front (preserving `filepath`, and `label` if present). The default `embedding_spec.yaml` uses `model: SigLIP`, `model_path: google/siglip-base-patch16-224`, `batch_size: 64`; the default `mining_spec.yaml` uses `topn: 5`, `knn_metric: cosine`, `filter_by_label: "false"` (quoted — the schema reads it as a string).
See `references/setup.md` for the full environment notes, `TAO_SKILL_BANK_PATH` handling, the path-mounting rationale, the `getpwuid` chown workaround, the CSV-to-parquet snippet, and the verbatim spec-file authoring blocks.
---
## Method
Three commands, in order. Each command's output parquet is the next command's input. Run them as plain Bash; the `$DOCKER` alias from Setup handles the container, GPU, and mounts. Every invocation follows the same shape: `-e <spec>` for the baked-in defaults, then a handful of Hydra overrides for the run-specific paths.
### Step 1 — Embed the target images
```bash
$DOCKER embedding image_embeddings \
-e <embedding_spec.yaml> \
input_parquet=<target_parquet> \
output_parquet=<target_embeddings_parquet>
```
Reads the gap-analysis / routing output and writes a parquet with `filepath`, `embedding`, and any extra metadata columns (e.g. `label`, `siamese_score`, `weakness`) carried forward verbatim from the input. Print the output schema (`pd.read_parquet(...).columns`) to stdout so the script-check hook can confirm the embedding column exists.
If you need to override `model` / `model_path` / `batch_size` for one run without editing the spec, append them as Hydra overrides (e.g. `model_path=...`).
### Step 2 — Embed the source pool
```bash
$DOCKER embedding image_embeddings \
-e <embedding_spec.yaml> \
input_parquet=<source_pool_parquet> \
output_parquet=<source_embeddings_parquet>
```
Same command shape as Step 1, applied to the source pool. Use the **identical** `embedding_spec.yaml` as Step 1, and do not override `model` / `model_path` / `batch_size` differently here — mismatched encoder configs across the two steps produce non-comparable embeddings.
### Step 3 — Mine nearest neighbours
```bash
$DOCKER tmm nearest_neighbors \
-e <mining_spec.yaml> \
source_parquet=<source_embeddings_parquet> \
target_parquet=<target_embeddings_parquet> \
output_parquet=<mined_parquet>
```
For each target embedding, finds the `topn` closest source embeddings under the chosen metric, deduplicates across targets, and writes a single-column (`filepath`) parquet of unique mined source paths. The container also drops a `mining_summary.txt` next to the output parquet with: query count, neighbour count, duplicates removed, and (when label filtering is on) kept-vs-dropped pair counts. Tweak `topn`, `knn_metric`, or `filter_by_label` via inline Hydra override when sweeping (e.g. `topn=10`) — no need to rewrite the spec.
When `filter_by_label=true` but one of the embedding parquets is missing the `label` column, the container logs a warning and proceeds without filtering. If the mined output looks larger than expected or contains cross-label pairs, scan the docker log for that warning before assuming the task did the right thing.
See `references/reference-invocation.md` for the minimal paste-and-edit end-to-end recipe (resolves `$DS_IMAGE`, writes both specs, runs all three steps, chowns outputs, and prints row counts) to run as a single streamed Bash block.
---
## Outputs and report
Write everything into a timestamped folder under the experiment / iteration directory. Get the real timestamp by running `date +%Y-%m-%d_%H%M%S` in Bash — do NOT hardcode or guess. If the user specifies a custom output path, use it directly but maintain the same internal layout. The packaging hook adds `mining_config/` and `claude_session.jsonl` automatically when `Mining_Report.md` is written.
The mined parquet is the artifact downstream training consumes. The two embedding parquets are intermediate but worth retaining — reusable across multiple mining runs against the same source pool, and the only place to look when a "looks unrelated" report needs encoder-level debugging.
See `references/outputs-and-reporting.md` for the full output-directory layout and the verbatim `Mining_Report.md` template (Verdict, Inputs, Encoder Consistency, Mining Run, Per-Label Breakdown, Output Sanity, Recommended Actions; keep it 600–1200 words).
---
## Common pitfalls
The most frequent failure is **mismatched encoders between the two embedding steps** — the single most common cause of garbage mining output; both steps must consume the same `embedding_spec.yaml`. Other recurring traps: passing `--user` (the `getpwuid` `KeyError`), skipping an embedding step, a missing `label` column silently no-oping `filter_by_label=true`, spec files outside `$WORKSPACE`, unresolved `???` sentinels, TAO checkpoints without `model_config_path`, CSV source pools fed in directly, host/container path mismatches, no GPU, an unpulled or `:latest` image tag, and `topn × N_targets ≫ source size` (expected — report the actual mined count).
See `references/troubleshooting.md` for the full pitfall list with the exact errors, causes, and fixes.
---
## Execution Order
1. Resolve `DS_IMAGE` from `versions.yaml` (`images.tao_toolkit.data_services`), then run `docker info`, `nvidia-smi`, and `docker image inspect "$DS_IMAGE"` (pulling if missing) once to confirm the environment. Abort with a clear message if any fail.
2. Run `date +%Y-%m-%d_%H%M%S` to get the timestamp; create `<output_dir>/mining_results/<timestamp>/`.
3. Write `embedding_spec.yaml` and `mining_spec.yaml` into the timestamped dir, filling in the encoder choice and mining knobs. Keep these under `$WORKSPACE` so the `-e` path resolves inside the container.
4. If the source pool is a CSV, convert to parquet first (preserve `filepath` and `label`).
5. Run Step 1 (embed targets) via `docker run … embedding image_embeddings -e embedding_spec.yaml input_parquet=… output_parquet=…`. Print the output parquet's row count and columns to stdout.
6. Run Step 2 (embed source pool) with the **identical** `embedding_spec.yaml` as Step 1. Print output row count and columns.
7. Run Step 3 (mine nearest neighbours) via `docker run … tmm nearest_neighbors -e mining_spec.yaml source_parquet=… target_parquet=… output_parquet=…`. Confirm `mining_summary.txt` was written next to `mined.parquet`.
8. Compute the per-label breakdown (Section 5) by joining the target embeddings parquet with the mined output on filepath, if both carry `label`.
9. Write `Mining_Report.md` last — writing it triggers the packaging hook, which copies session logs and skill config alongside.
所有檔案
14 個檔案安裝 tao-mine-aoi-images
請下載並將技能檔案解壓縮至您的 .claude/skills/ 目錄中。
下載 ZIP複製儲存庫並將技能檔案複製到您的專案中。
git clone https://github.com/NVIDIA/skills/tree/main/skills/tao-mine-aoi-images # Copy SKILL.md to your .claude/skills/ directory
複製





首頁
