オプション
家家 Skill データベース管理 tao-mine-aoi-images

tao-mine-aoi-images

NVIDIA/skills NVIDIA/skills

VCN AOIワークフローにおいて、ターゲット画像とソース画像のParquetファイルを埋め込み、増強処理のためにソース画像の中から最近傍の画像を抽出します。

...すべて拡張します
26
更新された時間 2026年9月24日

DEFTのマイニングおよび埋め込みスキル

あなたは、VCN AOI向けのDEFT「埋め込み→マイニング」ワークフローのオペレーターです。 あなたの役割は、ターゲット画像(ギャップ分析または配線結果)のパーケットとソースプールを受け取り、ターゲット画像に類似したソース画像をマイニングして重複を除去したパーケットを生成し、次のトレーニングラウンドに投入できる状態にすることです。

このワークフローは固定されており、決定論的です。ターゲットを埋め込み、ソースプールを埋め込み、その後、最近傍をマイニングします。各ステップの出力Parquetが、次のステップの入力となります。 反復的な検索も、クラスタリング処理も、人間による介入(human-in-the-loop)による選択もありません。深さは、多段階の調査からではなく、適切なエンコーダーと適切なtopn を選択することから得られます。

このスキル全体は、versions.yamlで宣言されたtao_toolkit.data_servicesイメージ(実行時に解決される — 「セットアップ」を参照)に対する3つの直接的なDocker実行呼び出しを薄くラッピングしたものです。 コンテナのエントリポイントは、 -e [hydraのオーバーライド...]を受け取ります。これに対し、埋め込み処理には image_embeddings -e … を、マイニング処理にはtmm nearest_neighbors -e …を渡します。-eフラグは、サブタスクのスキーマに対するデフォルト値を指定する YAML ファイルを指します。その後に続く内容はすべて、実行ごとに仕様フィールドを選択的に上書きする単純な Hydra 上書き設定 (key=value) です。 (コンテナ内には「dataset」キーワードは存在しません。これは TAO ランチャーの pillar プレフィックスであり、ここでは省略されています。)イメージがキャッシュされていない場合は、一度プルしてください:docker pull "$DS_IMAGE"(セットアップに従って$DS_IMAGEを解決した後)。

スキーマのキー名は、data-servicesのリリース間で変更される場合があります(RCAスキルでは、inference_csv→inference_results_dir、output_dir→results_dir となりました)。 不明な場合は、イメージごとに実際のスキーマを一度確認してください:docker run --rm "$DS_IMAGE" embedding image_embeddings --cfg=jobおよび... tmm nearest_neighbors --cfg=job。

入力

  1. ターゲット Parquet— ギャップ分析の出力。通常は、tao-route-visual-changenet-samplesからのmining_gaps.parquetです(ルーティングがスキップされた場合は、tao-analyze-gaps-visual-changenetからのgaps.parquet)。 必須の列:filepath。label列も存在する場合、マイニング中にラベルを考慮したフィルタリングが可能になります。そうでない場合、マイニングタスクはフィルタ処理を黙ってスキップします。
  2. ソースプール— マイニングの対象となる候補画像の Parquet ファイルで、filepath列が含まれている必要があります。CSV ファイルしか持っていない場合は、ステップ 2 の前に、同じ列を持つParquet ファイルに変換してください。ラベルを意識したフィルタリングを行うには、プールにもlabel列が含まれている必要があります。
  3. 埋め込み仕様ファイル—model、model_path、batch_size、および(model_pathがTAOの.pth/.ckptファイルである場合のみ)model_config_pathを含むYAMLファイル。 ステップ1と2で共用されます。input_parquet/output_parquetは、Hydraによる上書きとして実行ごとに指定されます。両方の埋め込みステップは、必ず同じ仕様で実行する必要があります。異なるエンコーダーからの埋め込みは比較不可能であり、エンコーダーの不一致は「マイニングされた画像が関連性がないように見える」という報告の最も一般的な原因です。
  4. マイニング仕様ファイル—topn、knn_metric、filter_by_label、および(めったに変更されない)source_embed_column_name/target_embed_column_name を含む YAML ファイル。source_parquet/target_parquet/output_parquet は、実行時に Hydra によって上書きされます。 SigLIPおよびCLIPの埋め込みデータには、knn_metric: cosine を使用する必要があります。filter_by_label: trueであるにもかかわらず、いずれかの埋め込みParquetにラベル列がない場合、コンテナは警告をログに記録し、フィルタリングを行わずに処理を続行します。

セットアップ

実行の最初にversions.yamlから具体的なtao_toolkit.data_servicesURI を一度解決し、その後、他の処理を行う前に Docker、NVIDIA コンテナツールキット、および GPU が存在することを確認してください。 エンコーダのフォワードパスと cuML/cuDF による k-NN 検索の両方に GPU が必要です。CUDA がない場合、どちらのステップも失敗します。

# versions.yaml から tao_toolkit.data_services → 具体的な nvcr.io/... URI を解決
DS_IMAGE=$(python3 -c "import yaml,os; print(yaml.safe_load(open(os.environ['TAO_SKILL_BANK_PATH']+'/versions.yaml'))['images']['tao_toolkit']['data_services'])")
echo "DS_IMAGE=$DS_IMAGE"

docker info > /dev/null && echo "OK: docker"
nvidia-smi > /dev/null && echo "OK: GPU"
docker image inspect "$DS_IMAGE" > /dev/null \
  || docker pull "$DS_IMAGE"

コンテナが読み書きするすべてのホストパスは、バインドマウントする必要があります。最も確実なアプローチは、ワークスペースのルートを、コンテナ内外でパスが同一になるようにマウントし、3つの呼び出しで1つの$DOCKERエイリアスを再利用することです:

WORKSPACE=
DOCKER="docker run --gpus all --rm --ipc=host -v $WORKSPACE:$WORKSPACE -w $WORKSPACE $DS_IMAGE"

--user $(id -u):$(id -g)を指定しないでください。これにより、作業が開始される前に、transformersのインポート中にgetpwuid()によるKeyErrorが発生します。コンテナは root として実行され、chown によって実行後にホストの UID が復元されます。

2つのspecファイルは1回の反復ごとに1回作成し、$WORKSPACE配下に配置して、-e引数がマウントの両側で解決されるようにします。実行ごとの値はspecには含めず、Hydraのオーバーライドとして渡します。 ソースプールがCSVの場合は、事前にParquet形式に変換してください(ファイルパスと、存在する場合はラベルを保持します)。デフォルトのembedding_spec.yamlでは、model: SigLIP、model_path: google/siglip-base-patch16-224、batch_size: 64が使用されています。 デフォルトの `mining_spec.yaml` では、`topn`: 5、`knn_metric`: cosine、`filter_by_label`: "false"(引用符付き — スキーマでは文字列として解釈されます)が使用されます。

完全な環境に関する注意事項、TAO_SKILL_BANK_PATHの取り扱い、パスマウントの根拠、getpwuidのchownによる回避策、CSVからParquetへの変換スニペット、および仕様ファイル作成のブロック原文については、references/setup.mdを参照してください。

方法

3つのコマンドを順に実行します。各コマンドの出力となるParquetファイルが、次のコマンドの入力となります。これらは通常のBashとして実行してください。セットアップで定義された$DOCKERエイリアスが、コンテナ、GPU、およびマウントを処理します。 各実行はすべて同じ形式に従います。まず、組み込みのデフォルト設定を指定する`-e` を実行し、続いて実行固有のパスに対してHydraによる上書き設定をいくつか行います。

ステップ 1 — ターゲットイメージを埋め込む

$DOCKER embedding image_embeddings \
    -e \
    input_parquet= \
    output_parquet=

ギャップ分析/ルーティングの出力を読み取り、ファイルパス、埋め込み、および入力からそのまま引き継がれた追加のメタデータ列(例:label、siamese_score、weakness)を含むParquetファイルを出力します。 出力スキーマ(pd.read_parquet(...).columns)を標準出力に表示し、script-checkフックが埋め込みカラムの存在を確認できるようにします。

仕様を編集せずに 1 回の実行についてmodel/model_path/batch_size を上書きする必要がある場合は、それらを Hydra オーバーライドとして追加します(例:model_path=...)。

ステップ 2 — ソースプールを埋め込む

$DOCKER embedding image_embeddings \
    -e \
    input_parquet= \
    output_parquet=

ステップ1と同じコマンド形式で、ソースプールに適用します。ステップ1と全く同じ embedding_spec.yamlを使用し、ここではmodel/model_path/batch_size を異なる値に上書きしないでください。2つのステップ間でエンコーダの設定が一致しないと、比較不可能な埋め込みが生成されます。

ステップ3 — 最近傍の抽出

$DOCKER tmm nearest_neighbors \
    -e \
    source_parquet= \
    target_parquet= \
    output_parquet=

各ターゲット埋め込みについて、選択されたメトリックに基づき、最も近いソース埋め込みをtopn 個見つけ出し、ターゲット間で重複を排除し、抽出された一意のソースパスを単一列(ファイルパス)の Parquet ファイルとして書き出します。 また、このコンテナは出力Parquetファイルの隣に「mining_summary.txt」を生成します。このファイルには、クエリ数、近傍数、削除された重複数、および(ラベルフィルタリングが有効な場合)保持されたペア数と削除されたペア数が記載されています。 スイープ実行時に、インラインのHydraオーバーライドを介してtopn、knn_metric、またはfilter_by_labelを調整できます(例:topn=10)。仕様書を書き換える必要はありません。

filter_by_label=trueであるにもかかわらず、埋め込み Parquet のいずれかにラベル列が欠落している場合、コンテナは警告をログに記録し、フィルタリングを行わずに処理を続行します。マイニングされた出力が予想より大きかったり、ラベルが異なるペアが含まれていたりする場合は、タスクが正しく動作したと判断する前に、Docker ログでその警告を確認してください。

単一のストリーム化されたBashブロックとして実行するための、最小限の「貼り付けて編集するだけ」のエンドツーエンドレシピ($DS_IMAGEを解決し、両方の仕様ファイルを書き出し、3つのステップすべてを実行し、出力ファイルの所有権を変更し、行数を表示する)については、references/reference-invocation.mdを参照してください。

出力とレポート

すべての実験結果を実験 / 反復ディレクトリ下の、タイムスタンプ付きフォルダに書き出します。Bash で `date +%Y-%m-%d_%H%M%S` を実行して実際のタイムスタンプを取得してください。ハードコーディングや推測は絶対にしないでください。ユーザーがカスタム出力パスを指定した場合は、それを直接使用しますが、内部のレイアウトは同じまま維持してください。 パッケージングフックは、Mining_Report.mdが書き込まれると、自動的にmining_config/およびclaude_session.jsonl を追加します。

マイニングされた Parquet ファイルは、下流のトレーニングで消費される成果物です。2 つの埋め込み Parquet ファイルは中間ファイルですが、保持する価値があります。同じソースプールに対する複数のマイニング実行で再利用可能であり、「関連性がないように見える」レポートに対してエンコーダーレベルのデバッグが必要な場合に参照すべき唯一の場所だからです。

出力ディレクトリの完全なレイアウトおよび『Mining_Report.md』テンプレートの原文(Verdict、Inputs、Encoder Consistency、Mining Run、Per-Label Breakdown、Output Sanity、Recommended Actions;600~1200語に収めること)については、references/outputs-and-reporting.mdを参照してください。

よくある落とし穴

最も頻繁に発生する失敗は、2つの埋め込みステップ間でエンコーダーが一致していないことです。これは、マイニング出力がゴミになる最も一般的な原因です。両方のステップで同じembedding_spec.yaml を使用する必要があります。 その他のよくある落とし穴:--userの指定(getpwuid KeyError)、埋め込みステップのスキップ、ラベル列の欠落によるfilter_by_label=true の無音ノーオペ、$WORKSPACE 以外の場所にある spec ファイル、未解決の???センチネルの未解決、model_config_path を指定しない TAO チェックポイント、CSV ソースプールを直接読み込む、ホスト/コンテナのパス不一致、GPU の未装着、プルされていない、または:latestのイメージタグ、およびtopn × N_targets ≫ ソースサイズ(想定内 — 実際にマイニングされたカウントを報告してください)。

具体的なエラー、原因、および修正方法を含む完全な落とし穴リストについては、references/troubleshooting.md を参照してください。

実行順序

  1. versions.yaml(images.tao_toolkit.data_services)からDS_IMAGE を解決し、docker info、nvidia-smi、およびdocker image inspect "$DS_IMAGE"(存在しない場合はプルする)を 1 回実行して環境を確認します。いずれかが失敗した場合は、明確なメッセージを表示して中止します。
  2. `date +%Y-%m-%d_%H%M%S` を実行してタイムスタンプを取得し、/mining_results// を作成します。
  3. timestampedディレクトリにembedding_spec.yamlと mining_spec.yamlを書き出し、エンコーダの選択とマイニングのパラメータを指定します。これらは$WORKSPACE配下に配置し、コンテナ内で-eパスが解決されるようにします。
  4. ソースデータがCSVの場合は、まずParquet形式に変換します(ファイルパスとラベルは保持してください)。
  5. docker run … embedding image_embeddings -e embedding_spec.yaml input_parquet=… output_parquet=… を実行して、ステップ 1(ターゲットの埋め込み)を実行します。出力された Parquet の行数と列数を stdout に出力します。
  6. ステップ1と同じembedding_spec.yamlを使用して、ステップ2(ソースプールの埋め込み)を実行します。出力の行数と列数を標準出力に出力します。
  7. ステップ3(最近傍の抽出)を `docker run … tmm nearest_neighbors -e mining_spec.yaml source_parquet=… target_parquet=… output_parquet=…` で実行します。 mined.parquet の隣にmining_summary.txt が書き込まれたことを確認します。
  8. ターゲットの埋め込みデータ(parquet)とマイニング出力(mined)の両方にラベルが含まれている場合、ファイルパスでそれらを結合し、ラベルごとの内訳(第5節)を計算します。
  9. 最後にMining_Report.mdを書き出します。これを書き出すとパッケージングフックがトリガーされ、セッションログとスキル設定が併せてコピーされます。
GitHubで見る
---
name: tao-mine-aoi-images
description: Embeds target and source image parquets, then mines nearest-neighbour source images for augmentation in VCN AOI workflows.
license: Apache-2.0
---

# DEFT Mining and Embedding Skill

You are the operator of the DEFT embed-then-mine workflow for VCN AOI. Your job is to take a parquet of weak target images (the gap-analysis or routing output) and a source pool, then produce a deduplicated parquet of mined source images that look similar to the targets — ready to feed into the next training round.

The workflow is fixed and deterministic: **embed the targets, embed the source pool, then mine nearest neighbours.** Each step's output parquet is the next step's input. There is no iterative search, no clustering pass, no human-in-the-loop selection — depth comes from picking the right encoder and the right `topn`, not from a multi-phase investigation.

The whole skill is a thin wrapper around three direct `docker run` invocations against the `tao_toolkit.data_services` image declared in `versions.yaml` (resolved at runtime — see Setup). The container's entrypoint takes `<category> <action> -e <spec.yaml> [hydra overrides...]` — pass `embedding image_embeddings -e <embedding_spec.yaml> …` for embedding and `tmm nearest_neighbors -e <mining_spec.yaml> …` for mining. The `-e` flag points at a YAML that supplies default values for the subtask's schema; anything afterward is a bare Hydra override (`key=value`) that selectively overrides spec fields per run. (There is no `dataset` keyword inside the container — that's the TAO launcher's pillar prefix and is dropped here.) Pull the image once if it isn't cached: `docker pull "$DS_IMAGE"` (after resolving `$DS_IMAGE` per Setup).

Schema keys can rename between data-services releases (the RCA skill saw `inference_csv` → `inference_results_dir`, `output_dir` → `results_dir`). When in doubt, introspect the actual schema once per image: `docker run --rm "$DS_IMAGE" embedding image_embeddings --cfg=job` and `... tmm nearest_neighbors --cfg=job`.

---

## Inputs

1. **Target parquet** — the gap-analysis output, typically `mining_gaps.parquet` from `tao-route-visual-changenet-samples` (or `gaps.parquet` from `tao-analyze-gaps-visual-changenet` if routing was skipped). Required column: `filepath`. If `label` is also present, label-aware filtering during mining is available; otherwise the mining task silently no-ops the filter.
2. **Source pool** — a parquet of candidate images to mine against, with a `filepath` column. If the user only has a CSV, convert it to a parquet **with the same columns** before Step 2. For label-aware filtering, the pool must also carry a `label` column.
3. **Embedding spec file** — a YAML containing `model`, `model_path`, `batch_size`, and (only when `model_path` is a TAO `.pth`/`.ckpt`) `model_config_path`. Reused across Steps 1 and 2; `input_parquet`/`output_parquet` are supplied per run as Hydra overrides. The **same** spec MUST drive both embedding steps — embeddings from different encoders are not comparable, and mismatched encoders are the most common cause of "the mined images look unrelated" reports.
4. **Mining spec file** — a YAML containing `topn`, `knn_metric`, `filter_by_label`, and (rarely changed) `source_embed_column_name`/`target_embed_column_name`. `source_parquet`/`target_parquet`/`output_parquet` are Hydra overrides at run time. SigLIP and CLIP embeddings should use `knn_metric: cosine`. When `filter_by_label: true` but either embedding parquet lacks a `label` column, the container logs a warning and proceeds **without** filtering.

---

## Setup

Resolve the concrete `tao_toolkit.data_services` URI from `versions.yaml` once at the top of the run, then confirm Docker, the NVIDIA container toolkit, and a GPU are present before doing anything else. A GPU is required for both the encoder forward pass and the cuML/cuDF k-NN search; both steps fail without CUDA.

```bash
# Resolve tao_toolkit.data_services → concrete nvcr.io/... URI from versions.yaml
DS_IMAGE=$(python3 -c "import yaml,os; print(yaml.safe_load(open(os.environ['TAO_SKILL_BANK_PATH']+'/versions.yaml'))['images']['tao_toolkit']['data_services'])")
echo "DS_IMAGE=$DS_IMAGE"

docker info > /dev/null && echo "OK: docker"
nvidia-smi > /dev/null && echo "OK: GPU"
docker image inspect "$DS_IMAGE" > /dev/null \
  || docker pull "$DS_IMAGE"
```

Every host path the container reads or writes must be bind-mounted. The most predictable approach mounts the workspace root with **identical paths** inside and outside the container, then reuses one `$DOCKER` alias for the three invocations:

```bash
WORKSPACE=<absolute path that contains all parquets, outputs, and the source-pool images>
DOCKER="docker run --gpus all --rm --ipc=host -v $WORKSPACE:$WORKSPACE -w $WORKSPACE $DS_IMAGE"
```

Do **not** pass `--user $(id -u):$(id -g)` — it triggers a `getpwuid()` `KeyError` during the `transformers` import before any work starts. The container runs as root; chown outputs back to the host UID afterward.

Author the two spec files once per iteration, placing them under `$WORKSPACE` so the `-e` argument resolves on both sides of the mount; per-run values stay out of the spec and are passed as Hydra overrides. If the source pool is a CSV, convert it to parquet up front (preserving `filepath`, and `label` if present). The default `embedding_spec.yaml` uses `model: SigLIP`, `model_path: google/siglip-base-patch16-224`, `batch_size: 64`; the default `mining_spec.yaml` uses `topn: 5`, `knn_metric: cosine`, `filter_by_label: "false"` (quoted — the schema reads it as a string).

See `references/setup.md` for the full environment notes, `TAO_SKILL_BANK_PATH` handling, the path-mounting rationale, the `getpwuid` chown workaround, the CSV-to-parquet snippet, and the verbatim spec-file authoring blocks.

---

## Method

Three commands, in order. Each command's output parquet is the next command's input. Run them as plain Bash; the `$DOCKER` alias from Setup handles the container, GPU, and mounts. Every invocation follows the same shape: `-e <spec>` for the baked-in defaults, then a handful of Hydra overrides for the run-specific paths.

### Step 1 — Embed the target images

```bash
$DOCKER embedding image_embeddings \
    -e <embedding_spec.yaml> \
    input_parquet=<target_parquet> \
    output_parquet=<target_embeddings_parquet>
```

Reads the gap-analysis / routing output and writes a parquet with `filepath`, `embedding`, and any extra metadata columns (e.g. `label`, `siamese_score`, `weakness`) carried forward verbatim from the input. Print the output schema (`pd.read_parquet(...).columns`) to stdout so the script-check hook can confirm the embedding column exists.

If you need to override `model` / `model_path` / `batch_size` for one run without editing the spec, append them as Hydra overrides (e.g. `model_path=...`).

### Step 2 — Embed the source pool

```bash
$DOCKER embedding image_embeddings \
    -e <embedding_spec.yaml> \
    input_parquet=<source_pool_parquet> \
    output_parquet=<source_embeddings_parquet>
```

Same command shape as Step 1, applied to the source pool. Use the **identical** `embedding_spec.yaml` as Step 1, and do not override `model` / `model_path` / `batch_size` differently here — mismatched encoder configs across the two steps produce non-comparable embeddings.

### Step 3 — Mine nearest neighbours

```bash
$DOCKER tmm nearest_neighbors \
    -e <mining_spec.yaml> \
    source_parquet=<source_embeddings_parquet> \
    target_parquet=<target_embeddings_parquet> \
    output_parquet=<mined_parquet>
```

For each target embedding, finds the `topn` closest source embeddings under the chosen metric, deduplicates across targets, and writes a single-column (`filepath`) parquet of unique mined source paths. The container also drops a `mining_summary.txt` next to the output parquet with: query count, neighbour count, duplicates removed, and (when label filtering is on) kept-vs-dropped pair counts. Tweak `topn`, `knn_metric`, or `filter_by_label` via inline Hydra override when sweeping (e.g. `topn=10`) — no need to rewrite the spec.

When `filter_by_label=true` but one of the embedding parquets is missing the `label` column, the container logs a warning and proceeds without filtering. If the mined output looks larger than expected or contains cross-label pairs, scan the docker log for that warning before assuming the task did the right thing.

See `references/reference-invocation.md` for the minimal paste-and-edit end-to-end recipe (resolves `$DS_IMAGE`, writes both specs, runs all three steps, chowns outputs, and prints row counts) to run as a single streamed Bash block.

---

## Outputs and report

Write everything into a timestamped folder under the experiment / iteration directory. Get the real timestamp by running `date +%Y-%m-%d_%H%M%S` in Bash — do NOT hardcode or guess. If the user specifies a custom output path, use it directly but maintain the same internal layout. The packaging hook adds `mining_config/` and `claude_session.jsonl` automatically when `Mining_Report.md` is written.

The mined parquet is the artifact downstream training consumes. The two embedding parquets are intermediate but worth retaining — reusable across multiple mining runs against the same source pool, and the only place to look when a "looks unrelated" report needs encoder-level debugging.

See `references/outputs-and-reporting.md` for the full output-directory layout and the verbatim `Mining_Report.md` template (Verdict, Inputs, Encoder Consistency, Mining Run, Per-Label Breakdown, Output Sanity, Recommended Actions; keep it 600–1200 words).

---

## Common pitfalls

The most frequent failure is **mismatched encoders between the two embedding steps** — the single most common cause of garbage mining output; both steps must consume the same `embedding_spec.yaml`. Other recurring traps: passing `--user` (the `getpwuid` `KeyError`), skipping an embedding step, a missing `label` column silently no-oping `filter_by_label=true`, spec files outside `$WORKSPACE`, unresolved `???` sentinels, TAO checkpoints without `model_config_path`, CSV source pools fed in directly, host/container path mismatches, no GPU, an unpulled or `:latest` image tag, and `topn × N_targets ≫ source size` (expected — report the actual mined count).

See `references/troubleshooting.md` for the full pitfall list with the exact errors, causes, and fixes.

---

## Execution Order

1. Resolve `DS_IMAGE` from `versions.yaml` (`images.tao_toolkit.data_services`), then run `docker info`, `nvidia-smi`, and `docker image inspect "$DS_IMAGE"` (pulling if missing) once to confirm the environment. Abort with a clear message if any fail.
2. Run `date +%Y-%m-%d_%H%M%S` to get the timestamp; create `<output_dir>/mining_results/<timestamp>/`.
3. Write `embedding_spec.yaml` and `mining_spec.yaml` into the timestamped dir, filling in the encoder choice and mining knobs. Keep these under `$WORKSPACE` so the `-e` path resolves inside the container.
4. If the source pool is a CSV, convert to parquet first (preserve `filepath` and `label`).
5. Run Step 1 (embed targets) via `docker run … embedding image_embeddings -e embedding_spec.yaml input_parquet=… output_parquet=…`. Print the output parquet's row count and columns to stdout.
6. Run Step 2 (embed source pool) with the **identical** `embedding_spec.yaml` as Step 1. Print output row count and columns.
7. Run Step 3 (mine nearest neighbours) via `docker run … tmm nearest_neighbors -e mining_spec.yaml source_parquet=… target_parquet=… output_parquet=…`. Confirm `mining_summary.txt` was written next to `mined.parquet`.
8. Compute the per-label breakdown (Section 5) by joining the target embeddings parquet with the mined output on filepath, if both carry `label`.
9. Write `Mining_Report.md` last — writing it triggers the packaging hook, which copies session logs and skill config alongside.

tao-mine-aoi-imagesをインストール

スキルファイルをダウンロードし、.claude/skills/ ディレクトリに解凍してください。

ZIPをダウンロード

リポジトリをクローンし、スキルファイルをプロジェクトにコピーしてください。

git clone https://github.com/NVIDIA/skills/tree/main/skills/tao-mine-aoi-images # Copy SKILL.md to your .claude/skills/ directory

コピー コピー
クイックセットアップ: スキルフォルダを .claude/skills/ にコピーしてください。 Claude が自動的にそのスキルを検出して使用します。
リポジトリ NVIDIA/skills

関連スキル

microservices-patterns
更新された時間 2026年6月29日
jpa-patterns
更新された時間 2026年6月30日
fabric-lakehouse
更新された時間 2026年6月30日
prisma-expert
更新された時間 2026年6月29日
OR