옵션
집집 Skill 데이터베이스 관리 tao-mine-aoi-images

tao-mine-aoi-images

NVIDIA/skills NVIDIA/skills

VCN AOI 워크플로우에서 대상 및 소스 이미지 파케트를 임베드한 다음, 증강을 위해 가장 가까운 이웃 소스 이미지를 추출합니다.

...모든 것을 확장하십시오
26
업데이트 된 시간 2026년 9월 24일

DEFT 마이닝 및 임베딩 기술

귀하는 VCN AOI를 위한 DEFT 임베딩 후 마이닝 워크플로우의 운영자입니다. 귀하의 임무는 약한 타겟 이미지(갭 분석 또는 라우팅 출력)와 소스 풀로 구성된 파케트를 받아, 타겟과 유사하게 보이는 마이닝된 소스 이미지의 중복 제거된 파케트를 생성하여 다음 훈련 단계에 투입할 수 있도록 준비하는 것입니다.

이 워크플로는 고정적이고 결정론적입니다. 타겟을 임베딩하고, 소스 풀을 임베딩한 다음, 가장 가까운 이웃을 마이닝합니다. 각 단계의 출력 파케트는 다음 단계의 입력으로 사용됩니다. 반복적인 검색, 클러스터링 단계, 사람의 개입이 필요한 선택 과정은 없습니다. 깊이는 다단계 조사 과정을 통해서가 아니라, 올바른 인코더와 적절한 topn을 선택하는 데서 비롯됩니다.

전체 스킬은 versions.yaml에 선언된 tao_toolkit.data_services 이미지를 대상으로 하는 세 번의 직접적인 Docker 실행 호출을 감싸는 얇은 래퍼에 불과합니다(실행 시점에 해결됨 — ‘설정’ 참조). 컨테이너의 엔트리포인트는 -e [hydra 재정의...]를 받아, 임베딩을 위해 image_embeddings -e …을, 마이닝을 위해 tmm nearest_neighbors -e …을 전달합니다. -e 플래그는 하위 작업의 스키마에 대한 기본값을 제공하는 YAML 파일을 가리키며, 그 뒤에 오는 모든 항목은 실행마다 spec 필드를 선택적으로 재정의하는 순수한 Hydra 재정의(key=value)입니다. (컨테이너 내부에는 dataset 키워드가 없습니다. 이는 TAO 런처의 pillar 접두사이며, 여기서는 생략됩니다.) 이미지가 캐시되어 있지 않으면 한 번 이미지를 가져옵니다: docker pull "$DS_IMAGE" (Setup에 따라 $DS_IMAGE를 해결한 후).

스키마 키는 데이터 서비스 릴리스 간에 이름이 변경될 수 있습니다(RCA 스킬의 경우 inference_csv → inference_results_dir, output_dir → results_dir로 변경됨). 확실하지 않은 경우, 이미지당 한 번씩 실제 스키마를 확인하세요: docker run --rm "$DS_IMAGE" embedding image_embeddings --cfg=job 및 ... tmm nearest_neighbors --cfg=job.

입력

  1. 대상 파케트 — 갭 분석 결과, 일반적으로 tao-route-visual-changenet-samples 의 mining_gaps.parquet (또는 라우팅이 건너뛴 경우 tao-analyze-gaps-visual-changenet 의 gaps.parquet ). 필수 열: filepath. label 열도 함께 있는 경우, 마이닝 과정에서 레이블 기반 필터링을 사용할 수 있습니다. 그렇지 않으면 마이닝 작업에서 필터 처리를 수행하지 않고 무시합니다.
  2. 소스 풀 — 마이닝에 사용할 후보 이미지의 파케트 파일로, filepath 열이 포함되어 있어야 합니다. CSV 파일만 있는 경우, 2단계 전에 동일한 열을 가진 파케트 파일로 변환하십시오. 라벨 기반 필터링을 사용하려면 풀에 label 열도 포함되어야 합니다.
  3. 임베딩 사양 파일 — model, model_path, batch_size 및 ( model_path가 TAO .pth/.ckpt인 경우에만) model_config_path를 포함하는 YAML 파일입니다. 1단계와 2단계에서 재사용되며,input_parquet/output_parquet는 Hydra가 재정의하므로 실행마다 제공됩니다. 동일한 사양 파일이 두 임베딩 단계를 모두 제어해야 합니다. 서로 다른 인코더에서 생성된 임베딩은 비교할 수 없으며, 인코더가 일치하지 않는 경우 “채굴된 이미지가 관련성이 없어 보인다”는 보고가 발생하는 가장 흔한 원인입니다.
  4. 마이닝 사양 파일 — topn, knn_metric, filter_by_label 및 (거의 변경되지 않는)source_embed_column_name/target_embed_column_name을 포함하는 YAML 파일입니다.source_parquet/target_parquet/output_parquet는 실행 시점에 Hydra에 의해 재정의됩니다. SigLIP 및 CLIP 임베딩은 knn_metric: cosine을 사용해야 합니다. filter_by_label: true이지만 임베딩 파케트 중 하나에 레이블 열이 없는 경우, 컨테이너는 경고를 기록하고 필터링 없이 진행합니다.

설정

실행 시작 시 versions.yaml 에서 구체적인 tao_toolkit.data_services URI를 한 번만 해결한 다음, 다른 작업을 수행하기 전에 Docker, NVIDIA 컨테이너 툴킷 및 GPU가 있는지 확인하십시오. 인코더 포워드 패스와 cuML/cuDF k-NN 검색 모두에 GPU가 필요합니다. CUDA가 없으면 두 단계 모두 실패합니다.

# versions.yaml에서 tao_toolkit.data_services → 구체적인 nvcr.io/... URI를 해결합니다.
DS_IMAGE=$(python3 -c "import yaml,os; print(yaml.safe_load(open(os.environ['TAO_SKILL_BANK_PATH']+'/versions.yaml'))['images']['tao_toolkit']['data_services'])")
echo "DS_IMAGE=$DS_IMAGE"

docker info > /dev/null && echo "OK: docker"
nvidia-smi > /dev/null && echo "OK: GPU"
docker image inspect "$DS_IMAGE" > /dev/null \
  || docker pull "$DS_IMAGE"

컨테이너가 읽거나 쓰는 모든 호스트 경로는 바인드 마운트되어야 합니다. 가장 예측 가능한 방법은 컨테이너 내부와 외부의 경로가 동일한 상태로 작업 공간 루트를 마운트한 다음, 세 가지 호출에 대해 하나의 $DOCKER 별칭을 재사용하는 것입니다:

WORKSPACE=
DOCKER="docker run --gpus all --rm --ipc=host -v $WORKSPACE:$WORKSPACE -w $WORKSPACE $DS_IMAGE"

--user $(id -u):$(id -g) 옵션을 전달하지 마십시오. 작업이 시작되기도 전에 transformers 임포트 과정에서 getpwuid() KeyError가 발생합니다. 컨테이너는 root 권한으로 실행되며, 이후 chown 명령어가 호스트 UID로 다시 설정합니다.

반복마다 두 개의 spec 파일을 한 번씩 작성하고, $WORKSPACE 아래에 배치하여 -e 인수가 마운트 양쪽에서 모두 해결되도록 하십시오. 실행별 값은 spec에 포함하지 말고 Hydra 오버라이드로 전달하십시오. 소스 풀이 CSV인 경우, 미리 Parquet 형식으로 변환하십시오( 파일 경로와 라벨이 있는 경우 라벨을 유지). 기본 embedding_spec.yaml은 model: SigLIP, model_path: google/siglip-base-patch16-224, batch_size: 64를 사용합니다. 기본 mining_spec.yaml은 topn: 5, knn_metric: cosine, filter_by_label: "false" (따옴표 포함 — 스키마에서 이를 문자열로 인식함)를 사용합니다.

전체 환경 관련 참고 사항, TAO_SKILL_BANK_PATH 처리 방법, 경로 마운팅의 근거, getpwuid chown 우회 방법, CSV-to-parquet 스니펫, 그리고 사양 파일 작성 블록의 원문 내용은 references/setup.md를 참조하십시오.

방법

순서대로 세 가지 명령어를 실행합니다. 각 명령어의 출력 파케트(Parquet) 파일이 다음 명령어의 입력으로 사용됩니다. 일반 Bash로 실행하면 되며, Setup에서 정의된 $DOCKER 별칭이 컨테이너, GPU 및 마운트를 처리합니다. 모든 호출은 동일한 형식을 따릅니다: 내장된 기본값을 사용하려면 -e 을 입력하고, 실행별 경로를 지정하려면 몇 가지 Hydra 오버라이드를 추가합니다.

1단계 — 대상 이미지 삽입

$DOCKER embedding image_embeddings \
    -e  \
    input_parquet= \
    output_parquet=

갭 분석/라우팅 출력을 읽어들이고, 파일 경로, 임베딩 및 입력에서 그대로 가져온 추가 메타데이터 열(예: label, siamese_score, weakness)을 포함한 파케트 파일을 작성합니다. 스크립트 확인 후크가 임베딩 열이 존재하는지 확인할 수 있도록 출력 스키마(pd.read_parquet(...).columns)를 표준 출력으로 출력합니다.

사양을 편집하지 않고 한 번의 실행에 대해 model / model_path / batch_size를 재정의해야 하는 경우, 이를 Hydra 재정의(예: model_path=...)로 추가하십시오.

2단계 — 소스 풀 임베딩

$DOCKER embedding image_embeddings \
    -e  \
    input_parquet= \
    output_parquet=

1단계와 동일한 명령어 형식을 사용하여 소스 풀에 적용합니다. 1단계와 동일한 embedding_spec.yaml을 사용해야 하며, 여기서 model / model_path / batch_size를 다르게 재정의해서는 안 됩니다. 두 단계에서 인코더 구성이 일치하지 않으면 임베딩을 비교할 수 없게 됩니다.

3단계 — 가장 가까운 이웃 추출

$DOCKER tmm nearest_neighbors \
    -e  \
    source_parquet= \
    target_parquet= \
    output_parquet=

각 대상 임베딩에 대해, 선택된 메트릭 기준 상위 n개의 가장 가까운 소스 임베딩을 찾고, 대상 간 중복을 제거한 후, 추출된 고유한 소스 경로를 단일 열(파일 경로) 파케트 파일로 작성합니다. 또한 이 컨테이너는 출력 파케트 파일 옆에 mining_summary.txt 파일을 생성하며, 여기에는 쿼리 수, 이웃 수, 제거된 중복 수, 그리고 (라벨 필터링이 활성화된 경우) 유지된 쌍과 제거된 쌍의 개수가 포함됩니다. 스윕 시 인라인 Hydra 오버라이드를 통해 topn, knn_metric 또는 filter_by_label을 조정할 수 있습니다(예: topn=10). 사양을 다시 작성할 필요는 없습니다.

filter_by_label=true인 경우, 임베딩 Parquet 파일 중 하나에 레이블 열이 누락되어 있으면 컨테이너는 경고를 기록하고 필터링 없이 진행합니다. 마이닝된 출력 결과가 예상보다 크거나 레이블이 다른 쌍이 포함된 경우, 작업이 정상적으로 수행되었다고 단정하기 전에 Docker 로그에서 해당 경고를 확인하십시오.

단일 스트리밍 Bash 블록으로 실행할 수 있는 최소한의 ‘복사-붙여넣기-편집’ 방식의 엔드투엔드 레시피( `$DS_IMAGE`를 해결하고, 두 사양 파일을 모두 작성하며, 세 단계를 모두 실행하고, 출력 파일의 소유권을 변경하며, 행 수를 출력함)에 대해서는 references/reference-invocation.md를 참조하십시오.

출력 및 보고서

모든 결과를 실험 / 반복 디렉터리 아래의 타임스탬프가 포함된 폴더에 기록하십시오. Bash에서 date +%Y-%m-%d_%H%M%S 명령을 실행하여 실제 타임스탬프를 가져오십시오 — 절대 하드코딩하거나 추측하지 마십시오. 사용자가 사용자 정의 출력 경로를 지정한 경우, 이를 직접 사용하되 내부 레이아웃은 동일하게 유지하십시오. 패키징 훅은 Mining_Report.md가 작성되면 mining_config/ 및 claude_session.jsonl을 자동으로 추가합니다.

마이닝된 파케트는 다운스트림 훈련에서 사용하는 아티팩트입니다. 두 개의 임베딩 파케트는 중간 결과물이지만 보관할 가치가 있습니다. 동일한 소스 풀에 대한 여러 마이닝 실행에서 재사용할 수 있을 뿐만 아니라, "무관해 보이는" 보고서에 대해 인코더 수준 디버깅이 필요할 때 확인해야 할 유일한 위치이기 때문입니다.

전체 출력 디렉터리 구조와 Mining_Report.md 템플릿 원문(Verdict, Inputs, Encoder Consistency, Mining Run, Per-Label Breakdown, Output Sanity, Recommended Actions; 600~1200단어 범위 유지)은 references/outputs-and-reporting.md를 참조하십시오 .

흔히 발생하는 문제점

가장 빈번한 오류는 두 임베딩 단계 간 인코더가 일치하지 않는 경우로, 이는 마이닝 출력 결과가 비정상적으로 나오는 가장 흔한 원인입니다. 두 단계 모두 동일한 embedding_spec.yaml을 사용해야 합니다. 그 밖의 반복적으로 발생하는 함정으로는 --user 전달( getpwuid KeyError), 임베딩 단계 생략, 레이블 열 누락으로 인해 filter_by_label=true가 아무런 동작도 하지 않는 경우, $WORKSPACE 외부에 있는 spec 파일, 해결되지 않은 ??? 센티넬 미해결, model_config_path가 없는 TAO 체크포인트, CSV 소스 풀을 직접 입력, 호스트/컨테이너 경로 불일치, GPU 없음, 이미지 태그를 가져오지 않았거나 :latest 태그 사용, topn × N_targets ≫ 소스 크기 (예상된 현상 — 실제 마이닝된 개수를 보고하십시오).

정확한 오류, 원인 및 해결 방법이 포함된 전체 문제점 목록은 references/troubleshooting.md를 참조하십시오.

실행 순서

  1. versions.yaml (images.tao_toolkit.data_services)에서 DS_IMAGE를 확인한 후, docker info, nvidia-smi, docker image inspect "$DS_IMAGE" (이미지가 없는 경우 다운로드)를 한 번씩 실행하여 환경을 확인합니다. 실패하는 항목이 있으면 명확한 메시지와 함께 중단하십시오.
  2. ` date +%Y-%m-%d_%H%M%S `를 실행하여 타임스탬프를 얻고, ` /mining_results/` 및 `/` 디렉터리를 생성합니다.
  3. 임베딩 사양(embedding_spec.yaml )과 마이닝 사양(mining_spec.yaml ) 파일을 타임스탬프가 포함된 디렉터리에 작성하고, 인코더 선택 사항과 마이닝 조정 매개변수를 입력합니다. -e 경로가 컨테이너 내에서 올바르게 인식되도록 이 파일들을 $WORKSPACE 아래에 보관하십시오.
  4. 소스 데이터가 CSV인 경우, 먼저 Parquet 형식으로 변환합니다( 파일 경로 와 레이블은 유지).
  5. docker run … embedding image_embeddings -e embedding_spec.yaml input_parquet=… output_parquet=… 명령어를 통해 1단계(타깃 임베딩)를 실행합니다. 출력된 파케트의 행 수와 열 수를 stdout에 출력합니다.
  6. 1단계와 동일한 embedding_spec.yaml 을 사용하여 2단계(소스 풀 임베딩)를 실행합니다. 출력 Parquet의 행 수와 열 수를 출력합니다.
  7. docker run … tmm nearest_neighbors -e mining_spec.yaml source_parquet=… target_parquet=… output_parquet=… 명령어를 통해 3단계(가장 가까운 이웃 마이닝)를 실행합니다 . mining_summary.txt 파일이 mined.parquet 파일 옆에 생성되었는지 확인합니다.
  8. 두 파일 모두 레이블을 포함하고 있다면, target embeddings parquet 파일을 mined 출력 파일과 파일 경로를 기준으로 조인하여 레이블별 분포(5절)를 계산합니다.
  9. 마지막으로 Mining_Report.md 파일을 작성합니다. 이 파일을 작성하면 패키징 훅이 트리거되어 세션 로그와 스킬 구성이 함께 복사됩니다.
GitHub에서 보기
---
name: tao-mine-aoi-images
description: Embeds target and source image parquets, then mines nearest-neighbour source images for augmentation in VCN AOI workflows.
license: Apache-2.0
---

# DEFT Mining and Embedding Skill

You are the operator of the DEFT embed-then-mine workflow for VCN AOI. Your job is to take a parquet of weak target images (the gap-analysis or routing output) and a source pool, then produce a deduplicated parquet of mined source images that look similar to the targets — ready to feed into the next training round.

The workflow is fixed and deterministic: **embed the targets, embed the source pool, then mine nearest neighbours.** Each step's output parquet is the next step's input. There is no iterative search, no clustering pass, no human-in-the-loop selection — depth comes from picking the right encoder and the right `topn`, not from a multi-phase investigation.

The whole skill is a thin wrapper around three direct `docker run` invocations against the `tao_toolkit.data_services` image declared in `versions.yaml` (resolved at runtime — see Setup). The container's entrypoint takes `<category> <action> -e <spec.yaml> [hydra overrides...]` — pass `embedding image_embeddings -e <embedding_spec.yaml> …` for embedding and `tmm nearest_neighbors -e <mining_spec.yaml> …` for mining. The `-e` flag points at a YAML that supplies default values for the subtask's schema; anything afterward is a bare Hydra override (`key=value`) that selectively overrides spec fields per run. (There is no `dataset` keyword inside the container — that's the TAO launcher's pillar prefix and is dropped here.) Pull the image once if it isn't cached: `docker pull "$DS_IMAGE"` (after resolving `$DS_IMAGE` per Setup).

Schema keys can rename between data-services releases (the RCA skill saw `inference_csv` → `inference_results_dir`, `output_dir` → `results_dir`). When in doubt, introspect the actual schema once per image: `docker run --rm "$DS_IMAGE" embedding image_embeddings --cfg=job` and `... tmm nearest_neighbors --cfg=job`.

---

## Inputs

1. **Target parquet** — the gap-analysis output, typically `mining_gaps.parquet` from `tao-route-visual-changenet-samples` (or `gaps.parquet` from `tao-analyze-gaps-visual-changenet` if routing was skipped). Required column: `filepath`. If `label` is also present, label-aware filtering during mining is available; otherwise the mining task silently no-ops the filter.
2. **Source pool** — a parquet of candidate images to mine against, with a `filepath` column. If the user only has a CSV, convert it to a parquet **with the same columns** before Step 2. For label-aware filtering, the pool must also carry a `label` column.
3. **Embedding spec file** — a YAML containing `model`, `model_path`, `batch_size`, and (only when `model_path` is a TAO `.pth`/`.ckpt`) `model_config_path`. Reused across Steps 1 and 2; `input_parquet`/`output_parquet` are supplied per run as Hydra overrides. The **same** spec MUST drive both embedding steps — embeddings from different encoders are not comparable, and mismatched encoders are the most common cause of "the mined images look unrelated" reports.
4. **Mining spec file** — a YAML containing `topn`, `knn_metric`, `filter_by_label`, and (rarely changed) `source_embed_column_name`/`target_embed_column_name`. `source_parquet`/`target_parquet`/`output_parquet` are Hydra overrides at run time. SigLIP and CLIP embeddings should use `knn_metric: cosine`. When `filter_by_label: true` but either embedding parquet lacks a `label` column, the container logs a warning and proceeds **without** filtering.

---

## Setup

Resolve the concrete `tao_toolkit.data_services` URI from `versions.yaml` once at the top of the run, then confirm Docker, the NVIDIA container toolkit, and a GPU are present before doing anything else. A GPU is required for both the encoder forward pass and the cuML/cuDF k-NN search; both steps fail without CUDA.

```bash
# Resolve tao_toolkit.data_services → concrete nvcr.io/... URI from versions.yaml
DS_IMAGE=$(python3 -c "import yaml,os; print(yaml.safe_load(open(os.environ['TAO_SKILL_BANK_PATH']+'/versions.yaml'))['images']['tao_toolkit']['data_services'])")
echo "DS_IMAGE=$DS_IMAGE"

docker info > /dev/null && echo "OK: docker"
nvidia-smi > /dev/null && echo "OK: GPU"
docker image inspect "$DS_IMAGE" > /dev/null \
  || docker pull "$DS_IMAGE"
```

Every host path the container reads or writes must be bind-mounted. The most predictable approach mounts the workspace root with **identical paths** inside and outside the container, then reuses one `$DOCKER` alias for the three invocations:

```bash
WORKSPACE=<absolute path that contains all parquets, outputs, and the source-pool images>
DOCKER="docker run --gpus all --rm --ipc=host -v $WORKSPACE:$WORKSPACE -w $WORKSPACE $DS_IMAGE"
```

Do **not** pass `--user $(id -u):$(id -g)` — it triggers a `getpwuid()` `KeyError` during the `transformers` import before any work starts. The container runs as root; chown outputs back to the host UID afterward.

Author the two spec files once per iteration, placing them under `$WORKSPACE` so the `-e` argument resolves on both sides of the mount; per-run values stay out of the spec and are passed as Hydra overrides. If the source pool is a CSV, convert it to parquet up front (preserving `filepath`, and `label` if present). The default `embedding_spec.yaml` uses `model: SigLIP`, `model_path: google/siglip-base-patch16-224`, `batch_size: 64`; the default `mining_spec.yaml` uses `topn: 5`, `knn_metric: cosine`, `filter_by_label: "false"` (quoted — the schema reads it as a string).

See `references/setup.md` for the full environment notes, `TAO_SKILL_BANK_PATH` handling, the path-mounting rationale, the `getpwuid` chown workaround, the CSV-to-parquet snippet, and the verbatim spec-file authoring blocks.

---

## Method

Three commands, in order. Each command's output parquet is the next command's input. Run them as plain Bash; the `$DOCKER` alias from Setup handles the container, GPU, and mounts. Every invocation follows the same shape: `-e <spec>` for the baked-in defaults, then a handful of Hydra overrides for the run-specific paths.

### Step 1 — Embed the target images

```bash
$DOCKER embedding image_embeddings \
    -e <embedding_spec.yaml> \
    input_parquet=<target_parquet> \
    output_parquet=<target_embeddings_parquet>
```

Reads the gap-analysis / routing output and writes a parquet with `filepath`, `embedding`, and any extra metadata columns (e.g. `label`, `siamese_score`, `weakness`) carried forward verbatim from the input. Print the output schema (`pd.read_parquet(...).columns`) to stdout so the script-check hook can confirm the embedding column exists.

If you need to override `model` / `model_path` / `batch_size` for one run without editing the spec, append them as Hydra overrides (e.g. `model_path=...`).

### Step 2 — Embed the source pool

```bash
$DOCKER embedding image_embeddings \
    -e <embedding_spec.yaml> \
    input_parquet=<source_pool_parquet> \
    output_parquet=<source_embeddings_parquet>
```

Same command shape as Step 1, applied to the source pool. Use the **identical** `embedding_spec.yaml` as Step 1, and do not override `model` / `model_path` / `batch_size` differently here — mismatched encoder configs across the two steps produce non-comparable embeddings.

### Step 3 — Mine nearest neighbours

```bash
$DOCKER tmm nearest_neighbors \
    -e <mining_spec.yaml> \
    source_parquet=<source_embeddings_parquet> \
    target_parquet=<target_embeddings_parquet> \
    output_parquet=<mined_parquet>
```

For each target embedding, finds the `topn` closest source embeddings under the chosen metric, deduplicates across targets, and writes a single-column (`filepath`) parquet of unique mined source paths. The container also drops a `mining_summary.txt` next to the output parquet with: query count, neighbour count, duplicates removed, and (when label filtering is on) kept-vs-dropped pair counts. Tweak `topn`, `knn_metric`, or `filter_by_label` via inline Hydra override when sweeping (e.g. `topn=10`) — no need to rewrite the spec.

When `filter_by_label=true` but one of the embedding parquets is missing the `label` column, the container logs a warning and proceeds without filtering. If the mined output looks larger than expected or contains cross-label pairs, scan the docker log for that warning before assuming the task did the right thing.

See `references/reference-invocation.md` for the minimal paste-and-edit end-to-end recipe (resolves `$DS_IMAGE`, writes both specs, runs all three steps, chowns outputs, and prints row counts) to run as a single streamed Bash block.

---

## Outputs and report

Write everything into a timestamped folder under the experiment / iteration directory. Get the real timestamp by running `date +%Y-%m-%d_%H%M%S` in Bash — do NOT hardcode or guess. If the user specifies a custom output path, use it directly but maintain the same internal layout. The packaging hook adds `mining_config/` and `claude_session.jsonl` automatically when `Mining_Report.md` is written.

The mined parquet is the artifact downstream training consumes. The two embedding parquets are intermediate but worth retaining — reusable across multiple mining runs against the same source pool, and the only place to look when a "looks unrelated" report needs encoder-level debugging.

See `references/outputs-and-reporting.md` for the full output-directory layout and the verbatim `Mining_Report.md` template (Verdict, Inputs, Encoder Consistency, Mining Run, Per-Label Breakdown, Output Sanity, Recommended Actions; keep it 600–1200 words).

---

## Common pitfalls

The most frequent failure is **mismatched encoders between the two embedding steps** — the single most common cause of garbage mining output; both steps must consume the same `embedding_spec.yaml`. Other recurring traps: passing `--user` (the `getpwuid` `KeyError`), skipping an embedding step, a missing `label` column silently no-oping `filter_by_label=true`, spec files outside `$WORKSPACE`, unresolved `???` sentinels, TAO checkpoints without `model_config_path`, CSV source pools fed in directly, host/container path mismatches, no GPU, an unpulled or `:latest` image tag, and `topn × N_targets ≫ source size` (expected — report the actual mined count).

See `references/troubleshooting.md` for the full pitfall list with the exact errors, causes, and fixes.

---

## Execution Order

1. Resolve `DS_IMAGE` from `versions.yaml` (`images.tao_toolkit.data_services`), then run `docker info`, `nvidia-smi`, and `docker image inspect "$DS_IMAGE"` (pulling if missing) once to confirm the environment. Abort with a clear message if any fail.
2. Run `date +%Y-%m-%d_%H%M%S` to get the timestamp; create `<output_dir>/mining_results/<timestamp>/`.
3. Write `embedding_spec.yaml` and `mining_spec.yaml` into the timestamped dir, filling in the encoder choice and mining knobs. Keep these under `$WORKSPACE` so the `-e` path resolves inside the container.
4. If the source pool is a CSV, convert to parquet first (preserve `filepath` and `label`).
5. Run Step 1 (embed targets) via `docker run … embedding image_embeddings -e embedding_spec.yaml input_parquet=… output_parquet=…`. Print the output parquet's row count and columns to stdout.
6. Run Step 2 (embed source pool) with the **identical** `embedding_spec.yaml` as Step 1. Print output row count and columns.
7. Run Step 3 (mine nearest neighbours) via `docker run … tmm nearest_neighbors -e mining_spec.yaml source_parquet=… target_parquet=… output_parquet=…`. Confirm `mining_summary.txt` was written next to `mined.parquet`.
8. Compute the per-label breakdown (Section 5) by joining the target embeddings parquet with the mined output on filepath, if both carry `label`.
9. Write `Mining_Report.md` last — writing it triggers the packaging hook, which copies session logs and skill config alongside.

tao-mine-aoi-images 설치

스킬 파일을 다운로드하여 .claude/skills/ 디렉터리에 압축을 풀어주세요.

ZIP 다운로드

저장소를 클론하고 스킬 파일을 프로젝트에 복사하세요.

git clone https://github.com/NVIDIA/skills/tree/main/skills/tao-mine-aoi-images # Copy SKILL.md to your .claude/skills/ directory

복사 복사
빠른 설정: 스킬 폴더를 .claude/skills/로 복사하세요. Claude가 해당 스킬을 자동으로 감지하여 사용할 것입니다.
저장소 NVIDIA/skills

관련 스킬

microservices-patterns
업데이트 된 시간 2026년 6월 29일
jpa-patterns
업데이트 된 시간 2026년 6월 30일
fabric-lakehouse
업데이트 된 시간 2026년 6월 30일
prisma-expert
업데이트 된 시간 2026년 6월 29일
OR