オプション

tao-train-grounding-dino

NVIDIA/skills NVIDIA/skills

固定されたクラス語彙を持たずに、テキストプロンプトで記述されたオブジェクトを検出するGrounding DINOモデルに対し、トレーニング、評価、エクスポート、量子化、および推論実行を行う。

...すべて拡張します
0
更新された時間 2026年9月25日

DINOのグラウンディング

オープンセット物体検出のためのGrounding DINO。DINOスタイルの検出とBERTテキストエンコーダーを組み合わせ、言語主導型の検出を実現します。固定されたクラス語彙に依存せず、テキストプロンプトで記述された物体を検出します。

Grounding DINOの完全な重みを使用する場合は `train.pretrained_model_path` を、バックボーンのみを使用する場合は `model.pretrained_backbone_path` を設定してください。

TAO DeployのTensorRTアクション(gen_trt_engine、TensorRTevaluate、およびTensorRTinference)については、まずreferences/tao-deploy-grounding-dino.mdを参照してください。 デプロイ仕様テンプレートは、このスキルのreferences/フォルダ内に、spec_template_deploy_*.yamlというプレフィックスで格納されています。

データクラスのスキーマ

生成された TAO Core スキーマは、schemas/.schema.json にパッケージ化されており、利用可能なアクションはschemas/manifest.jsonにリストされています。また、生成された各スキーマは、スキーマのトップレベルにあるdefaultフィールドから、references/spec_template_.yamlも生成します。 AutoMLの有効化は、references/skill_info.yaml内のmodelレイヤーで、automl_enabledを介して宣言されます。実行可能なAutoMLでは、schemas/train.schema.jsonおよびreferences/spec_template_train.yamlが存在し、パースされる必要があります。automl_default_parameters、automl_disabled_parameters、デフォルト値、最小/最大境界、列挙型、オプションの重み、数学的条件、依存関係、および一般的なパラメータについては、パッケージ化されたtrainスキーマを使用してください。実行時に~/tao-coreが存在することを想定しないでください。メンテナンス担当者は、スキルバンクをパッケージ化する前にスキーマ/テンプレートを再生成します。

トレーニング・アクション・ポリシー

このモデルは、モデル層で AutoML が有効になっています。トレーニング段階のリクエストを処理する前に、`references/skill_info.yaml` を読み込み、明示的な `automl_policy` 値またはユーザーのワークフローリクエストのいずれかから、実行時のオーバーライドを解決してください。デフォルトでは`automl_policy: on`を使用し、新しい起動プロンプトでは`on`および`off` のみを公開してください。 「AutoMLをオフにする」、「AutoMLを無効にする」、「HPOなし」、または「プレーンなトレーニング」といったフレーズは、その実行のみにおいてautoml_policy: offとして扱ってください。automl_policy: on、automl_enabled: true であり、かつschemas/train.schema.jsonとreferences/spec_template_train.yamlの両方がパッケージ化されている場合、train アクションはデフォルトで、このモデルのskill_dir を使用してtao-skill-bank:tao-run-automl経由でルーティングされます。 データセット、仕様、出力ディレクトリ、GPU/プラットフォーム設定、親チェックポイント、およびautoml_policy に関するワークフロー/アプリケーションのオーバーライドは保持します。automl_policy: offの場合、またはパッケージ化された train スキーマ/テンプレートが存在しない場合にのみ、直接モデルトレーニングを使用します。スキーマが存在しない場合、スキーマが生成されるまで、このモデルでは AutoML が有効であるが実行不可能であることを報告します。

evaluate、inference、export、deploy フローなどのトレーニング以外のアクションは、このモデルスキル内に残ります。実行ごとのautoml_policyのオーバーライドによって、モデルのメタデータが変更されることはありません。

トレーニングの要件

  • データセットの種類:object_detection
  • フォーマット:odvg、coco、raw
  • モニタリング指標:val_mAP50

アクションごとのデータセット要件

アクション 仕様のキー ソース ファイル リスト?
評価 データセット.test_data_sources eval_dataset image_dir: images.tar.gz, json_file: annotations.json いいえ
推論 dataset.infer_data_sources.image_dir 推論データセット images.tar.gz はい
推論 dataset.infer_data_sources.captions ワークフローのプロンプト プロンプト一覧 はい
量子化 データセット.train_data_sources train_datasets image_dir: images.tar.gz, json_file: annotations_odvg.jsonl, label_map: annotations_odvg_labelmap.json はい
quantize dataset.val_data_sources eval_dataset image_dir: images.tar.gz, json_file: annotations.json いいえ
quantize dataset.quant_calibration_data_sources キャリブレーション/評価データセット image_dir: images.tar.gz, json_file: annotations.json なしクオンタイズdataset.quant_calibration_data_sourcesキャリブ
train dataset.train_data_sources train_datasets image_dir: images.tar.gz, json_file: annotations_odvg.jsonl, label_map: annotations_odvg_labelmap.json はい
train dataset.val_data_sources eval_dataset image_dir: images.tar.gz, json_file: annotations.json いいえ

ランナーは `images.tar.gz` のようなイメージアーカイブをソースとして使用できますが、ローカルの Docker TAO CLI 仕様では、`image_dir` を解凍済みのイメージディレクトリに指定する必要があります。 スキルメタデータは、これらのアーカイブベースのイメージソースを `runtime: extracted_folder` とマークするため、新しいランナーは TAO を起動する前に アーカイブを解凍することができます。

一般的な仕様の上書き

データソースのオーバーライドはすべてのアクションで必須です。エージェントは、上記の「アクションごとのデータセット要件」テーブルに基づいてデータソースのパスを構築し、それらを `spec_overrides` に含めなければなりません。

S3_TRAIN = "s3://bucket/data/train"
S3_EVAL = "s3://bucket/data/eval"

train(必須のデータソース):

{
    "train.num_epochs": 10,
    "train.checkpoint_interval": 10,
    "train.validation_interval": 10,
    "train.num_gpus": 1,
    "dataset.train_data_sources": [{"image_dir": f"{S3_TRAIN}/images.tar.gz", "json_file": f"{S3_TRAIN}/annotations_odvg.jsonl", "label_map": f"{S3_TRAIN}/annotations_odvg_labelmap.json"}],
    "dataset.val_data_sources": {"image_dir": f"{S3_EVAL}/images.tar.gz", "json_file": f"{S3_EVAL}/annotations.json"},
}

deploy/gen_trt_engine(references/tao-deploy-grounding-dino.md を使用):

{
    "gen_trt_engine.onnx_file": "",
    "gen_trt_engine.trt_engine": "",
    "gen_trt_engine.tensorrt.data_type": "FP16",
}

推論(必須のデータソース):

{
    "inference.checkpoint": "<選択したトレーニング/AutoMLチェックポイント>",
    "dataset.infer_data_sources.image_dir": [f"{S3_EVAL}/images.tar.gz"],
    "dataset.infer_data_sources.captions": [
        "消火器",
        "コーン",
        "カート",
        "フォークリフト"
    ],
}

評価(必須のデータソース):

{
    "evaluate.checkpoint": "<選択したトレーニング/AutoMLチェックポイント>",
    "dataset.test_data_sources": {"image_dir": f"{S3_EVAL}/images.tar.gz", "json_file": f"{S3_EVAL}/annotations.json"},
}

量子化(必須のデータソース):

{
    "quantize.model_path": "",
    "dataset.train_data_sources": [{"image_dir": f"{S3_TRAIN}/images.tar.gz", "json_file": f"{S3_TRAIN}/annotations_odvg.jsonl", "label_map": f"{S3_TRAIN}/annotations_odvg_labelmap.json"}],
    "dataset.val_data_sources": {"image_dir": f"{S3_EVAL}/images.tar.gz", "json_file": f"{S3_EVAL}/annotations.json"},
    "dataset.quant_calibration_data_sources": {"image_dir": f"{S3_EVAL}/images.tar.gz", "json_file": f"{S3_EVAL}/annotations.json"},
}

評価用データセット

オプション。トレーニングではODVG形式を使用できますが、mAPの評価にはCOCO形式のアノテーションが使用されます。

重要なパラメータ

  • model.backbone: デフォルトは swin_tiny_224_1k です。resnet_50 やその他の Swin バリアントもサポートしています。Swin は一般的にグラウンディングタスクでより優れた性能を発揮します。
  • model.text_encoder_type: テキストエンコーディング用のBERTモデル。デフォルトはbert-base-uncased。max_text_lenのデフォルト値は256です。
  • model.max_text_len: これをデータセットのラベル/トークン 位置マップと一致させてください。対応する ラベルマップが同じ長さで再生成されない限り、スモークテストのためにこれを短縮しないでください。そうしないと、トークン確率と位置 マップの間で行列の形状が一致せず、検証が 失敗する可能性があります。
  • train.optim.lr: 学習率。デフォルトは 2e-4。lr_backbone は 2e-5。fp16/fp32 に加え、bf16 精度もサポートしています。
  • dataset.max_labels: トレーニング中の画像あたりの最大ラベル数。デフォルトは 50。アノテーションが密なデータセットの場合はこの値を増やしてください。
  • model.num_queries: オブジェクトクエリ数。オープンボキャブラリの性質上、デフォルトは 900(DINO の 300 より多い)です。
  • model.num_queries / model.num_select: バッチ内のマッチしたODVGターゲットの数に対して、 num_queriesを十分に高い値に設定してください。 20 のような非常に小さい smoke 値は、 高密度な画像でのハンガリアンターゲットインデックス作成中に失敗する可能性があります。データセットが 画像あたりのオブジェクト数が少ないことが分かっている場合を除き、Grounding DINO の smoke 実行を最小限に抑えるには、 少なくとも 100 を使用してください。
  • train.optim.lr_steps: MultiStep LRのスケジュール。デフォルト [10]。

マルチGPU / マルチノード

起動方法:Lightning 管理 (単一のPythonプロセス、Lightning がワーカーを起動)。

仕様キー 説明 デフォルト
train.num_gpus GPUの数 1
train.gpu_ids GPUデバイスのインデックス [0]
train.num_nodes ノード数 1
train.distributed_strategy ddpまたはfsdp ddp

DINOと同様のDDP/FSDPの動作。マルチノード化には、オーケストレーターによって設定された環境変数WORLD_SIZE、NODE_RANK、MASTER_ADDR、MASTER_PORTが必要です。

エクスポート / TRT のデフォルト設定

  • エクスポート入力:960x544(他のODモデルよりも大きい)、opset 17。 スモークテストでは、Grounding-DINOのエクスポート仕様をテンプレートのエクスポート解像度のままにしておくこと。 エクスポート画像サイズを128x128などの非常に小さいサイズに縮小すると、 torch.onnx.exportの実行中に、コントラストティブテキストヘッドにおいて PyTorch ONNXのシェイプ推論アサーションが発生する可能性がある。
  • 親となる PyTorchgrounding_dinoCLI は、train、evaluate、 inference、export、およびquantize をサポートしています。TensorRT エンジンの生成、 TensorRT 推論、および TensorRT 評価は、references/tao-deploy-grounding-dino.md を通じて実行してください。
  • TRTデータ型:FP32、FP16のみ —INT8はサポートされていません
  • TRT ワークスペース:8192 MB(他の OD モデルよりも 8 倍大きい)
  • TRT max_batch_size: 4

ハードウェア

GPUは最低1基、推奨は4基。GPUあたり24GB以上のVRAM(A100推奨)。Grounding DINOは、テキストエンコーダー(BERT)を使用しているため、標準のDINOよりも処理負荷が高くなります。24GB以上のGPUメモリを推奨します。 16GBのGPUを使用する場合は、batch_sizeを小さくしてください。

エラーパターン

CUDAメモリ不足: batch_size を縮小してください(4 → 2 → 1)。BERTテキストエンコーダーは、ビジョンバックボーンに加えて、かなりのメモリオーバーヘッドをもたらします。

Val アノテーションのカテゴリ ID:損失を正しく計算するには、検証用アノテーションのカテゴリ ID を 0 から開始する必要があります。必要に応じて、アノテーション形式の変換を行ってください。

テキストエンコーダの読み込みエラー:コンテナが bert-base-uncased の重みをダウンロードできることを確認するか、ローカルパスを指定してください。

TAO Toolkit 7.0.0-rc-226 での PyTorch チェックポイントによる量子化が失敗します: コンテナの Grounding-DINO 量子化スクリプトが、チェックポイントの読み込み時に cap_lists=Noneを渡しているため、post_process.py でエラーが発生します。 ONNX量子化では、 エクスポートされたONNXアーティファクトとCOCOキャリブレーションデータが使用されますが、デフォルトのrc-226 PyTorchイメージにはmodelopt.onnx.quantizationモジュールが含まれていません。これは イメージ/SDKのブロック要因として扱い、チェックポイント解決の問題とは見なさないでください。

post_process.py内で mat1 と mat2 のシェイプを乗算できません。テキスト トークンの長さとラベルの位置マップに不整合があります。これは通常、 model.max_text_len がデフォルトの 256 未満に上書きされている一方で、データセットの ラベルマップが依然として長さ 256 の位置マップを使用しているためです。model.max_text_lenを元に戻すか、 同じ長さでラベルマップを再生成してください。

criterion.pyの次元 0 においてインデックスが範囲外です:model.num_queries が、現在のバッチ内のマッチした ODVG ターゲットに対して小さすぎます。 model.num_queriesを増やし、model.num_selectもそれに対応するようにしてください。

images.tar.gz/.jpgに関する NotADirectoryError: 直接実行された TAO CLI が アーカイブパスをディレクトリとしてトラバースしようとしています。アーカイブを解凍し、 関連するimage_dirフィールドに解凍後の画像フォルダを設定してください。アーカイブベースの スキルデータソースは、この理由から実行時に extracted_folderを使用します。

仕様パラメータ / 親モデルの推論

モデル固有の推論マッピングは、config.json ではなく、この MD ファイルに記述する必要があります。生成されたランナーは、このセクションを読み取り、create_job() の前に SDK ヘルパーを使用してマッピングを適用する必要があります。これは、従来のマイクロサービスにおけるinfer_params.pyのフローを踏襲しています。

TAOCoreのgrounding_dino.config.jsonからの推論マッピング:

アクション 仕様フィールド 推論関数 意味
evaluate encryption_key key 暗号化キー
評価する evaluate.checkpoint 親モデル 親ジョブの結果フォルダから推論されたモデルファイル
evaluate evaluate.trt_engine parent_model 親ジョブの結果フォルダから推論されたモデルファイル
evaluate results_dir output_dir 現在のジョブの結果ディレクトリ
export encryption_key key 暗号化キー
エクスポート export.checkpoint 親モデル 親ジョブの結果フォルダから推論されたモデルファイル
export export.onnx_file create_onnx_file ONNX 出力パス
エクスポート results_dir 出力ディレクトリ 現在のジョブの結果ディレクトリ
推論 encryption_key key encryption_keyencryption key
推論 推論のチェックポイント 親モデル 親ジョブの結果フォルダから推論されたモデルファイル
推論 inference.trt_engine parent_model 親ジョブの結果フォルダから推論されたモデルファイル
推論 results_dir output_dir 現在のジョブの結果ディレクトリ
quantize encryption_key key encryption_keykey
quantize quantize.model_path 親モデル 親ジョブの結果フォルダから推測されたモデルファイル
quantize results_dir output_dir 現在のジョブの結果ディレクトリ
train encryption_key key encryption_key暗号化キー
train model.pretrained_backbone_path ptm_if_no_resume_model 再開用チェックポイントが存在しない場合のPTM
train results_dir output_dir 現在のジョブの結果ディレクトリ
train train.pretrained_model_path ptm_if_no_resume_model 再開用チェックポイントが存在しない場合のPTM
train train.resume_training_checkpoint_path resume_model 現在のジョブの結果フォルダから推論されたモデルファイル

parent_modelまたはparent_model_folder の場合、親となる train/export/AutoML の子ジョブ ID をparent_job_id として渡します。SDK は親の結果フォルダを一覧表示し、チェックポイントのアートファクトをフィルタリングして、選択されたモデルファイルまたはフォルダを返します。 これらのマッピングをconfig.jsonに再度追加したり、生成されたランナースクリプトを修正してチェックポイントのパスを推測したりしないでください。

SDK リゾルバーの外で Grounding-DINO チェックポイントを選択する場合は、 目的のエポック/ステップのアーティファクトと完全に一致するようにしてください。例: model_epoch_000_step_00046.pth。gdino_model_latest.pthシンボリックリンクは、 「latest」が明示的に要求された場合にのみ有効です。model.backbone、model.num_queries、model.num_select、 model.num_feature_levels、model.max_text_len、およびエクスポート入力の解像度といった 構造的なモデル設定を、 evaluate、inference、export、および deploy 仕様へと引き継ぎ、チェックポイントと エンジンの形状が一致するようにしてください。

デプロイ

  • tao-deploy-grounding-dino
GitHubで見る
---
name: tao-train-grounding-dino
description: Trains, evaluates, exports, quantizes, and runs inference for a Grounding DINO model that detects objects described by text prompts without a fixed class vocabulary.
license: Apache-2.0
---

# Grounding DINO

Grounding DINO for open-set object detection. Combines DINO-style detection with BERT text encoder for language-guided detection. Detects objects described by text prompts without fixed class vocabulary.

Set train.pretrained_model_path for full Grounding DINO weights or model.pretrained_backbone_path for backbone-only.

For TAO Deploy TensorRT actions (`gen_trt_engine`, TensorRT `evaluate`, and TensorRT `inference`), read `references/tao-deploy-grounding-dino.md` first. Deploy spec templates live in this skill's `references/` folder with the `spec_template_deploy_*.yaml` prefix.

## Dataclass Schemas

Generated TAO Core schemas are packaged in `schemas/<action>.schema.json`, with `schemas/manifest.json` listing available actions. Each generated schema also emits `references/spec_template_<action>.yaml` from the schema top-level `default` field. AutoML enablement is declared at the model layer in `references/skill_info.yaml` via `automl_enabled`. Runnable AutoML still requires `schemas/train.schema.json` and `references/spec_template_train.yaml` to exist and parse. Use the packaged train schema for `automl_default_parameters`, `automl_disabled_parameters`, defaults, min/max bounds, enums, option weights, math conditions, dependencies, and popular parameters. Do not expect `~/tao-core` at runtime; maintainers regenerate schemas/templates before packaging the skill bank.

## Train Action Policy

This model is AutoML-enabled at the model layer. Before handling any train-stage request, read `references/skill_info.yaml` and resolve the run override from either an explicit `automl_policy` value or the user's workflow request. Use `automl_policy: on` by default and only expose `on` / `off` in new launch prompts. Treat phrases like "turn off AutoML", "disable AutoML", "no HPO", or "plain training" as `automl_policy: off` for this run only. When `automl_policy: on`, `automl_enabled: true`, and both `schemas/train.schema.json` and `references/spec_template_train.yaml` are packaged, route the train action through `tao-skill-bank:tao-run-automl` by default with this model's `skill_dir`. Preserve workflow/application overrides for datasets, specs, output directories, GPU/platform settings, parent checkpoints, and `automl_policy`. Use direct model training only when `automl_policy: off` or the packaged train schema/template is missing; in the missing-schema case, report that AutoML is enabled but not runnable for this model until schemas are generated.

Non-train actions such as `evaluate`, `inference`, `export`, and deploy flows stay in this model skill. The per-run `automl_policy` override does not change model metadata.

## Training Requirements

- **Dataset type:** object_detection
- **Formats:** odvg, coco, raw
- **Monitoring metric:** val_mAP50

### Per-Action Dataset Requirements

| Action | Spec Key | Source | Files | List? |
|---|---|---|---|---|
| evaluate | dataset.test_data_sources | eval_dataset | image_dir: images.tar.gz, json_file: annotations.json | No |
| inference | dataset.infer_data_sources.image_dir | inference_dataset | images.tar.gz | Yes |
| inference | dataset.infer_data_sources.captions | workflow prompts | prompt list | Yes |
| quantize | dataset.train_data_sources | train_datasets | image_dir: images.tar.gz, json_file: annotations_odvg.jsonl, label_map: annotations_odvg_labelmap.json | Yes |
| quantize | dataset.val_data_sources | eval_dataset | image_dir: images.tar.gz, json_file: annotations.json | No |
| quantize | dataset.quant_calibration_data_sources | calibration/eval dataset | image_dir: images.tar.gz, json_file: annotations.json | No |
| train | dataset.train_data_sources | train_datasets | image_dir: images.tar.gz, json_file: annotations_odvg.jsonl, label_map: annotations_odvg_labelmap.json | Yes |
| train | dataset.val_data_sources | eval_dataset | image_dir: images.tar.gz, json_file: annotations.json | No |

The runner may source image archives as `images.tar.gz`, but direct local
Docker TAO CLI specs must point `image_dir` to an extracted image directory.
Skill metadata marks these archive-backed image sources with
`runtime: extracted_folder` so a fresh runner can unpack the archive before
launching TAO.

### Typical Spec Overrides

Data source overrides are **mandatory for every action** — the agent MUST construct data source paths from the Per-Action Dataset Requirements table above and include them in `spec_overrides`.

```python
S3_TRAIN = "s3://bucket/data/train"
S3_EVAL = "s3://bucket/data/eval"
```

**train (mandatory data sources):**
```python
{
    "train.num_epochs": 10,
    "train.checkpoint_interval": 10,
    "train.validation_interval": 10,
    "train.num_gpus": 1,
    "dataset.train_data_sources": [{"image_dir": f"{S3_TRAIN}/images.tar.gz", "json_file": f"{S3_TRAIN}/annotations_odvg.jsonl", "label_map": f"{S3_TRAIN}/annotations_odvg_labelmap.json"}],
    "dataset.val_data_sources": {"image_dir": f"{S3_EVAL}/images.tar.gz", "json_file": f"{S3_EVAL}/annotations.json"},
}
```

**deploy/gen_trt_engine (use `references/tao-deploy-grounding-dino.md`):**
```python
{
    "gen_trt_engine.onnx_file": "<exported_onnx_uri>",
    "gen_trt_engine.trt_engine": "<output_engine_path>",
    "gen_trt_engine.tensorrt.data_type": "FP16",
}
```

**inference (mandatory data sources):**
```python
{
    "inference.checkpoint": "<selected train/AutoML checkpoint>",
    "dataset.infer_data_sources.image_dir": [f"{S3_EVAL}/images.tar.gz"],
    "dataset.infer_data_sources.captions": [
        "fire extinguisher",
        "cone",
        "cart",
        "forklift"
    ],
}
```

**evaluate (mandatory data sources):**
```python
{
    "evaluate.checkpoint": "<selected train/AutoML checkpoint>",
    "dataset.test_data_sources": {"image_dir": f"{S3_EVAL}/images.tar.gz", "json_file": f"{S3_EVAL}/annotations.json"},
}
```

**quantize (mandatory data sources):**
```python
{
    "quantize.model_path": "<selected train checkpoint or exported ONNX model>",
    "dataset.train_data_sources": [{"image_dir": f"{S3_TRAIN}/images.tar.gz", "json_file": f"{S3_TRAIN}/annotations_odvg.jsonl", "label_map": f"{S3_TRAIN}/annotations_odvg_labelmap.json"}],
    "dataset.val_data_sources": {"image_dir": f"{S3_EVAL}/images.tar.gz", "json_file": f"{S3_EVAL}/annotations.json"},
    "dataset.quant_calibration_data_sources": {"image_dir": f"{S3_EVAL}/images.tar.gz", "json_file": f"{S3_EVAL}/annotations.json"},
}
```
## Eval Dataset

Optional. Validation uses COCO-format annotations for mAP even though training can use ODVG format.

## Important Parameters

- **model.backbone**: Default swin_tiny_224_1k. Also supports resnet_50 and other Swin variants. Swin generally performs better for grounding tasks.
- **model.text_encoder_type**: BERT model for text encoding. Default bert-base-uncased. max_text_len defaults to 256.
- **model.max_text_len**: Keep this aligned with the dataset label/token
  position maps. Do not shrink it for smoke tests unless the corresponding
  label maps are regenerated with the same length; otherwise validation can
  fail with a matrix shape mismatch between token probabilities and position
  maps.
- **train.optim.lr**: Learning rate. Default 2e-4. lr_backbone 2e-5. Supports bf16 precision in addition to fp16/fp32.
- **dataset.max_labels**: Maximum labels per image during training. Default 50. Increase for dense annotation datasets.
- **model.num_queries**: Object queries. Default 900 (higher than DINO's 300) due to open-vocabulary nature.
- **model.num_queries / model.num_select**: Keep `num_queries` high enough
  for the number of matched ODVG targets in a batch. Very small smoke values
  such as 20 can fail during Hungarian target indexing on dense images; use at
  least 100 for minimal Grounding DINO smoke runs unless the dataset is known
  to have fewer objects per image.
- **train.optim.lr_steps**: MultiStep LR schedule. Default [10].

## Multi-GPU / Multi-Node

**Launch method:** Lightning-managed (single `python` process, Lightning spawns workers).

| Spec Key | Description | Default |
|----------|-------------|---------|
| `train.num_gpus` | Number of GPUs | 1 |
| `train.gpu_ids` | GPU device indices | [0] |
| `train.num_nodes` | Number of nodes | 1 |
| `train.distributed_strategy` | `ddp` or `fsdp` | `ddp` |

Same DDP/FSDP behavior as DINO. Multi-node requires `WORLD_SIZE`, `NODE_RANK`, `MASTER_ADDR`, `MASTER_PORT` env vars set by orchestrator.

## Export / TRT Defaults

- Export input: 960x544 (larger than other OD models), opset 17. Keep
  Grounding-DINO export specs at the template export resolution for smoke tests;
  reducing export to very small image sizes such as 128x128 can trigger a
  PyTorch ONNX shape-inference assertion in the contrastive text head during
  `torch.onnx.export`.
- The parent PyTorch `grounding_dino` CLI supports `train`, `evaluate`,
  `inference`, `export`, and `quantize`. Run TensorRT engine generation,
  TensorRT inference, and TensorRT evaluation through `references/tao-deploy-grounding-dino.md`.
- TRT data types: FP32, FP16 only — **INT8 is NOT supported**
- TRT workspace: 8192 MB (8x larger than other OD models)
- TRT max_batch_size: 4

## Hardware

Minimum 1 GPU(s), recommended 4 GPU(s). 24GB+ (A100 recommended) VRAM per GPU. Grounding DINO is heavier than standard DINO due to the text encoder (BERT). 24GB+ GPU memory recommended. Reduce batch_size for 16GB GPUs.

## Error Patterns

**CUDA out of memory**: Reduce batch_size (4 -> 2 -> 1). The BERT text encoder adds significant memory overhead on top of the vision backbone.

**Val annotation category IDs**: Validation annotations should have category IDs starting from 0 for correct loss computation. Use annotation format conversion if needed.

**Text encoder loading error**: Ensure the container has access to download bert-base-uncased weights or provide a local path.

**Quantize with a PyTorch checkpoint fails in TAO Toolkit 7.0.0-rc-226**:
The container's Grounding-DINO quantize script passes `cap_lists=None` when
loading a checkpoint, which fails in `post_process.py`. ONNX quantization uses
the exported ONNX artifact and COCO calibration data, but the default rc-226
PyTorch image also lacks the `modelopt.onnx.quantization` module. Treat this as
an image/SDK blocker, not a checkpoint resolver issue.

**mat1 and mat2 shapes cannot be multiplied in `post_process.py`**: The text
token length and label position maps are inconsistent, commonly because
`model.max_text_len` was overridden below the default 256 while the dataset
label maps still use 256-length position maps. Restore `model.max_text_len` or
regenerate the label maps with the same length.

**index is out of bounds for dimension 0 in `criterion.py`**: `model.num_queries`
is too small for the matched ODVG targets in the current batch. Increase
`model.num_queries` and keep `model.num_select` compatible with it.

**NotADirectoryError with `images.tar.gz/<image>.jpg`**: The direct TAO CLI is
trying to traverse an archive path as a directory. Extract the archive and set
the relevant `image_dir` field to the extracted image folder; archive-backed
skill data sources use `runtime: extracted_folder` for this reason.

## Spec Param / Parent Model Inference

Model-specific inference mappings belong in this MD file, not in `config.json`. Generated runners should read this section and apply the mappings with SDK helpers before `create_job()`. This mirrors the old microservices `infer_params.py` flow.

Inference mappings from TAO Core `grounding_dino.config.json`:

| Action | Spec Field | Inference Function | Meaning |
|---|---|---|---|
| evaluate | `encryption_key` | `key` | encryption key |
| evaluate | `evaluate.checkpoint` | `parent_model` | model file inferred from the parent job results folder |
| evaluate | `evaluate.trt_engine` | `parent_model` | model file inferred from the parent job results folder |
| evaluate | `results_dir` | `output_dir` | current job results directory |
| export | `encryption_key` | `key` | encryption key |
| export | `export.checkpoint` | `parent_model` | model file inferred from the parent job results folder |
| export | `export.onnx_file` | `create_onnx_file` | output ONNX path |
| export | `results_dir` | `output_dir` | current job results directory |
| inference | `encryption_key` | `key` | encryption key |
| inference | `inference.checkpoint` | `parent_model` | model file inferred from the parent job results folder |
| inference | `inference.trt_engine` | `parent_model` | model file inferred from the parent job results folder |
| inference | `results_dir` | `output_dir` | current job results directory |
| quantize | `encryption_key` | `key` | encryption key |
| quantize | `quantize.model_path` | `parent_model` | model file inferred from the parent job results folder |
| quantize | `results_dir` | `output_dir` | current job results directory |
| train | `encryption_key` | `key` | encryption key |
| train | `model.pretrained_backbone_path` | `ptm_if_no_resume_model` | PTM when no resume checkpoint exists |
| train | `results_dir` | `output_dir` | current job results directory |
| train | `train.pretrained_model_path` | `ptm_if_no_resume_model` | PTM when no resume checkpoint exists |
| train | `train.resume_training_checkpoint_path` | `resume_model` | model file inferred from the current job results folder |

For `parent_model` or `parent_model_folder`, pass the upstream train/export/AutoML child job id as `parent_job_id`. The SDK lists the parent result folder, filters checkpoint artifacts, and returns the selected model file or folder. Do not add these mappings back to `config.json` and do not patch generated runner scripts to guess checkpoint paths.

When selecting a Grounding-DINO checkpoint outside the SDK resolver, match the
intended epoch/step artifact exactly, for example
`model_epoch_000_step_00046.pth`. The `gdino_model_latest.pth` symlink is valid
only when latest is explicitly requested. Carry structural model settings such
as `model.backbone`, `model.num_queries`, `model.num_select`,
`model.num_feature_levels`, `model.max_text_len`, and export input resolution
forward into evaluate, inference, export, and deploy specs so checkpoint and
engine shapes match.

## Deployment

- [tao-deploy-grounding-dino](references/tao-deploy-grounding-dino.md)

tao-train-grounding-dinoをインストール

スキルファイルをダウンロードし、.claude/skills/ ディレクトリに解凍してください。

ZIPをダウンロード

リポジトリをクローンし、スキルファイルをプロジェクトにコピーしてください。

git clone https://github.com/NVIDIA/skills/tree/main/skills/tao-train-grounding-dino # Copy SKILL.md to your .claude/skills/ directory

コピー コピー
クイックセットアップ: スキルフォルダを .claude/skills/ にコピーしてください。 Claude が自動的にそのスキルを検出して使用します。
リポジトリ NVIDIA/skills

関連スキル

web-search
更新された時間 2026年6月29日
webapp-testing
更新された時間 2026年6月29日
lark-base
更新された時間 2026年7月5日
agentmail
更新された時間 2026年6月29日
OR