选项
首页首页 Skill 数据科学与机器学习 tao-train-grounding-dino

tao-train-grounding-dino

NVIDIA/skills NVIDIA/skills

对一个Grounding DINO模型进行训练、评估、导出、量化并执行推理,该模型能够根据文本提示检测物体,且不依赖固定的类别词汇表。

...展开全部
0
更新时间 2026-09-25

DINO的基线模型

适用于开放集物体检测的 Grounding DINO。该模型将 DINO 风格的检测与 BERT 文本编码器相结合,实现语言引导式检测。可检测由文本提示描述的物体,且无需固定的类别词汇表。

若需使用完整的 Grounding DINO 权重,请设置 `train.pretrained_model_path`;若仅需使用骨干网络,请设置 `model.pretrained_backbone_path`。

对于 TAO Deploy 的 TensorRT 操作(gen_trt_engine、TensorRTevaluate 和 TensorRTinference),请先阅读references/tao-deploy-grounding-dino.md。 部署规范模板位于该技能的references/文件夹中,文件名前缀为spec_template_deploy_*.yaml。

数据类模式

生成的 TAO Core 模式打包在schemas/.schema.json 中,其中schemas/manifest.json列出了可用的操作。每个生成的模式还会从模式顶层的default字段生成references/spec_template_.yaml 文件。 AutoML 的启用通过references/skill_info.yaml中的automl_enabled 在模型层进行声明。可运行的 AutoML 仍需要schemas/train.schema.json和references/spec_template_train.yaml文件存在并被解析。 请使用打包的训练模式来定义automl_default_parameters、automl_disabled_parameters、默认值、最小/最大边界、枚举、选项权重、数学条件、依赖关系以及常用参数。运行时不要依赖~/tao-core;维护人员会在打包技能库之前重新生成模式和模板。

训练操作策略

该模型在模型层启用了 AutoML。在处理任何训练阶段的请求之前,请读取references/skill_info.yaml,并根据显式的automl_policy值或用户的工作流请求解析运行覆盖设置。默认使用automl_policy: on,仅在新启动提示中提供on /off 选项。 将“关闭 AutoML”、“禁用 AutoML”、“不使用 HPO”或“普通训练”等短语视为仅针对本次运行的automl_policy: off。 当`automl_policy: on`、`automl_enabled: true`,且`schemas/train.schema.json` 和 `references/spec_template_train.yaml` 均已打包时,默认将训练操作通过 `tao-skill-bank:tao-run-automl` 路由,并使用该模型的`skill_dir`。 保留工作流/应用程序对数据集、规范、输出目录、GPU/平台设置、父检查点以及automl_policy 的覆盖设置。 仅当 `automl_policy: off` 或打包的训练模式/模板缺失时,才使用直接模型训练;在模式缺失的情况下,应报告该模型已启用 AutoML 但无法运行,直到生成模式为止。

诸如评估、推理、导出和部署流程等非训练操作仍保留在此模型技能中。每次运行的automl_policy覆盖设置不会更改模型元数据。

训练要求

  • 数据集类型:object_detection
  • 格式:odvg、coco、raw
  • 监控指标:val_mAP50

按操作划分的数据集要求

动作 规范键 来源 文件 列表?
评估 数据集.test_data_sources eval_dataset image_dir:images.tar.gz,json_file:annotations.json 否
推理 数据集.推理数据源.图像目录 推理数据集 images.tar.gz 是
推理 dataset.infer_data_sources.captions 工作流提示 提示列表 是
量化 数据集.训练数据源 train_datasets image_dir: images.tar.gz, json_file: annotations_odvg.jsonl, label_map: annotations_odvg_labelmap.json 是
量化 数据集.val_data_sources eval_dataset image_dir: images.tar.gz, json_file: annotations.json 否
量化 数据集.量化校准数据源 校准/评估数据集 image_dir:images.tar.gz,json_file:annotations.json 否
训练 数据集.train_data_sources train_datasets image_dir:images.tar.gz,json_file:annotations_odvg.jsonl,label_map:annotations_odvg_labelmap.json 是
train dataset.val_data_sources eval_dataset image_dir:images.tar.gz,json_file:annotations.json 否

运行器可将image.tar.gz 等镜像归档文件作为镜像源,但直接在本地 使用的 Docker TAO CLI 规范必须将image_dir指向已解压的镜像目录。 技能元数据会为这些基于归档文件的镜像源标记 runtime: extracted_folder,以便新的运行器在 启动 TAO 之前能先解压该归档文件。

典型的规范覆盖

每个操作都必须覆盖数据源——代理必须根据上文的“按操作划分的数据集要求”表构建数据源路径,并将它们包含在spec_overrides 中。

S3_TRAIN = "s3://bucket/data/train"
S3_EVAL = "s3://bucket/data/eval"

训练(必填数据源):

{
    "train.num_epochs": 10,
    "train.checkpoint_interval": 10,
    "train.validation_interval": 10,
    "train.num_gpus": 1,
    "dataset.train_data_sources": [{"image_dir": f"{S3_TRAIN}/images.tar.gz", "json_file": f"{S3_TRAIN}/annotations_odvg.jsonl", "label_map": f"{S3_TRAIN}/annotations_odvg_labelmap.json"}],
    "dataset.val_data_sources": {"image_dir": f"{S3_EVAL}/images.tar.gz", "json_file": f"{S3_EVAL}/annotations.json"},
}

deploy/gen_trt_engine(使用references/tao-deploy-grounding-dino.md):

{
    "gen_trt_engine.onnx_file": "",
    "gen_trt_engine.trt_engine": "",
    "gen_trt_engine.tensorrt.data_type": "FP16",
}

推理(必填数据源):

{
    "inference.checkpoint": "<选定的训练/AutoML检查点>",
    "dataset.infer_data_sources.image_dir": [f"{S3_EVAL}/images.tar.gz"],
    "dataset.infer_data_sources.captions": [
        "灭火器",
        "锥形路标",
        "手推车",
        "叉车"
    ],
}

评估(必填数据源):

{
    "evaluate.checkpoint": "<选定的训练/AutoML 检查点>",
    "dataset.test_data_sources": {"image_dir": f"{S3_EVAL}/images.tar.gz", "json_file": f"{S3_EVAL}/annotations.json"},
}

量化(必填数据源):

{
    "quantize.model_path": "",
    "dataset.train_data_sources": [{"image_dir": f"{S3_TRAIN}/images.tar.gz", "json_file": f"{S3_TRAIN}/annotations_odvg.jsonl", "label_map": f"{S3_TRAIN}/annotations_odvg_labelmap.json"}],
    "dataset.val_data_sources": {"image_dir": f"{S3_EVAL}/images.tar.gz", "json_file": f"{S3_EVAL}/annotations.json"},
    "dataset.quant_calibration_data_sources": {"image_dir": f"{S3_EVAL}/images.tar.gz", "json_file": f"{S3_EVAL}/annotations.json"},
}

评估数据集

可选。尽管训练可以使用ODVG格式,但评估时mAP指标需使用COCO格式的标注。

重要参数

  • model.backbone:默认值为 swin_tiny_224_1k。同时也支持 resnet_50 及其他 Swin 变体。Swin 在定位任务中通常表现更佳。
  • model.text_encoder_type:用于文本编码的 BERT 模型。默认为 bert-base-uncased。max_text_len 默认为 256。
  • model.max_text_len:请确保此参数与数据集的标签/令牌 位置映射保持一致。除非相应的 标签映射已按相同长度重新生成,否则请勿在快速测试中缩短该值;否则验证可能会 因令牌概率与位置 映射之间的矩阵形状不匹配而失败。
  • train.optim.lr:学习率。默认值为 2e-4。lr_backbone 为 2e-5。除 fp16/fp32 外,还支持 bf16 精度。
  • dataset.max_labels:训练过程中每张图像的最大标签数。默认值为 50。对于标注密集的数据集,请增加此值。
  • model.num_queries:对象查询次数。默认值为 900(高于 DINO 的 300),这是由于其开放词汇表的特性。
  • model.num_queries / model.num_select:应将num_queries设置得足够高, 以满足批次中匹配的 ODVG 目标数量。 非常小的 smoke 值 (例如 20)在对高密度图像进行匈牙利目标索引时可能会失败;除非已知数据集 中每张图像的物体数量较少,否则对于最基本的 Grounding DINO smoke 运行,请使用 至少 100。
  • train.optim.lr_steps:多步学习率(LR)调度。默认值 [10]。

多GPU / 多节点

启动方式:Lightning 管理(单个Python进程,由 Lightning 生成工作进程)。

参数说明 描述 默认值
train.num_gpus GPU 数量 1
train.gpu_ids GPU 设备索引 [0]
train.num_nodes 节点数 1
train.distributed_strategy ddp或fsdp ddp

与 DINO 具有相同的 DDP/FSDP 行为。多节点模式需要由编排器设置WORLD_SIZE、NODE_RANK、MASTER_ADDR、MASTER_PORT环境变量。

导出 / TRT 默认值

  • 导出输入:960x544(大于其他 OD 模型),操作集 17。保留 Grounding-DINO 的导出规格保持在模板导出分辨率下进行初步测试; 若将导出图像尺寸缩小至 128x128 等极小尺寸,可能会在 torch.onnx.export 过程中触发对比文本头中的 PyTorch ONNX 形状推理断言。
  • 父级 PyTorchgrounding_dinoCLI 支持训练、评估、 推理、导出和 量化。通过references/tao-deploy-grounding-dino.md 运行 TensorRT 引擎生成、 TensorRT 推理和 TensorRT 评估。
  • TRT 数据类型:仅支持FP32、FP16 —不支持 INT8
  • TRT 工作区:8192 MB(比其他 OD 模型大 8 倍)
  • TRT max_batch_size:4

硬件

至少 1 个 GPU,建议 4 个 GPU。每个 GPU 需配备 24GB 以上(建议使用 A100)的显存。由于采用了文本编码器(BERT),Grounding DINO 的计算负载比标准 DINO 更高。建议使用 24GB 以上的 GPU 显存。 若使用 16GB GPU,请减小 batch_size。

错误模式

CUDA 内存不足:减少 batch_size(4 → 2 → 1)。BERT 文本编码器会在视觉骨干网络的基础上增加显著的内存开销。

Val 标注类别 ID:为确保损失计算正确,验证标注的类别 ID 应从 0 开始。如有需要,请使用标注格式转换。

文本编码器加载错误:确保容器能够下载 bert-base-uncased 权重,或提供本地路径。

在 TAO Toolkit 7.0.0-rc-226 中使用 PyTorch 检查点进行量化失败: 容器的 Grounding-DINO 量化脚本在 加载检查点时传入了cap_lists=None,导致post_process.py 中的操作失败。 ONNX量化使用 导出的ONNX构建产物和COCO校准数据,但默认的rc-226 PyTorch镜像也缺少modelopt.onnx.quantization模块。请将此问题视为 镜像/SDK层面的阻塞问题,而非检查点解析器的问题。

在post_process.py中无法对 mat1 和 mat2 的形状进行乘法运算:文本 令牌长度与标签位置映射不一致,通常是因为 model.max_text_len被重写为小于默认值 256,而数据集 标签映射仍使用长度为 256 的位置映射。 请恢复model.max_text_len的默认值,或 使用相同长度重新生成标签映射。

criterion.py中维度 0 的索引超出范围:model.num_queries 对于当前批次中匹配的 ODVG 目标而言过小。请增加 model.num_queries,并确保model.num_select与其保持兼容。

images.tar.gz/.jpg出现 NotADirectoryError:直接 TAO CLI 正 试图将归档路径作为目录进行遍历。请解压归档文件,并将 相关的image_dir字段设置为解压后的图像文件夹;基于归档的 技能数据源正因如此使用runtime:extracted_folder。

规格参数 / 父模型推理

特定于模型的推理映射应置于此 MD 文件中,而非config.json 中。生成的运行器应在调用 create_job() 之前读取本节内容,并借助 SDK 辅助函数应用这些映射。这与旧版微服务中的infer_params.py流程一致。

来自 TAO Core `grounding_dino.config.json` 的推理映射:

操作 规范字段 推断函数 含义
evaluate 加密密钥 密钥 加密密钥
评估 evaluate.checkpoint 父模型 从父任务结果文件夹推导出的模型文件
evaluate evaluate.trt_engine parent_model 从父任务结果文件夹推导出的模型文件
evaluate results_dir output_dir 当前作业结果目录
导出 加密密钥 key 加密密钥
导出 export.checkpoint 父模型 从父任务结果文件夹推导出的模型文件
export export.onnx_file create_onnx_file ONNX 输出路径
导出 结果目录 output_dir 当前作业结果目录
推理 加密密钥 key 加密密钥
推理 推理检查点 父模型 从父任务结果文件夹推导出的模型文件
推理 inference.trt_engine 父模型 从父任务结果文件夹推导出的模型文件
推断 结果目录 output_dir 当前任务结果目录
量化 加密密钥 key 加密密钥
量化 量化.模型路径 父模型 从父任务结果文件夹推导出的模型文件
量化 results_dir output_dir 当前作业结果目录
train 加密密钥 密钥 加密密钥
train 模型.预训练骨干路径 ptm_if_no_resume_model 当不存在恢复检查点时的 PTM
train results_dir output_dir 当前作业结果目录
train train.pretrained_model_path ptm_if_no_resume_model 当不存在恢复检查点时的 PTM
train train.resume_training_checkpoint_path resume_model 从当前作业结果文件夹推导出的模型文件

对于parent_model或parent_model_folder,请将上游 train/export/AutoML 子任务的 ID 作为parent_job_id 传入。SDK 会列出父任务的结果文件夹,过滤检查点资源,并返回选定的模型文件或文件夹。 请勿将这些映射关系重新添加回config.json,也请勿修改生成的运行器脚本以推测检查点路径。

在 SDK 解析器之外选择 Grounding-DINO 检查点时,请与 目标 epoch/step 构建结果完全匹配,例如 model_epoch_000_step_00046.pth。gdino_model_latest.pth符号链接仅在 显式请求 latest 时才有效。 将结构化模型设置(如 model.backbone、model.num_queries、model.num_select、 model.num_feature_levels、model.max_text_len 以及导出输入分辨率) 向前传递至 evaluate、inference、export 和 deploy 规范中,以确保检查点和 引擎的结构保持一致。

部署

  • tao-deploy-grounding-dino
在 GitHub 上查看
---
name: tao-train-grounding-dino
description: Trains, evaluates, exports, quantizes, and runs inference for a Grounding DINO model that detects objects described by text prompts without a fixed class vocabulary.
license: Apache-2.0
---

# Grounding DINO

Grounding DINO for open-set object detection. Combines DINO-style detection with BERT text encoder for language-guided detection. Detects objects described by text prompts without fixed class vocabulary.

Set train.pretrained_model_path for full Grounding DINO weights or model.pretrained_backbone_path for backbone-only.

For TAO Deploy TensorRT actions (`gen_trt_engine`, TensorRT `evaluate`, and TensorRT `inference`), read `references/tao-deploy-grounding-dino.md` first. Deploy spec templates live in this skill's `references/` folder with the `spec_template_deploy_*.yaml` prefix.

## Dataclass Schemas

Generated TAO Core schemas are packaged in `schemas/<action>.schema.json`, with `schemas/manifest.json` listing available actions. Each generated schema also emits `references/spec_template_<action>.yaml` from the schema top-level `default` field. AutoML enablement is declared at the model layer in `references/skill_info.yaml` via `automl_enabled`. Runnable AutoML still requires `schemas/train.schema.json` and `references/spec_template_train.yaml` to exist and parse. Use the packaged train schema for `automl_default_parameters`, `automl_disabled_parameters`, defaults, min/max bounds, enums, option weights, math conditions, dependencies, and popular parameters. Do not expect `~/tao-core` at runtime; maintainers regenerate schemas/templates before packaging the skill bank.

## Train Action Policy

This model is AutoML-enabled at the model layer. Before handling any train-stage request, read `references/skill_info.yaml` and resolve the run override from either an explicit `automl_policy` value or the user's workflow request. Use `automl_policy: on` by default and only expose `on` / `off` in new launch prompts. Treat phrases like "turn off AutoML", "disable AutoML", "no HPO", or "plain training" as `automl_policy: off` for this run only. When `automl_policy: on`, `automl_enabled: true`, and both `schemas/train.schema.json` and `references/spec_template_train.yaml` are packaged, route the train action through `tao-skill-bank:tao-run-automl` by default with this model's `skill_dir`. Preserve workflow/application overrides for datasets, specs, output directories, GPU/platform settings, parent checkpoints, and `automl_policy`. Use direct model training only when `automl_policy: off` or the packaged train schema/template is missing; in the missing-schema case, report that AutoML is enabled but not runnable for this model until schemas are generated.

Non-train actions such as `evaluate`, `inference`, `export`, and deploy flows stay in this model skill. The per-run `automl_policy` override does not change model metadata.

## Training Requirements

- **Dataset type:** object_detection
- **Formats:** odvg, coco, raw
- **Monitoring metric:** val_mAP50

### Per-Action Dataset Requirements

| Action | Spec Key | Source | Files | List? |
|---|---|---|---|---|
| evaluate | dataset.test_data_sources | eval_dataset | image_dir: images.tar.gz, json_file: annotations.json | No |
| inference | dataset.infer_data_sources.image_dir | inference_dataset | images.tar.gz | Yes |
| inference | dataset.infer_data_sources.captions | workflow prompts | prompt list | Yes |
| quantize | dataset.train_data_sources | train_datasets | image_dir: images.tar.gz, json_file: annotations_odvg.jsonl, label_map: annotations_odvg_labelmap.json | Yes |
| quantize | dataset.val_data_sources | eval_dataset | image_dir: images.tar.gz, json_file: annotations.json | No |
| quantize | dataset.quant_calibration_data_sources | calibration/eval dataset | image_dir: images.tar.gz, json_file: annotations.json | No |
| train | dataset.train_data_sources | train_datasets | image_dir: images.tar.gz, json_file: annotations_odvg.jsonl, label_map: annotations_odvg_labelmap.json | Yes |
| train | dataset.val_data_sources | eval_dataset | image_dir: images.tar.gz, json_file: annotations.json | No |

The runner may source image archives as `images.tar.gz`, but direct local
Docker TAO CLI specs must point `image_dir` to an extracted image directory.
Skill metadata marks these archive-backed image sources with
`runtime: extracted_folder` so a fresh runner can unpack the archive before
launching TAO.

### Typical Spec Overrides

Data source overrides are **mandatory for every action** — the agent MUST construct data source paths from the Per-Action Dataset Requirements table above and include them in `spec_overrides`.

```python
S3_TRAIN = "s3://bucket/data/train"
S3_EVAL = "s3://bucket/data/eval"
```

**train (mandatory data sources):**
```python
{
    "train.num_epochs": 10,
    "train.checkpoint_interval": 10,
    "train.validation_interval": 10,
    "train.num_gpus": 1,
    "dataset.train_data_sources": [{"image_dir": f"{S3_TRAIN}/images.tar.gz", "json_file": f"{S3_TRAIN}/annotations_odvg.jsonl", "label_map": f"{S3_TRAIN}/annotations_odvg_labelmap.json"}],
    "dataset.val_data_sources": {"image_dir": f"{S3_EVAL}/images.tar.gz", "json_file": f"{S3_EVAL}/annotations.json"},
}
```

**deploy/gen_trt_engine (use `references/tao-deploy-grounding-dino.md`):**
```python
{
    "gen_trt_engine.onnx_file": "<exported_onnx_uri>",
    "gen_trt_engine.trt_engine": "<output_engine_path>",
    "gen_trt_engine.tensorrt.data_type": "FP16",
}
```

**inference (mandatory data sources):**
```python
{
    "inference.checkpoint": "<selected train/AutoML checkpoint>",
    "dataset.infer_data_sources.image_dir": [f"{S3_EVAL}/images.tar.gz"],
    "dataset.infer_data_sources.captions": [
        "fire extinguisher",
        "cone",
        "cart",
        "forklift"
    ],
}
```

**evaluate (mandatory data sources):**
```python
{
    "evaluate.checkpoint": "<selected train/AutoML checkpoint>",
    "dataset.test_data_sources": {"image_dir": f"{S3_EVAL}/images.tar.gz", "json_file": f"{S3_EVAL}/annotations.json"},
}
```

**quantize (mandatory data sources):**
```python
{
    "quantize.model_path": "<selected train checkpoint or exported ONNX model>",
    "dataset.train_data_sources": [{"image_dir": f"{S3_TRAIN}/images.tar.gz", "json_file": f"{S3_TRAIN}/annotations_odvg.jsonl", "label_map": f"{S3_TRAIN}/annotations_odvg_labelmap.json"}],
    "dataset.val_data_sources": {"image_dir": f"{S3_EVAL}/images.tar.gz", "json_file": f"{S3_EVAL}/annotations.json"},
    "dataset.quant_calibration_data_sources": {"image_dir": f"{S3_EVAL}/images.tar.gz", "json_file": f"{S3_EVAL}/annotations.json"},
}
```
## Eval Dataset

Optional. Validation uses COCO-format annotations for mAP even though training can use ODVG format.

## Important Parameters

- **model.backbone**: Default swin_tiny_224_1k. Also supports resnet_50 and other Swin variants. Swin generally performs better for grounding tasks.
- **model.text_encoder_type**: BERT model for text encoding. Default bert-base-uncased. max_text_len defaults to 256.
- **model.max_text_len**: Keep this aligned with the dataset label/token
  position maps. Do not shrink it for smoke tests unless the corresponding
  label maps are regenerated with the same length; otherwise validation can
  fail with a matrix shape mismatch between token probabilities and position
  maps.
- **train.optim.lr**: Learning rate. Default 2e-4. lr_backbone 2e-5. Supports bf16 precision in addition to fp16/fp32.
- **dataset.max_labels**: Maximum labels per image during training. Default 50. Increase for dense annotation datasets.
- **model.num_queries**: Object queries. Default 900 (higher than DINO's 300) due to open-vocabulary nature.
- **model.num_queries / model.num_select**: Keep `num_queries` high enough
  for the number of matched ODVG targets in a batch. Very small smoke values
  such as 20 can fail during Hungarian target indexing on dense images; use at
  least 100 for minimal Grounding DINO smoke runs unless the dataset is known
  to have fewer objects per image.
- **train.optim.lr_steps**: MultiStep LR schedule. Default [10].

## Multi-GPU / Multi-Node

**Launch method:** Lightning-managed (single `python` process, Lightning spawns workers).

| Spec Key | Description | Default |
|----------|-------------|---------|
| `train.num_gpus` | Number of GPUs | 1 |
| `train.gpu_ids` | GPU device indices | [0] |
| `train.num_nodes` | Number of nodes | 1 |
| `train.distributed_strategy` | `ddp` or `fsdp` | `ddp` |

Same DDP/FSDP behavior as DINO. Multi-node requires `WORLD_SIZE`, `NODE_RANK`, `MASTER_ADDR`, `MASTER_PORT` env vars set by orchestrator.

## Export / TRT Defaults

- Export input: 960x544 (larger than other OD models), opset 17. Keep
  Grounding-DINO export specs at the template export resolution for smoke tests;
  reducing export to very small image sizes such as 128x128 can trigger a
  PyTorch ONNX shape-inference assertion in the contrastive text head during
  `torch.onnx.export`.
- The parent PyTorch `grounding_dino` CLI supports `train`, `evaluate`,
  `inference`, `export`, and `quantize`. Run TensorRT engine generation,
  TensorRT inference, and TensorRT evaluation through `references/tao-deploy-grounding-dino.md`.
- TRT data types: FP32, FP16 only — **INT8 is NOT supported**
- TRT workspace: 8192 MB (8x larger than other OD models)
- TRT max_batch_size: 4

## Hardware

Minimum 1 GPU(s), recommended 4 GPU(s). 24GB+ (A100 recommended) VRAM per GPU. Grounding DINO is heavier than standard DINO due to the text encoder (BERT). 24GB+ GPU memory recommended. Reduce batch_size for 16GB GPUs.

## Error Patterns

**CUDA out of memory**: Reduce batch_size (4 -> 2 -> 1). The BERT text encoder adds significant memory overhead on top of the vision backbone.

**Val annotation category IDs**: Validation annotations should have category IDs starting from 0 for correct loss computation. Use annotation format conversion if needed.

**Text encoder loading error**: Ensure the container has access to download bert-base-uncased weights or provide a local path.

**Quantize with a PyTorch checkpoint fails in TAO Toolkit 7.0.0-rc-226**:
The container's Grounding-DINO quantize script passes `cap_lists=None` when
loading a checkpoint, which fails in `post_process.py`. ONNX quantization uses
the exported ONNX artifact and COCO calibration data, but the default rc-226
PyTorch image also lacks the `modelopt.onnx.quantization` module. Treat this as
an image/SDK blocker, not a checkpoint resolver issue.

**mat1 and mat2 shapes cannot be multiplied in `post_process.py`**: The text
token length and label position maps are inconsistent, commonly because
`model.max_text_len` was overridden below the default 256 while the dataset
label maps still use 256-length position maps. Restore `model.max_text_len` or
regenerate the label maps with the same length.

**index is out of bounds for dimension 0 in `criterion.py`**: `model.num_queries`
is too small for the matched ODVG targets in the current batch. Increase
`model.num_queries` and keep `model.num_select` compatible with it.

**NotADirectoryError with `images.tar.gz/<image>.jpg`**: The direct TAO CLI is
trying to traverse an archive path as a directory. Extract the archive and set
the relevant `image_dir` field to the extracted image folder; archive-backed
skill data sources use `runtime: extracted_folder` for this reason.

## Spec Param / Parent Model Inference

Model-specific inference mappings belong in this MD file, not in `config.json`. Generated runners should read this section and apply the mappings with SDK helpers before `create_job()`. This mirrors the old microservices `infer_params.py` flow.

Inference mappings from TAO Core `grounding_dino.config.json`:

| Action | Spec Field | Inference Function | Meaning |
|---|---|---|---|
| evaluate | `encryption_key` | `key` | encryption key |
| evaluate | `evaluate.checkpoint` | `parent_model` | model file inferred from the parent job results folder |
| evaluate | `evaluate.trt_engine` | `parent_model` | model file inferred from the parent job results folder |
| evaluate | `results_dir` | `output_dir` | current job results directory |
| export | `encryption_key` | `key` | encryption key |
| export | `export.checkpoint` | `parent_model` | model file inferred from the parent job results folder |
| export | `export.onnx_file` | `create_onnx_file` | output ONNX path |
| export | `results_dir` | `output_dir` | current job results directory |
| inference | `encryption_key` | `key` | encryption key |
| inference | `inference.checkpoint` | `parent_model` | model file inferred from the parent job results folder |
| inference | `inference.trt_engine` | `parent_model` | model file inferred from the parent job results folder |
| inference | `results_dir` | `output_dir` | current job results directory |
| quantize | `encryption_key` | `key` | encryption key |
| quantize | `quantize.model_path` | `parent_model` | model file inferred from the parent job results folder |
| quantize | `results_dir` | `output_dir` | current job results directory |
| train | `encryption_key` | `key` | encryption key |
| train | `model.pretrained_backbone_path` | `ptm_if_no_resume_model` | PTM when no resume checkpoint exists |
| train | `results_dir` | `output_dir` | current job results directory |
| train | `train.pretrained_model_path` | `ptm_if_no_resume_model` | PTM when no resume checkpoint exists |
| train | `train.resume_training_checkpoint_path` | `resume_model` | model file inferred from the current job results folder |

For `parent_model` or `parent_model_folder`, pass the upstream train/export/AutoML child job id as `parent_job_id`. The SDK lists the parent result folder, filters checkpoint artifacts, and returns the selected model file or folder. Do not add these mappings back to `config.json` and do not patch generated runner scripts to guess checkpoint paths.

When selecting a Grounding-DINO checkpoint outside the SDK resolver, match the
intended epoch/step artifact exactly, for example
`model_epoch_000_step_00046.pth`. The `gdino_model_latest.pth` symlink is valid
only when latest is explicitly requested. Carry structural model settings such
as `model.backbone`, `model.num_queries`, `model.num_select`,
`model.num_feature_levels`, `model.max_text_len`, and export input resolution
forward into evaluate, inference, export, and deploy specs so checkpoint and
engine shapes match.

## Deployment

- [tao-deploy-grounding-dino](references/tao-deploy-grounding-dino.md)

安装 tao-train-grounding-dino

下载技能文件并将其解压到 .claude/skills/ 目录中。

下载ZIP

克隆仓库并复制技能文件到您的项目中。

git clone https://github.com/NVIDIA/skills/tree/main/skills/tao-train-grounding-dino # Copy SKILL.md to your .claude/skills/ directory

复制 复制
快速设置: 将技能文件夹复制到 .claude/skills/ Claude 会自动检测并使用该技能
仓库 NVIDIA/skills

相关技能

web-search
更新时间 2026-06-29
webapp-testing
更新时间 2026-06-29
lark-base
更新时间 2026-07-05
agentmail
更新时间 2026-06-29
OR