选项
首页首页 Skill 开发运营和 CI/CD vss-deploy-dense-captioning

vss-deploy-dense-captioning

NVIDIA/skills NVIDIA/skills

部署一个独立的 RT-VLM 高密度字幕生成微服务,并测试其用于文件上传、字幕生成、流媒体、聊天自动补全以及 Kafka 集成的 REST API 端点。

...展开全部
2
更新时间 2026-09-27

目的

将 RT-VLM 高密度字幕生成微服务独立部署,并测试其暴露的每个端点(文件上传、生成字幕、流添加/删除、聊天自动补全、Kafka 主题)。

先决条件

对于 RT-VLM 的独立部署:

  • Docker、Docker Compose、NVIDIA Container Toolkit 以及一台可访问的 GPU。
  • $NGC_CLI_API_KEY中需包含 NGC 注册表凭据,用于通过docker login nvcr.io 登录、 拉取镜像以及下载本地 NGC 模型/构建产物。
  • curl、jq 以及用于独立部署 Compose 副本的任何可写工作目录。

针对现有服务的 API 调用:

  • 正在运行的 RT-VLM 服务,可通过$BASE_URL 访问。
  • $RTVI_VLM_API_KEY或$NGC_CLI_API_KEY 中的 Bearer 令牌,具体取决于 服务的配置方式。

对于完整的 VSS 配置文件部署:

  • 请使用../vss-deploy-profile/SKILL.md;此技能不支持部署完整的 VSS 配置文件。

操作指南

请遵循下方的路由表和分步工作流。每个以“工作流”、“快速入门”或“流程”结尾的章节均应自上而下依次执行。详细参考资料位于references/ 目录中;除非未来修订版指定了具体的辅助工具,否则请直接执行文档中记录的工作流。

示例

已验证的端到端示例保存在evals/目录下(每个*.json清单文件包含一个可运行的场景),并内嵌在下方各工作流的curl代码块中。运行nv-base validate --agent-eval进行三级评估,即可重现这些示例。

限制

  • 需要通过此技能部署的独立 RT-VLM 服务,或 调用方可访问的现有 RT-VLM 服务。
  • NGC 托管的模型和 NIM 可能受速率限制、GPU 内存要求以及许可限制的约束。
  • 并发数、GPU 内存和存储限制取决于主机硬件以及配置文件的 compose 文件。
  • 请将NGC_CLI_API_KEY、RTVI_VLM_API_KEY 和.env文件排除在 git 之外,并避免将其记录在日志中;请勿回显凭据值,也不要将其包含在最终响应中。
  • Docker 组访问权限和sudo实质上属于 root 级权限。当无法使用无密码 sudo 时,请在部署参考中使用非交互式sudo -n保护机制,并在需要主机所有者操作时停止使用。

故障排除

  • 错误:REST 调用返回“连接被拒绝”。原因:目标微服务未运行。解决方案:探测/docs或/health;通过vss-deploy-profile或相应的vss-deploy-*技能重新部署。
  • 错误:从 NGC 拉取时出现 HTTP 401/403 错误。原因:缺少或过期的NGC_CLI_API_KEY。解决方案:执行docker login nvcr.io并重新导出密钥后再重试。
  • 错误:容器内存不足(OOM)或模型加载失败。原因:所选配置文件所需的 GPU 内存不足。解决方案:切换至更小的配置,或通过 `docker compose down` 释放 GPU 资源。

部署和使用 RT-VLM 密集式字幕生成(VSS 3.2)

RT-VLM 是 NVIDIA 的实时视觉语言微服务:解码视频(文件或 RTSP),将其分割为片段,运行 VLM(cosmos-reason1、cosmos-reason2 或任何 OpenAI兼容模型),通过SSE/HTTP将密集式字幕流回传,并将 字幕、事件警报和错误发布到Kafka。当尚未运行完整的VSS配置文件时,请使用此技能部署 独立的RT-VLM服务,随后调用 其/v1/...API 进行字幕生成、文件上传、直播管理、健康 检查、NIM 兼容的聊天补全或 Prometheus 指标操作。API 参考: https://docs.nvidia.com/vss/latest/real-time-vlm-api.html。

部署路由

如果用户请求部署完整的 VSS 配置文件,请使用 ../vss-deploy-profile/SKILL.md。该技能 负责配置文件路由、generated.env、resolved.yml、多服务规模配置以及 全栈部署/拆解。

如果用户请求独立的 RT-VLM 高密度字幕生成,或者当前没有 VSS 配置文件 正在运行,请在调用 API 之前使用 references/deploy-rt-vlm-service.md 中的独立 RT-VLM 流程。 这遵循了与 vss-deploy-profile 相同的以 Compose 为中心的模式:收集上下文、运行预检查、基于本地副本操作、 使用Docker Compose 配置进行干跑、审查、部署,然后等待健康状态确认。

独立部署流程

请始终遵循此顺序。切勿跳过干跑步骤。

# 1. 将 deploy/docker/services/rtvi/rtvi-vlm/rtvi-vlm-docker-compose.yml
#    复制到任何可写入的独立工作目录中。
# 2. 从该 Compose 副本中推导出 RTVI_VLM_IMAGE_TAG。
# 3. 从该副本中移除仅适用于独立部署的悬空 depends_on 代码块。
# 4. 创建一个被 git 忽略的 .env 文件,其中包含所需的 RT-VLM 值。
# 5. 准备主机绑定路径,例如 $VSS_DATA_DIR/data_log/vst/clip_storage。
#    使用 `sudo -n` 修复所有权问题;如果无法使用无密码 sudo,
#    请暂停操作,并请主机所有者手动运行打印出的命令。
# 6. docker compose --env-file .env -f rtvi-vlm-docker-compose.yml config --quiet
# 7. 使用 `docker pull` 拉取确切的 RT-VLM 镜像标签。
# 8. 执行 `docker compose ... up -d rtvi-vlm`,等待就绪后进行初步测试。

在执行任何 `pull` 或 `up` 操作之前,请先运行预检查;在此处停止并修复故障, 然后再调试 RT-VLM 本身:

nvidia-smi --query-gpu=index,name --format=csv,noheader
nvidia-container-cli info
docker compose version
docker run --rm --gpus all nvidia/cuda:12.4.0-base-ubuntu22.04 nvidia-smi

对于独立的单文件部署,请勿直接运行原始的 deploy/docker/services/rtvi/rtvi-vlm/rtvi-vlm-docker-compose.yml:该文件 包含对同级 VLM/NIM 服务的depends_on引用,这些引用仅 在完整的 VSS/met-blueprints Compose 项目中定义。独立部署示例 展示了如何复制 Compose 文件、从中推导出当前镜像标签、移除 `depends_on` 块,并在启动前验证结果。

对于代理驱动的验证,切勿让sudo提示符进入交互模式。在执行任何 特权所有权操作或 Docker 操作之前,请使用references/deploy-rt-vlm-service.md 中的非交互式保护机制: 优先使用普通docker;否则使用sudo -n docker;若sudo -n失败,请 使用主机所有者的精确手动命令终止操作,而非通过 交互式 sudo 或降低权限来重试。

如果 Docker 28 及以上版本中docker pull因 containerd snapshotter/unpack 错误失败, 请在重试前,按照独立部署参考文档中的说明,在/etc/docker/daemon.json 中应用containerd-snapshotter=false的修复方案。

独立部署的.env最小配置值:

主机环境变量 何时需要 用途
NGC_CLI_API_KEY 独立部署路径 从 NGC 注册表拉取镜像以及下载 NGC 模型/工件
RTVI_VLM_API_KEY或NGC_CLI_API_KEY 经过身份验证的 API 调用 服务运行后的 RT-VLM 承载认证
RTVI_VLM_PORT 始终 映射到容器8000的主机 API 端口
HOST_IP 始终 Kafka 引导主机 (${HOST_IP}:9092)
VSS_DATA_DIR 始终 必需的 clip-storage 绑定挂载
RTVI_VLM_MODEL_TO_USE 独立部署时始终启用 后端选择器;使用cosmos-reason2作为默认本地模型,或使用openai-compat作为远程/同级端点
RTVI_VLM_MODEL_PATH 本地自托管模型 基于源代码的 Cosmos Reason 2 路径:ngc:nim/nvidia/cosmos-reason2-8b:hf-1208
RTVI_VLM_ENDPOINT RTVI_VLM_MODEL_TO_USE=openai-compat 远程/同级 OpenAI 兼容 VLM 端点
VLM_NAME RTVI_VLM_MODEL_TO_USE=openai-compat 该端点公开的模型/部署名称

配置

export BASE_URL="http://localhost:${RTVI_VLM_PORT:-8018}"  # 主机端 RT-VLM 端口
export API_KEY="${NGC_CLI_API_KEY:-${RTVI_VLM_API_KEY:-}}" # 主机端 curl 命令使用的承载令牌
: "${API_KEY:?在调用需身份验证的端点前,请先设置 NGC_CLI_API_KEY 或 RTVI_VLM_API_KEY}"

以下每个请求均使用Authorization: Bearer $API_KEY。健康检查端点 (/v1/health/*,/v1/ready,/v1/live,/v1/startup) 通常无需身份验证即可正常工作。

使用前的快速测试:

curl -fsS "$BASE_URL/v1/health/ready"
MODEL_ID="$(curl -fsS "$BASE_URL/v1/models" -H "Authorization: Bearer $API_KEY" | jq -r '.data[0].id // .id')"
curl -fsS "$BASE_URL/openapi.json" | jq -r '.paths | keys[]' | sort

RTSP 示例流守护程序

当任务或评估中提及RTSP_SAMPLE_URL 时,请将该精确的环境 变量视为必需输入。 在探测或 注册任何流之前,请验证该变量是否已设置且不为空;若缺失,则显示明确的失败信息并终止。请 勿从 NvStreamer、VIOS、样本数据包或任何其他 备用方案中推导替代值,因为这会验证与调用方请求不同的流。

: "${RTSP_SAMPLE_URL:?在进行 RTSP 验证前,请将 RTSP_SAMPLE_URL 设置为可访问的 RTSP 样本流}"
case "$RTSP_SAMPLE_URL" in
  rtsp://*) ;;
  *) echo "RTSP_SAMPLE_URL 必须是 rtsp:// 格式的 URL,获取到:$RTSP_SAMPLE_URL" >&2; exit 1 ;;
esac

if command -v ffprobe >/dev/null 2>&1; then
  ffprobe -v error -rtsp_transport tcp \
    -select_streams v:0 -show_entries stream=codec_type \
    -of csv=p=0 "$RTSP_SAMPLE_URL" | grep -qx video
elif command -v gst-discoverer-1.0 >/dev/null 2>&1; then
  gst-discoverer-1.0 "$RTSP_SAMPLE_URL" | grep -qi 'video'
else
  echo "请在进行 RTSP 验证前安装 ffprobe 或 gst-discoverer-1.0。" >&2
  exit 1
fi

快速入门 — 从本地视频生成密集式字幕

# 1. 上传视频,获取其文件 ID
FILE_ID=$(curl -fsS -X POST "$BASE_URL/v1/files" \
  -H "Authorization: Bearer $API_KEY" \
  -F "file=@/path/to/warehouse.mp4" \
  -F "purpose=vision" \
  -F "media_type=video" | jq -r '.id')

# 2. 生成字幕和警报(分块响应的 SSE 流)
curl -N -X POST "$BASE_URL/v1/generate_captions" \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d "{
    \"id\": \"$FILE_ID\",
    \"prompt\": \"为该仓库视频的每个 10 秒片段编写简明扼要的字幕。\",
    \"model\": \"$MODEL_ID\",
    \"chunk_duration\": 10,
    \"stream\": true
  }"

API 接口

在调用可选端点之前,请以实时 OpenAPI 作为权威来源:

curl -fsS "$BASE_URL/openapi.json" | jq -r '.paths | keys[]' | sort

VSS 3.2 的核心路径包括:

  • POST /v1/files用于多部分媒体上传;将返回的文件ID传递给 字幕生成功能,并在完成后删除该文件。
  • POST /v1/generate_captions用于文件或流的字幕生成。请使用 GET /v1/models 返回的确切模型 ID;诸如cosmos-reason2之类的别名是 后端选择器,而非请求模型 ID。
  • POST /v1/streams/add、GET /v1/streams/get-stream-info 以及 DELETE /v1/streams/delete/{stream_id}用于 RTSP 生命周期管理。从results[0].id 中解析流 ID。
  • POST /v1/chat/completions用于 OpenAI 兼容的文本和多模态调用。 当前 26.05 版本对纯文本的/v1/completions 请求返回 HTTP 400 状态码;在验证旧版行为时, 请将此情况视为预期行为。
  • GET /v1/health/ready、/v1/models、/v1/assets/stats 和/v1/metrics 用于服务探测。除非 OpenAPI 中列出,否则不要假设/v1/license存在。

详细的端点模式、响应结构、CV 风格的单流端点 以及 26.05 兼容性说明位于 references/api-surface-26.05.md 中。

常见工作流

  • 存储文件字幕:使用POST /v1/files 上传,调用 /v1/generate_captions并传入返回的文件 ID,使用stream=true启用 SSE, 随后删除文件以释放存储空间。
  • RTSP 实时字幕生成:当调用方提供RTSP_SAMPLE_URL 时,请使用该 确切 URL,并在注册前运行RTSP 样本流守护程序。 当RTSP_SAMPLE_URL为 空时,请勿 从 NvStreamer 或 VIOS 派生替代流;应立即返回失败。注册前必须 存在实际的视频流/字幕条目;添加流、为其生成字幕,然后注销该流。
  • 警报提示:包含一条确定性的“检测到异常:是/否”行。 Kafka 发布是服务器端的配置,作为 HTTP 响应的补充, 相关文档详见references/kafka-workflows.md。
  • Kafka 验证:主题名称应信任vss-rtvi-vlm生产环境。 在完整的 VSS 警报实时配置中,使用现有的 VSS Kafka 容器 mdx-kafka用于 CLI 检查和最终的事件消费者命令。对于 独立验证,请使用发布${HOST_IP}:9092 的代理;切勿 在未经用户确认的情况下停止或替换现有的代理。

错误参考

常见原因:400 表示请求格式或模型 ID 无效,401/403 表示 缺失或错误的 Bearer 令牌,404 表示文件/流已被删除或端点不受支持, 413 表示上传文件过大,422 表示模式验证失败,429 表示 并发过高,500 表示推理/运行时失败,503 表示启动仍在 进行中。请检查Docker 日志 vss-rtvi-vlm以排查服务端故障。

在 GitHub 上查看
---
name: vss-deploy-dense-captioning
description: Deploy a standalone RT-VLM dense-captioning microservice and exercise its REST API endpoints for file upload, caption generation, streaming, chat completions, and Kafka integration.
license: Apache-2.0
---
## Purpose

Stand up the RT-VLM dense-captioning microservice on its own and exercise every endpoint it exposes (file upload, generate_captions, stream add/delete, chat-completions, Kafka topics).

## Prerequisites

For standalone RT-VLM deployment:
- Docker, Docker Compose, NVIDIA Container Toolkit, and a visible GPU.
- NGC registry credentials in `$NGC_CLI_API_KEY` for `docker login nvcr.io`,
  image pulls, and local NGC model/artifact downloads.
- `curl`, `jq`, and any writable working directory for the standalone compose copy.

For API calls against an existing service:
- Running RT-VLM service reachable at `$BASE_URL`.
- Bearer token in `$RTVI_VLM_API_KEY` or `$NGC_CLI_API_KEY`, depending on how the
  service was configured.

For full VSS profile deployment:
- Use `../vss-deploy-profile/SKILL.md`; this skill does not deploy full VSS profiles.

## Instructions

Follow the routing tables and step-by-step workflows below. Each section that ends in *workflow*, *quick start*, or *flow* is intended to be executed top-to-bottom. Detailed reference material lives in `references/`; execute the documented workflows directly unless a future revision names a concrete helper.

## Examples

Worked end-to-end examples are kept under `evals/` (each `*.json` manifest contains a runnable scenario) and inline in the per-workflow `curl` blocks below. Run a Tier-3 evaluation with `nv-base validate <this-skill-dir> --agent-eval` to replay them.

## Limitations

- Requires either a standalone RT-VLM service deployed via this skill or an
  existing RT-VLM service reachable from the caller.
- NGC-hosted models and NIMs may be subject to rate-limits, GPU memory requirements, and license restrictions.
- Concurrency, GPU memory, and storage limits depend on the host hardware and the profile's compose file.
- Keep `NGC_CLI_API_KEY`, `RTVI_VLM_API_KEY`, and `.env` files out of git and out of logs; do not echo credential values or include them in final responses.
- Docker group access and `sudo` are effectively root-level privileges. Use the non-interactive `sudo -n` guard in the deploy reference and stop for host-owner action when passwordless sudo is unavailable.

## Troubleshooting

- **Error**: REST call returns connection refused. **Cause**: target microservice not running. **Solution**: probe `/docs` or `/health`; redeploy via `vss-deploy-profile` or the matching `vss-deploy-*` skill.
- **Error**: HTTP 401/403 from NGC pulls. **Cause**: missing/expired `NGC_CLI_API_KEY`. **Solution**: `docker login nvcr.io` and re-export the key before retrying.
- **Error**: container OOM or model fails to load. **Cause**: insufficient GPU memory for the selected profile. **Solution**: switch to a smaller variant or free GPUs via `docker compose down`.

# Deploy and Use RT-VLM Dense Captioning (VSS 3.2)

RT-VLM is NVIDIA's real-time vision-language microservice: decode video (file or
RTSP), segment it into chunks, run a VLM (`cosmos-reason1`, `cosmos-reason2`, or any
OpenAI-compatible model), stream dense captions back over SSE/HTTP, and publish
captions, incident alerts, and errors to Kafka. Use this skill to deploy the
standalone RT-VLM service when a full VSS profile is not already running, then call
its `/v1/...` API for caption generation, file upload, live-stream management, health
checks, NIM-compatible chat completions, or Prometheus metrics. API reference:
<https://docs.nvidia.com/vss/latest/real-time-vlm-api.html>.

## Deployment Routing

If the user asks to deploy a full VSS profile, use
[`../vss-deploy-profile/SKILL.md`](../vss-deploy-profile/SKILL.md). That skill
owns profile routing, `generated.env`, `resolved.yml`, multi-service sizing, and
full-stack deploy/teardown.

If the user asks for standalone RT-VLM dense captioning, or no VSS profile is
already running, use the standalone RT-VLM flow in
[`references/deploy-rt-vlm-service.md`](references/deploy-rt-vlm-service.md)
before calling the API. This follows the same compose-centric pattern as
`vss-deploy-profile`: gather context, run preflights, work from a local copy,
dry-run with `docker compose config`, review, deploy, then wait for health.

## Standalone Deployment Flow

Always follow this sequence. Never skip the dry-run.

```bash
# 1. Copy deploy/docker/services/rtvi/rtvi-vlm/rtvi-vlm-docker-compose.yml
#    into any writable standalone working directory.
# 2. Derive RTVI_VLM_IMAGE_TAG from that compose copy.
# 3. Strip the standalone-only dangling depends_on block from the copy.
# 4. Create a gitignored .env with the required RT-VLM values.
# 5. Prepare host bind paths such as $VSS_DATA_DIR/data_log/vst/clip_storage.
#    Use `sudo -n` for ownership fixes; if passwordless sudo is unavailable,
#    stop and ask the host owner to run the printed command manually.
# 6. docker compose --env-file .env -f rtvi-vlm-docker-compose.yml config --quiet
# 7. docker pull the exact RT-VLM image tag.
# 8. docker compose ... up -d rtvi-vlm, wait for ready, then smoke test.
```

Run preflights before any pull or `up`; stop and fix failures here before
debugging RT-VLM itself:

```bash
nvidia-smi --query-gpu=index,name --format=csv,noheader
nvidia-container-cli info
docker compose version
docker run --rm --gpus all nvidia/cuda:12.4.0-base-ubuntu22.04 nvidia-smi
```

For standalone single-file deployments, do not run the raw
`deploy/docker/services/rtvi/rtvi-vlm/rtvi-vlm-docker-compose.yml` directly: it
contains `depends_on` references to sibling VLM/NIM services that are only
defined in the full VSS/met-blueprints compose project. The standalone reference
shows how to copy the compose file, derive the current image tag from it, strip
the `depends_on` block, and validate the result before `up`.

For agent-driven validation, never let `sudo` prompt interactively. Before any
privileged ownership or Docker operation, use the non-interactive guard in
[`references/deploy-rt-vlm-service.md`](references/deploy-rt-vlm-service.md):
prefer plain `docker`; otherwise use `sudo -n docker`; if `sudo -n` fails, stop
with the exact manual command for the host owner instead of retrying with
interactive sudo or weakening permissions.

If `docker pull` fails with a containerd snapshotter/unpack error on Docker 28+,
apply the `/etc/docker/daemon.json` `containerd-snapshotter=false` fix in the
standalone reference before retrying.

Minimum standalone `.env` values:

| Host env var | Required when | Purpose |
|---|---|---|
| `NGC_CLI_API_KEY` | Standalone deploy path | NGC registry image pull and NGC model/artifact download |
| `RTVI_VLM_API_KEY` or `NGC_CLI_API_KEY` | Authenticated API calls | RT-VLM bearer auth after the service is running |
| `RTVI_VLM_PORT` | Always | Host API port mapped to container `8000` |
| `HOST_IP` | Always | Kafka bootstrap host (`${HOST_IP}:9092`) |
| `VSS_DATA_DIR` | Always | Required clip-storage bind mount |
| `RTVI_VLM_MODEL_TO_USE` | Always for standalone | Backend selector; use `cosmos-reason2` for the default local model or `openai-compat` for a remote/sibling endpoint |
| `RTVI_VLM_MODEL_PATH` | Local self-hosted model | Source-backed Cosmos Reason 2 path: `ngc:nim/nvidia/cosmos-reason2-8b:hf-1208` |
| `RTVI_VLM_ENDPOINT` | `RTVI_VLM_MODEL_TO_USE=openai-compat` | Remote/sibling OpenAI-compatible VLM endpoint |
| `VLM_NAME` | `RTVI_VLM_MODEL_TO_USE=openai-compat` | Model/deployment name exposed by that endpoint |

## Setup

```bash
export BASE_URL="http://localhost:${RTVI_VLM_PORT:-8018}"  # host-side RT-VLM port
export API_KEY="${NGC_CLI_API_KEY:-${RTVI_VLM_API_KEY:-}}" # bearer token used by host-side curl commands
: "${API_KEY:?Set NGC_CLI_API_KEY or RTVI_VLM_API_KEY before calling authenticated endpoints}"
```

Every request below uses `Authorization: Bearer $API_KEY`. Health endpoints
(`/v1/health/*`, `/v1/ready`, `/v1/live`, `/v1/startup`) typically work without auth.

**Smoke test before use:**
```bash
curl -fsS "$BASE_URL/v1/health/ready"
MODEL_ID="$(curl -fsS "$BASE_URL/v1/models" -H "Authorization: Bearer $API_KEY" | jq -r '.data[0].id // .id')"
curl -fsS "$BASE_URL/openapi.json" | jq -r '.paths | keys[]' | sort
```

## RTSP Sample Stream Guard

When a task or eval names `RTSP_SAMPLE_URL`, treat that exact environment
variable as a required input. Verify it is set and non-empty before probing or
registering any stream; if it is missing, stop with a clear failure message. Do
not derive a substitute from NvStreamer, VIOS, sample-data bundles, or any other
fallback, because that validates a different stream than the caller requested.

```bash
: "${RTSP_SAMPLE_URL:?Set RTSP_SAMPLE_URL to a reachable RTSP sample stream before RTSP validation}"
case "$RTSP_SAMPLE_URL" in
  rtsp://*) ;;
  *) echo "RTSP_SAMPLE_URL must be an rtsp:// URL, got: $RTSP_SAMPLE_URL" >&2; exit 1 ;;
esac

if command -v ffprobe >/dev/null 2>&1; then
  ffprobe -v error -rtsp_transport tcp \
    -select_streams v:0 -show_entries stream=codec_type \
    -of csv=p=0 "$RTSP_SAMPLE_URL" | grep -qx video
elif command -v gst-discoverer-1.0 >/dev/null 2>&1; then
  gst-discoverer-1.0 "$RTSP_SAMPLE_URL" | grep -qi 'video'
else
  echo "Install ffprobe or gst-discoverer-1.0 before RTSP validation." >&2
  exit 1
fi
```

## Quick Start — dense captions from a local video

```bash
# 1. Upload the video, capture its file id
FILE_ID=$(curl -fsS -X POST "$BASE_URL/v1/files" \
  -H "Authorization: Bearer $API_KEY" \
  -F "file=@/path/to/warehouse.mp4" \
  -F "purpose=vision" \
  -F "media_type=video" | jq -r '.id')

# 2. Generate captions + alerts (SSE stream of chunked responses)
curl -N -X POST "$BASE_URL/v1/generate_captions" \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d "{
    \"id\": \"$FILE_ID\",
    \"prompt\": \"Write a concise dense caption for each 10-second segment of this warehouse video.\",
    \"model\": \"$MODEL_ID\",
    \"chunk_duration\": 10,
    \"stream\": true
  }"
```

## API Surface

Use the live OpenAPI as the source of truth before calling optional endpoints:

```bash
curl -fsS "$BASE_URL/openapi.json" | jq -r '.paths | keys[]' | sort
```

Core paths for VSS 3.2 are:

- `POST /v1/files` for multipart media upload; pass the returned file `id` into
  caption generation and delete the file when finished.
- `POST /v1/generate_captions` for file or stream captioning. Use the exact
  model id returned by `GET /v1/models`; aliases such as `cosmos-reason2` are
  backend selectors, not request model ids.
- `POST /v1/streams/add`, `GET /v1/streams/get-stream-info`, and
  `DELETE /v1/streams/delete/{stream_id}` for RTSP lifecycle. Parse stream ids
  from `results[0].id`.
- `POST /v1/chat/completions` for OpenAI-compatible text and multimodal calls.
  Current 26.05 builds return HTTP 400 for text-only `/v1/completions`; treat
  that as expected when validating legacy behavior.
- `GET /v1/health/ready`, `/v1/models`, `/v1/assets/stats`, and `/v1/metrics`
  for service probes. Do not assume `/v1/license` exists unless OpenAPI lists it.

Detailed endpoint schemas, response shapes, CV-style singular stream endpoints,
and 26.05 compatibility notes live in
[`references/api-surface-26.05.md`](references/api-surface-26.05.md).

## Common Workflows

- Stored file captioning: upload with `POST /v1/files`, call
  `/v1/generate_captions` with the returned file id, use `stream=true` for SSE,
  then delete the file to release storage.
- RTSP live captioning: when the caller provides `RTSP_SAMPLE_URL`, use that
  exact URL and run the **RTSP Sample Stream Guard** before registration. Do not
  derive a replacement stream from NvStreamer or VIOS when `RTSP_SAMPLE_URL` is
  empty; fail fast instead. Require an actual video stream/caps entry before
  registration; add the stream, caption it, then unregister it.
- Alert prompts: include a deterministic `Anomaly Detected: Yes/No` line.
  Kafka publication is server-side config, additive to HTTP responses, and
  documented in [`references/kafka-workflows.md`](references/kafka-workflows.md).
- Kafka validation: trust the live `vss-rtvi-vlm` environment for topic names.
  In a full VSS alerts real-time profile, use the existing VSS Kafka container
  `mdx-kafka` for CLI checks and final incident-consumer commands. For
  standalone validation, use a broker that advertises `${HOST_IP}:9092`; never
  stop or replace a pre-existing broker without user confirmation.

## Error Reference

Common causes: 400 for invalid request shape or model id, 401/403 for missing
or wrong bearer token, 404 for deleted files/streams or unsupported endpoints,
413 for oversized uploads, 422 for schema validation, 429 for too much
concurrency, 500 for inference/runtime failures, and 503 while startup is still
in progress. Inspect `docker logs vss-rtvi-vlm` for service-side failures.

安装 vss-deploy-dense-captioning

下载技能文件并将其解压到 .claude/skills/ 目录中。

下载ZIP

克隆仓库并复制技能文件到您的项目中。

git clone https://github.com/NVIDIA/skills/tree/main/skills/vss-deploy-dense-captioning # Copy SKILL.md to your .claude/skills/ directory

复制 复制
快速设置: 将技能文件夹复制到 .claude/skills/ Claude 将自动检测并使用该技能
仓库 NVIDIA/skills

相关技能

klingai-upgrade-migration
更新时间 2026-07-03
Verification &amp; Quality Assurance
更新时间 2026-06-29
base44-cli
更新时间 2026-06-29
Railway CLI Management
更新时间 2026-07-02
OR