選項
首頁首頁 Skill 開發營運和 CI/CD vss-deploy-dense-captioning

vss-deploy-dense-captioning

NVIDIA/skills NVIDIA/skills

部署一個獨立運作的 RT-VLM 密集式字幕生成微服務,並測試其 REST API 端點,包括檔案上傳、字幕生成、串流、聊天自動完成功能,以及與 Kafka 的整合。

...展開全部
2
更新時間 2026-09-27

目的

將 RT-VLM 高密度字幕生成微服務獨立部署,並測試其所公開的每個端點(檔案上傳、generate_captions、串流新增/刪除、聊天自動完成、Kafka 主題)。

先決條件

針對 RT-VLM 獨立部署:

  • 需具備 Docker、Docker Compose、NVIDIA Container Toolkit 以及一顆可用的 GPU。
  • $NGC_CLI_API_KEY中須包含 NGC 註冊表憑證,以便透過 `docker login nvcr.io` 登入、 拉取映像檔,以及下載本機 NGC 模型/構件。
  • curl、jq,以及任何可供獨立部署的 Compose 副本寫入的工作目錄。

針對現有服務進行 API 呼叫時:

  • 可透過$BASE_URL 存取的 RT-VLM 服務。
  • $RTVI_VLM_API_KEY或$NGC_CLI_API_KEY 中的 Bearer 憑證,具體取決於 服務的配置方式。

若要部署完整的 VSS 設定檔:

  • 請使用../vss-deploy-profile/SKILL.md;此技能不會部署完整的 VSS 設定檔。

操作說明

請遵循以下路由表及逐步工作流程。每個以「工作流程」、「快速入門」或「流程」結尾的章節,皆應自上而下依序執行。詳細參考資料存放於references/ 目錄中;除非未來修訂版指定了具體的輔助程式,否則請直接執行文件中記載的工作流程。

範例

已驗證的端到端範例存放於evals/目錄下(每個*.json清單皆包含可執行的情境),並內嵌於下方的各工作流程curl區塊中。請執行 Tier-3 評估(使用nv-base validate --agent-eval)以重現這些範例。

限制

  • 需具備透過此技能部署的獨立 RT-VLM 服務,或 呼叫方可存取的現有 RT-VLM 服務。
  • 由 NGC 託管的模型和 NIM 可能會受到速率限制、GPU 記憶體需求以及授權限制的影響。
  • 並發數、GPU 記憶體及儲存空間的限制取決於主機硬體及配置檔的 compose 檔案。
  • 請將NGC_CLI_API_KEY、RTVI_VLM_API_KEY 及.env檔案排除在 git 及日誌之外;切勿回顯憑證值,亦勿將其包含在最終回應中。
  • Docker 群組存取權限與sudo實質上屬於 root 級別的權限。當無法使用無密碼 sudo 時,請在部署參考中使用非互動式sudo -n防護機制,並將「停止」操作保留給主機擁有者執行。

疑難排解

  • 錯誤:REST 呼叫傳回「連線遭拒絕」。原因:目標微服務未運行。解決方案:偵測/docs或/health;透過vss-deploy-profile或相應的vss-deploy-*技能重新部署。
  • 錯誤:NGC 拉取時出現 HTTP 401/403 錯誤。原因:NGC_CLI_API_KEY 遺失或已過期。解決方案:執行 `docker login nvcr.io` 並重新匯出金鑰後再嘗試。
  • 錯誤:容器發生 OOM 或模型無法載入。原因:所選設定檔的 GPU 記憶體不足。解決方案:切換至較小的變體,或透過 `docker compose down` 釋放 GPU 資源。

部署與使用 RT-VLM 密集式圖說生成(VSS 3.2)

RT-VLM 是 NVIDIA 的即時視覺語言微服務:解碼影片(檔案或 RTSP),將其分割成區塊,執行 VLM(cosmos-reason1、cosmos-reason2 或任何 與 OpenAI 相容的模型),透過 SSE/HTTP 串流傳回密集式字幕,並將 字幕、事件警示及錯誤發佈至 Kafka。當尚未運行完整的 VSS 配置檔時,請使用此技能部署 獨立的 RT-VLM 服務,然後呼叫 其/v1/...API 來進行字幕生成、檔案上傳、直播管理、健康 檢查、與 NIM 相容的聊天自動完成功能,或 Prometheus 指標。API 參考: https://docs.nvidia.com/vss/latest/real-time-vlm-api.html。

部署路由

若使用者要求部署完整的 VSS 配置檔,請使用 ../vss-deploy-profile/SKILL.md。該技能 負責配置檔路由、generated.env、resolved.yml、多服務規模配置,以及 全堆疊部署/拆除。

若使用者要求獨立運作的 RT-VLM 高密度字幕生成,或目前尚無 VSS 配置檔 正在執行,請在呼叫 API 之前,使用 references/deploy-rt-vlm-service.md 中的獨立 RT-VLM 流程。 此流程遵循與 vss-deploy-profile 相同的 Compose 為中心的模式:收集上下文、執行預檢查、基於本地副本操作、 使用Docker Compose 配置進行模擬執行、審查、部署,然後等待狀態確認。

獨立部署流程

請務必遵循此順序。切勿跳過模擬執行步驟。

# 1. 將 deploy/docker/services/rtvi/rtvi-vlm/rtvi-vlm-docker-compose.yml
#    複製到任何可寫入的獨立運作工作目錄中。
# 2. 從該 Compose 副本中推導出 RTVI_VLM_IMAGE_TAG。
# 3. 從該複製檔中移除僅適用於獨立部署的冗餘 depends_on 區塊。
# 4. 建立一個被 git 忽略的 .env 檔案,並填入所需的 RT-VLM 值。
# 5. 準備主機的綁定路徑,例如 $VSS_DATA_DIR/data_log/vst/clip_storage。
#    使用 `sudo -n` 修正所有權;若無法使用無密碼 sudo,
#    請暫停並請主機擁有者手動執行列印出的指令。
# 6. docker compose --env-file .env -f rtvi-vlm-docker-compose.yml config --quiet
# 7. 使用 `docker pull` 拉取確切的 RT-VLM 映像標籤。
# 8. 執行 `docker compose ... up -d rtvi-vlm`,等待狀態轉為「ready」,然後進行初步測試。

在執行任何 `pull` 或`up` 指令前,請先執行預檢;在此階段先停止並修正錯誤, 再進行 RT-VLM 本身的除錯:

nvidia-smi --query-gpu=index,name --format=csv,noheader
nvidia-container-cli info
docker compose version
docker run --rm --gpus all nvidia/cuda:12.4.0-base-ubuntu22.04 nvidia-smi

對於獨立的單一檔案部署,請勿直接執行原始的 deploy/docker/services/rtvi/rtvi-vlm/rtvi-vlm-docker-compose.yml:該檔案 包含對同級 VLM/NIM 服務的depends_on引用,這些服務僅 僅在完整的 VSS/met-blueprints Compose 專案中定義。此獨立部署範例 說明了如何複製 Compose 檔案、據此推導當前映像標籤、移除 `depends_on` 區塊,並在啟動前驗證結果。

對於代理驅動的驗證,切勿讓sudo提示框進入互動模式。在執行任何 特權擁有權設定或 Docker 操作之前,請使用references/deploy-rt-vlm-service.md 中 的非互動式防護機制: 優先使用純文字模式的docker;否則使用sudo -n docker;若sudo -n失敗,請直接 使用主機擁有者的精確手動命令終止操作,而非嘗試使用 互動式 sudo 或降低權限。

若在 Docker 28 以上版本中,docker pull因 containerd snapshotter/unpack 錯誤而失敗, 請在重新嘗試前,套用獨立部署參考文件中/etc/docker/daemon.json的 containerd-snapshotter=false修正方案。

獨立執行環境的.env最小設定值:

主機環境變數 何時需要 用途
NGC_CLI_API_KEY 獨立部署路徑 NGC 註冊表映像拉取及 NGC 模型/產出檔案下載
RTVI_VLM_API_KEY或NGC_CLI_API_KEY 經身份驗證的 API 呼叫 服務運行後的 RT-VLM 承載者驗證
RTVI_VLM_PORT 始終 主機 API 埠映射至容器8000
HOST_IP 始終 Kafka 啟動主機 (${HOST_IP}:9092)
VSS_DATA_DIR 始終 必需的 clip-storage 綁定掛載
RTVI_VLM_MODEL_TO_USE 獨立執行時始終啟用 後端選擇器;若要使用預設的本地模型,請使用cosmos-reason2;若要使用遠端或同級端點,請使用openai-compat
RTVI_VLM_MODEL_PATH 本地自託管模型 基於原始碼的 Cosmos Reason 2 路徑:ngc:nim/nvidia/cosmos-reason2-8b:hf-1208
RTVI_VLM_ENDPOINT RTVI_VLM_MODEL_TO_USE=openai-compat 遠端/同級 OpenAI 相容 VLM 端點
VLM_NAME RTVI_VLM_MODEL_TO_USE=openai-compat 該端點所公開的模型/部署名稱

設定

export BASE_URL="http://localhost:${RTVI_VLM_PORT:-8018}"  # 主機端 RT-VLM 連接埠
export API_KEY="${NGC_CLI_API_KEY:-${RTVI_VLM_API_KEY:-}}" # 主機端 curl 指令使用的 Bearer 憑證
: "${API_KEY:?在呼叫需驗證的端點前,請先設定 NGC_CLI_API_KEY 或 RTVI_VLM_API_KEY}"

以下每個請求均使用Authorization: Bearer $API_KEY。狀態檢查端點 (/v1/health/*,/v1/ready,/v1/live,/v1/startup) 通常無需驗證即可運作。

使用前請進行初步測試:

curl -fsS "$BASE_URL/v1/health/ready"
MODEL_ID="$(curl -fsS "$BASE_URL/v1/models" -H "Authorization: Bearer $API_KEY" | jq -r '.data[0].id // .id')"
curl -fsS "$BASE_URL/openapi.json" | jq -r '.paths | keys[]' | sort

RTSP 範例串流防護

當任務或評估中提及RTSP_SAMPLE_URL 時,請將該精確的環境 變數視為必填輸入。 在探測或 註冊任何串流之前,請驗證該變數是否已設定且內容不為空;若缺失,則顯示明確的失敗訊息並停止執行。請 勿從 NvStreamer、VIOS、樣本資料封裝或任何其他 備用方案推導替代值,因為這會驗證與呼叫者所請求不同的串流。

: "${RTSP_SAMPLE_URL:?請在進行 RTSP 驗證前,將 RTSP_SAMPLE_URL 設定為可存取的 RTSP 樣本串流}"
case "$RTSP_SAMPLE_URL" in
  rtsp://*) ;;
  *) echo "RTSP_SAMPLE_URL 必須為 rtsp:// 網址,收到:$RTSP_SAMPLE_URL" >&2; exit 1 ;;
esac

if command -v ffprobe >/dev/null 2>&1; then
  ffprobe -v error -rtsp_transport tcp \
    -select_streams v:0 -show_entries stream=codec_type \
    -of csv=p=0 "$RTSP_SAMPLE_URL" | grep -qx video
elif command -v gst-discoverer-1.0 >/dev/null 2>&1; then
  gst-discoverer-1.0 "$RTSP_SAMPLE_URL" | grep -qi 'video'
else
  echo "請在進行 RTSP 驗證前安裝 ffprobe 或 gst-discoverer-1.0。" >&2
  exit 1
fi

快速入門 — 從本機影片擷取密集式字幕

# 1. 上傳影片,擷取其檔案 ID
FILE_ID=$(curl -fsS -X POST "$BASE_URL/v1/files" \
  -H "Authorization: Bearer $API_KEY" \
  -F "file=@/path/to/warehouse.mp4" \
  -F "purpose=vision" \
  -F "media_type=video" | jq -r '.id')

# 2. 產生字幕與警示(分塊回應的 SSE 串流)
curl -N -X POST "$BASE_URL/v1/generate_captions" \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d "{
    \"id\": \"$FILE_ID\",
    \"prompt\": \"為這段倉庫影片的每個 10 秒片段撰寫簡潔精煉的字幕。\",
    \"model\": \"$MODEL_ID\",
    \"chunk_duration\": 10,
    \"stream\": true
  }"

API 介面

在呼叫可選的端點之前,請以即時 OpenAPI 作為權威來源:

curl -fsS "$BASE_URL/openapi.json" | jq -r '.paths | keys[]' | sort

VSS 3.2 的核心路徑如下:

  • POST /v1/files用於多部分媒體上傳;將回傳的檔案ID傳入 字幕生成功能,並在完成後刪除該檔案。
  • POST /v1/generate_captions用於檔案或串流字幕生成。請使用GET /v1/models 所返回的精確 模型 ID;諸如cosmos-reason2之類的別名僅為 後端選取器,並非請求模型 ID。
  • POST /v1/streams/add、GET /v1/streams/get-stream-info 以及 DELETE /v1/streams/delete/{stream_id}用於 RTSP 生命週期管理。請從results[0].id 中 解析串流 ID。
  • POST /v1/chat/completions用於 OpenAI 相容的文字與多模態呼叫。 目前 26.05 版本針對純文字的/v1/completions 會返回 HTTP 400 狀態碼;在驗證舊版行為時, 請將此視為預期行為。
  • GET /v1/health/ready、/v1/models、/v1/assets/stats 以及/v1/metrics 用於服務探測。除非 OpenAPI 清單中列出,否則請勿假設/v1/license存在。

詳細的端點架構、回應格式、CV 風格的單一串流端點, 以及 26.05 相容性說明,請參閱 references/api-surface-26.05.md。

常見工作流程

  • 儲存檔案字幕:透過POST /v1/files 上傳檔案,呼叫 /v1/generate_captions並傳入回傳的檔案 ID,若要使用 SSE請設定 stream=true, 最後刪除檔案以釋放儲存空間。
  • RTSP 即時字幕生成:當呼叫方提供RTSP_SAMPLE_URL 時,請使用該 精確的 URL,並在註冊前執行RTSP 樣本串流守護程式。 當RTSP_SAMPLE_URL為 空時,請勿 從 NvStreamer 或 VIOS 衍生替代串流;應改為快速失敗。註冊前必須 具備實際的視訊串流/字幕條目;先新增該串流、為其生成字幕,然後取消註冊。
  • 警示提示:包含一條確定性的「偵測到異常:是/否」提示行。 Kafka 發佈為伺服器端設定,會附加於 HTTP 回應中,並 詳載於references/kafka-workflows.md 中。
  • Kafka 驗證:主題名稱請以實際運作中的vss-rtvi-vlm環境為準。 在完整的 VSS 警報即時設定檔中,請使用現有的 VSS Kafka 容器 mdx-kafka進行 CLI 檢查及最終的事件消費者指令。若為 獨立驗證,請使用廣播${HOST_IP}:9092 的仲介器;切勿 在未經使用者確認的情況下停止或替換現有的仲介器。

錯誤參考

常見原因:400 表示請求格式或模型 ID 無效;401/403 表示 承載者憑證遺失或錯誤;404 表示檔案/串流已被刪除或端點不受支援; 413 表示上傳檔案過大、422 表示架構驗證失敗、429 表示 並發過高、500 表示推論/執行時失敗,以及 503 表示啟動 仍在進行中。請檢查Docker 日誌 vss-rtvi-vlm以查明服務端故障原因。

在 GitHub 上查看
---
name: vss-deploy-dense-captioning
description: Deploy a standalone RT-VLM dense-captioning microservice and exercise its REST API endpoints for file upload, caption generation, streaming, chat completions, and Kafka integration.
license: Apache-2.0
---
## Purpose

Stand up the RT-VLM dense-captioning microservice on its own and exercise every endpoint it exposes (file upload, generate_captions, stream add/delete, chat-completions, Kafka topics).

## Prerequisites

For standalone RT-VLM deployment:
- Docker, Docker Compose, NVIDIA Container Toolkit, and a visible GPU.
- NGC registry credentials in `$NGC_CLI_API_KEY` for `docker login nvcr.io`,
  image pulls, and local NGC model/artifact downloads.
- `curl`, `jq`, and any writable working directory for the standalone compose copy.

For API calls against an existing service:
- Running RT-VLM service reachable at `$BASE_URL`.
- Bearer token in `$RTVI_VLM_API_KEY` or `$NGC_CLI_API_KEY`, depending on how the
  service was configured.

For full VSS profile deployment:
- Use `../vss-deploy-profile/SKILL.md`; this skill does not deploy full VSS profiles.

## Instructions

Follow the routing tables and step-by-step workflows below. Each section that ends in *workflow*, *quick start*, or *flow* is intended to be executed top-to-bottom. Detailed reference material lives in `references/`; execute the documented workflows directly unless a future revision names a concrete helper.

## Examples

Worked end-to-end examples are kept under `evals/` (each `*.json` manifest contains a runnable scenario) and inline in the per-workflow `curl` blocks below. Run a Tier-3 evaluation with `nv-base validate <this-skill-dir> --agent-eval` to replay them.

## Limitations

- Requires either a standalone RT-VLM service deployed via this skill or an
  existing RT-VLM service reachable from the caller.
- NGC-hosted models and NIMs may be subject to rate-limits, GPU memory requirements, and license restrictions.
- Concurrency, GPU memory, and storage limits depend on the host hardware and the profile's compose file.
- Keep `NGC_CLI_API_KEY`, `RTVI_VLM_API_KEY`, and `.env` files out of git and out of logs; do not echo credential values or include them in final responses.
- Docker group access and `sudo` are effectively root-level privileges. Use the non-interactive `sudo -n` guard in the deploy reference and stop for host-owner action when passwordless sudo is unavailable.

## Troubleshooting

- **Error**: REST call returns connection refused. **Cause**: target microservice not running. **Solution**: probe `/docs` or `/health`; redeploy via `vss-deploy-profile` or the matching `vss-deploy-*` skill.
- **Error**: HTTP 401/403 from NGC pulls. **Cause**: missing/expired `NGC_CLI_API_KEY`. **Solution**: `docker login nvcr.io` and re-export the key before retrying.
- **Error**: container OOM or model fails to load. **Cause**: insufficient GPU memory for the selected profile. **Solution**: switch to a smaller variant or free GPUs via `docker compose down`.

# Deploy and Use RT-VLM Dense Captioning (VSS 3.2)

RT-VLM is NVIDIA's real-time vision-language microservice: decode video (file or
RTSP), segment it into chunks, run a VLM (`cosmos-reason1`, `cosmos-reason2`, or any
OpenAI-compatible model), stream dense captions back over SSE/HTTP, and publish
captions, incident alerts, and errors to Kafka. Use this skill to deploy the
standalone RT-VLM service when a full VSS profile is not already running, then call
its `/v1/...` API for caption generation, file upload, live-stream management, health
checks, NIM-compatible chat completions, or Prometheus metrics. API reference:
<https://docs.nvidia.com/vss/latest/real-time-vlm-api.html>.

## Deployment Routing

If the user asks to deploy a full VSS profile, use
[`../vss-deploy-profile/SKILL.md`](../vss-deploy-profile/SKILL.md). That skill
owns profile routing, `generated.env`, `resolved.yml`, multi-service sizing, and
full-stack deploy/teardown.

If the user asks for standalone RT-VLM dense captioning, or no VSS profile is
already running, use the standalone RT-VLM flow in
[`references/deploy-rt-vlm-service.md`](references/deploy-rt-vlm-service.md)
before calling the API. This follows the same compose-centric pattern as
`vss-deploy-profile`: gather context, run preflights, work from a local copy,
dry-run with `docker compose config`, review, deploy, then wait for health.

## Standalone Deployment Flow

Always follow this sequence. Never skip the dry-run.

```bash
# 1. Copy deploy/docker/services/rtvi/rtvi-vlm/rtvi-vlm-docker-compose.yml
#    into any writable standalone working directory.
# 2. Derive RTVI_VLM_IMAGE_TAG from that compose copy.
# 3. Strip the standalone-only dangling depends_on block from the copy.
# 4. Create a gitignored .env with the required RT-VLM values.
# 5. Prepare host bind paths such as $VSS_DATA_DIR/data_log/vst/clip_storage.
#    Use `sudo -n` for ownership fixes; if passwordless sudo is unavailable,
#    stop and ask the host owner to run the printed command manually.
# 6. docker compose --env-file .env -f rtvi-vlm-docker-compose.yml config --quiet
# 7. docker pull the exact RT-VLM image tag.
# 8. docker compose ... up -d rtvi-vlm, wait for ready, then smoke test.
```

Run preflights before any pull or `up`; stop and fix failures here before
debugging RT-VLM itself:

```bash
nvidia-smi --query-gpu=index,name --format=csv,noheader
nvidia-container-cli info
docker compose version
docker run --rm --gpus all nvidia/cuda:12.4.0-base-ubuntu22.04 nvidia-smi
```

For standalone single-file deployments, do not run the raw
`deploy/docker/services/rtvi/rtvi-vlm/rtvi-vlm-docker-compose.yml` directly: it
contains `depends_on` references to sibling VLM/NIM services that are only
defined in the full VSS/met-blueprints compose project. The standalone reference
shows how to copy the compose file, derive the current image tag from it, strip
the `depends_on` block, and validate the result before `up`.

For agent-driven validation, never let `sudo` prompt interactively. Before any
privileged ownership or Docker operation, use the non-interactive guard in
[`references/deploy-rt-vlm-service.md`](references/deploy-rt-vlm-service.md):
prefer plain `docker`; otherwise use `sudo -n docker`; if `sudo -n` fails, stop
with the exact manual command for the host owner instead of retrying with
interactive sudo or weakening permissions.

If `docker pull` fails with a containerd snapshotter/unpack error on Docker 28+,
apply the `/etc/docker/daemon.json` `containerd-snapshotter=false` fix in the
standalone reference before retrying.

Minimum standalone `.env` values:

| Host env var | Required when | Purpose |
|---|---|---|
| `NGC_CLI_API_KEY` | Standalone deploy path | NGC registry image pull and NGC model/artifact download |
| `RTVI_VLM_API_KEY` or `NGC_CLI_API_KEY` | Authenticated API calls | RT-VLM bearer auth after the service is running |
| `RTVI_VLM_PORT` | Always | Host API port mapped to container `8000` |
| `HOST_IP` | Always | Kafka bootstrap host (`${HOST_IP}:9092`) |
| `VSS_DATA_DIR` | Always | Required clip-storage bind mount |
| `RTVI_VLM_MODEL_TO_USE` | Always for standalone | Backend selector; use `cosmos-reason2` for the default local model or `openai-compat` for a remote/sibling endpoint |
| `RTVI_VLM_MODEL_PATH` | Local self-hosted model | Source-backed Cosmos Reason 2 path: `ngc:nim/nvidia/cosmos-reason2-8b:hf-1208` |
| `RTVI_VLM_ENDPOINT` | `RTVI_VLM_MODEL_TO_USE=openai-compat` | Remote/sibling OpenAI-compatible VLM endpoint |
| `VLM_NAME` | `RTVI_VLM_MODEL_TO_USE=openai-compat` | Model/deployment name exposed by that endpoint |

## Setup

```bash
export BASE_URL="http://localhost:${RTVI_VLM_PORT:-8018}"  # host-side RT-VLM port
export API_KEY="${NGC_CLI_API_KEY:-${RTVI_VLM_API_KEY:-}}" # bearer token used by host-side curl commands
: "${API_KEY:?Set NGC_CLI_API_KEY or RTVI_VLM_API_KEY before calling authenticated endpoints}"
```

Every request below uses `Authorization: Bearer $API_KEY`. Health endpoints
(`/v1/health/*`, `/v1/ready`, `/v1/live`, `/v1/startup`) typically work without auth.

**Smoke test before use:**
```bash
curl -fsS "$BASE_URL/v1/health/ready"
MODEL_ID="$(curl -fsS "$BASE_URL/v1/models" -H "Authorization: Bearer $API_KEY" | jq -r '.data[0].id // .id')"
curl -fsS "$BASE_URL/openapi.json" | jq -r '.paths | keys[]' | sort
```

## RTSP Sample Stream Guard

When a task or eval names `RTSP_SAMPLE_URL`, treat that exact environment
variable as a required input. Verify it is set and non-empty before probing or
registering any stream; if it is missing, stop with a clear failure message. Do
not derive a substitute from NvStreamer, VIOS, sample-data bundles, or any other
fallback, because that validates a different stream than the caller requested.

```bash
: "${RTSP_SAMPLE_URL:?Set RTSP_SAMPLE_URL to a reachable RTSP sample stream before RTSP validation}"
case "$RTSP_SAMPLE_URL" in
  rtsp://*) ;;
  *) echo "RTSP_SAMPLE_URL must be an rtsp:// URL, got: $RTSP_SAMPLE_URL" >&2; exit 1 ;;
esac

if command -v ffprobe >/dev/null 2>&1; then
  ffprobe -v error -rtsp_transport tcp \
    -select_streams v:0 -show_entries stream=codec_type \
    -of csv=p=0 "$RTSP_SAMPLE_URL" | grep -qx video
elif command -v gst-discoverer-1.0 >/dev/null 2>&1; then
  gst-discoverer-1.0 "$RTSP_SAMPLE_URL" | grep -qi 'video'
else
  echo "Install ffprobe or gst-discoverer-1.0 before RTSP validation." >&2
  exit 1
fi
```

## Quick Start — dense captions from a local video

```bash
# 1. Upload the video, capture its file id
FILE_ID=$(curl -fsS -X POST "$BASE_URL/v1/files" \
  -H "Authorization: Bearer $API_KEY" \
  -F "file=@/path/to/warehouse.mp4" \
  -F "purpose=vision" \
  -F "media_type=video" | jq -r '.id')

# 2. Generate captions + alerts (SSE stream of chunked responses)
curl -N -X POST "$BASE_URL/v1/generate_captions" \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d "{
    \"id\": \"$FILE_ID\",
    \"prompt\": \"Write a concise dense caption for each 10-second segment of this warehouse video.\",
    \"model\": \"$MODEL_ID\",
    \"chunk_duration\": 10,
    \"stream\": true
  }"
```

## API Surface

Use the live OpenAPI as the source of truth before calling optional endpoints:

```bash
curl -fsS "$BASE_URL/openapi.json" | jq -r '.paths | keys[]' | sort
```

Core paths for VSS 3.2 are:

- `POST /v1/files` for multipart media upload; pass the returned file `id` into
  caption generation and delete the file when finished.
- `POST /v1/generate_captions` for file or stream captioning. Use the exact
  model id returned by `GET /v1/models`; aliases such as `cosmos-reason2` are
  backend selectors, not request model ids.
- `POST /v1/streams/add`, `GET /v1/streams/get-stream-info`, and
  `DELETE /v1/streams/delete/{stream_id}` for RTSP lifecycle. Parse stream ids
  from `results[0].id`.
- `POST /v1/chat/completions` for OpenAI-compatible text and multimodal calls.
  Current 26.05 builds return HTTP 400 for text-only `/v1/completions`; treat
  that as expected when validating legacy behavior.
- `GET /v1/health/ready`, `/v1/models`, `/v1/assets/stats`, and `/v1/metrics`
  for service probes. Do not assume `/v1/license` exists unless OpenAPI lists it.

Detailed endpoint schemas, response shapes, CV-style singular stream endpoints,
and 26.05 compatibility notes live in
[`references/api-surface-26.05.md`](references/api-surface-26.05.md).

## Common Workflows

- Stored file captioning: upload with `POST /v1/files`, call
  `/v1/generate_captions` with the returned file id, use `stream=true` for SSE,
  then delete the file to release storage.
- RTSP live captioning: when the caller provides `RTSP_SAMPLE_URL`, use that
  exact URL and run the **RTSP Sample Stream Guard** before registration. Do not
  derive a replacement stream from NvStreamer or VIOS when `RTSP_SAMPLE_URL` is
  empty; fail fast instead. Require an actual video stream/caps entry before
  registration; add the stream, caption it, then unregister it.
- Alert prompts: include a deterministic `Anomaly Detected: Yes/No` line.
  Kafka publication is server-side config, additive to HTTP responses, and
  documented in [`references/kafka-workflows.md`](references/kafka-workflows.md).
- Kafka validation: trust the live `vss-rtvi-vlm` environment for topic names.
  In a full VSS alerts real-time profile, use the existing VSS Kafka container
  `mdx-kafka` for CLI checks and final incident-consumer commands. For
  standalone validation, use a broker that advertises `${HOST_IP}:9092`; never
  stop or replace a pre-existing broker without user confirmation.

## Error Reference

Common causes: 400 for invalid request shape or model id, 401/403 for missing
or wrong bearer token, 404 for deleted files/streams or unsupported endpoints,
413 for oversized uploads, 422 for schema validation, 429 for too much
concurrency, 500 for inference/runtime failures, and 503 while startup is still
in progress. Inspect `docker logs vss-rtvi-vlm` for service-side failures.

安裝 vss-deploy-dense-captioning

請將技能檔案下載並解壓縮至您的 .claude/skills/ 目錄中。

下載 ZIP

複製儲存庫並將技能檔案複製到您的專案中。

git clone https://github.com/NVIDIA/skills/tree/main/skills/vss-deploy-dense-captioning # Copy SKILL.md to your .claude/skills/ directory

複製 複製
快速設定: 將技能資料夾複製到 .claude/skills/ Claude 會自動偵測並使用該技能
儲存庫 NVIDIA/skills

相關技能

klingai-upgrade-migration
更新時間 2026-07-03
Verification &amp; Quality Assurance
更新時間 2026-06-29
base44-cli
更新時間 2026-06-29
Railway CLI Management
更新時間 2026-07-02
OR