vss-deploy-dense-captioning
NVIDIA/skills
部署一個獨立運作的 RT-VLM 密集式字幕生成微服務,並測試其 REST API 端點,包括檔案上傳、字幕生成、串流、聊天自動完成功能,以及與 Kafka 的整合。
...展開全部目的
將 RT-VLM 高密度字幕生成微服務獨立部署,並測試其所公開的每個端點(檔案上傳、generate_captions、串流新增/刪除、聊天自動完成、Kafka 主題)。
先決條件
針對 RT-VLM 獨立部署:
- 需具備 Docker、Docker Compose、NVIDIA Container Toolkit 以及一顆可用的 GPU。
$NGC_CLI_API_KEY中須包含 NGC 註冊表憑證,以便透過 `docker login nvcr.io` 登入、 拉取映像檔,以及下載本機 NGC 模型/構件。curl、jq,以及任何可供獨立部署的 Compose 副本寫入的工作目錄。
針對現有服務進行 API 呼叫時:
- 可透過
$BASE_URL存取的 RT-VLM 服務。 $RTVI_VLM_API_KEY或$NGC_CLI_API_KEY中的 Bearer 憑證,具體取決於 服務的配置方式。
若要部署完整的 VSS 設定檔:
- 請使用
../vss-deploy-profile/SKILL.md;此技能不會部署完整的 VSS 設定檔。
操作說明
請遵循以下路由表及逐步工作流程。每個以「工作流程」、「快速入門」或「流程」結尾的章節,皆應自上而下依序執行。詳細參考資料存放於references/ 目錄中;除非未來修訂版指定了具體的輔助程式,否則請直接執行文件中記載的工作流程。
範例
已驗證的端到端範例存放於evals/目錄下(每個*.json清單皆包含可執行的情境),並內嵌於下方的各工作流程curl區塊中。請執行 Tier-3 評估(使用nv-base validate)以重現這些範例。
限制
- 需具備透過此技能部署的獨立 RT-VLM 服務,或 呼叫方可存取的現有 RT-VLM 服務。
- 由 NGC 託管的模型和 NIM 可能會受到速率限制、GPU 記憶體需求以及授權限制的影響。
- 並發數、GPU 記憶體及儲存空間的限制取決於主機硬體及配置檔的 compose 檔案。
- 請將
NGC_CLI_API_KEY、RTVI_VLM_API_KEY及.env檔案排除在 git 及日誌之外;切勿回顯憑證值,亦勿將其包含在最終回應中。 - Docker 群組存取權限與
sudo實質上屬於 root 級別的權限。當無法使用無密碼 sudo 時,請在部署參考中使用非互動式sudo -n防護機制,並將「停止」操作保留給主機擁有者執行。
疑難排解
- 錯誤:REST 呼叫傳回「連線遭拒絕」。原因:目標微服務未運行。解決方案:偵測
/docs或/health;透過vss-deploy-profile或相應的vss-deploy-*技能重新部署。 - 錯誤:NGC 拉取時出現 HTTP 401/403 錯誤。原因:
NGC_CLI_API_KEY遺失或已過期。解決方案:執行 `docker login nvcr.io` 並重新匯出金鑰後再嘗試。 - 錯誤:容器發生 OOM 或模型無法載入。原因:所選設定檔的 GPU 記憶體不足。解決方案:切換至較小的變體,或透過 `
docker compose down` 釋放 GPU 資源。
部署與使用 RT-VLM 密集式圖說生成(VSS 3.2)
RT-VLM 是 NVIDIA 的即時視覺語言微服務:解碼影片(檔案或
RTSP),將其分割成區塊,執行 VLM(cosmos-reason1、cosmos-reason2 或任何
與 OpenAI 相容的模型),透過 SSE/HTTP 串流傳回密集式字幕,並將
字幕、事件警示及錯誤發佈至 Kafka。當尚未運行完整的 VSS 配置檔時,請使用此技能部署
獨立的 RT-VLM 服務,然後呼叫
其/v1/...API 來進行字幕生成、檔案上傳、直播管理、健康
檢查、與 NIM 相容的聊天自動完成功能,或 Prometheus 指標。API 參考:
https://docs.nvidia.com/vss/latest/real-time-vlm-api.html。
部署路由
若使用者要求部署完整的 VSS 配置檔,請使用
../vss-deploy-profile/SKILL.md。該技能
負責配置檔路由、generated.env、resolved.yml、多服務規模配置,以及
全堆疊部署/拆除。
若使用者要求獨立運作的 RT-VLM 高密度字幕生成,或目前尚無 VSS 配置檔
正在執行,請在呼叫 API 之前,使用
references/deploy-rt-vlm-service.md
中的獨立 RT-VLM 流程。 此流程遵循與
vss-deploy-profile 相同的 Compose 為中心的模式:收集上下文、執行預檢查、基於本地副本操作、
使用Docker Compose 配置進行模擬執行、審查、部署,然後等待狀態確認。
獨立部署流程
請務必遵循此順序。切勿跳過模擬執行步驟。
# 1. 將 deploy/docker/services/rtvi/rtvi-vlm/rtvi-vlm-docker-compose.yml
# 複製到任何可寫入的獨立運作工作目錄中。
# 2. 從該 Compose 副本中推導出 RTVI_VLM_IMAGE_TAG。
# 3. 從該複製檔中移除僅適用於獨立部署的冗餘 depends_on 區塊。
# 4. 建立一個被 git 忽略的 .env 檔案,並填入所需的 RT-VLM 值。
# 5. 準備主機的綁定路徑,例如 $VSS_DATA_DIR/data_log/vst/clip_storage。
# 使用 `sudo -n` 修正所有權;若無法使用無密碼 sudo,
# 請暫停並請主機擁有者手動執行列印出的指令。
# 6. docker compose --env-file .env -f rtvi-vlm-docker-compose.yml config --quiet
# 7. 使用 `docker pull` 拉取確切的 RT-VLM 映像標籤。
# 8. 執行 `docker compose ... up -d rtvi-vlm`,等待狀態轉為「ready」,然後進行初步測試。
在執行任何 `pull` 或`up` 指令前,請先執行預檢;在此階段先停止並修正錯誤,
再進行 RT-VLM 本身的除錯:
nvidia-smi --query-gpu=index,name --format=csv,noheader
nvidia-container-cli info
docker compose version
docker run --rm --gpus all nvidia/cuda:12.4.0-base-ubuntu22.04 nvidia-smi
對於獨立的單一檔案部署,請勿直接執行原始的
deploy/docker/services/rtvi/rtvi-vlm/rtvi-vlm-docker-compose.yml:該檔案
包含對同級 VLM/NIM 服務的depends_on引用,這些服務僅
僅在完整的 VSS/met-blueprints Compose 專案中定義。此獨立部署範例
說明了如何複製 Compose 檔案、據此推導當前映像標籤、移除
`depends_on` 區塊,並在啟動前驗證結果。
對於代理驅動的驗證,切勿讓sudo提示框進入互動模式。在執行任何
特權擁有權設定或 Docker 操作之前,請使用references/deploy-rt-vlm-service.md 中
的非互動式防護機制:
優先使用純文字模式的docker;否則使用sudo -n docker;若sudo -n失敗,請直接
使用主機擁有者的精確手動命令終止操作,而非嘗試使用
互動式 sudo 或降低權限。
若在 Docker 28 以上版本中,docker pull因 containerd snapshotter/unpack 錯誤而失敗,
請在重新嘗試前,套用獨立部署參考文件中/etc/docker/daemon.json的
containerd-snapshotter=false修正方案。
獨立執行環境的.env最小設定值:
| 主機環境變數 | 何時需要 | 用途 |
|---|---|---|
NGC_CLI_API_KEY |
獨立部署路徑 | NGC 註冊表映像拉取及 NGC 模型/產出檔案下載 |
RTVI_VLM_API_KEY或NGC_CLI_API_KEY |
經身份驗證的 API 呼叫 | 服務運行後的 RT-VLM 承載者驗證 |
RTVI_VLM_PORT |
始終 | 主機 API 埠映射至容器8000 |
HOST_IP |
始終 | Kafka 啟動主機 (${HOST_IP}:9092) |
VSS_DATA_DIR |
始終 | 必需的 clip-storage 綁定掛載 |
RTVI_VLM_MODEL_TO_USE |
獨立執行時始終啟用 | 後端選擇器;若要使用預設的本地模型,請使用cosmos-reason2;若要使用遠端或同級端點,請使用openai-compat |
RTVI_VLM_MODEL_PATH |
本地自託管模型 | 基於原始碼的 Cosmos Reason 2 路徑:ngc:nim/nvidia/cosmos-reason2-8b:hf-1208 |
RTVI_VLM_ENDPOINT |
RTVI_VLM_MODEL_TO_USE=openai-compat |
遠端/同級 OpenAI 相容 VLM 端點 |
VLM_NAME |
RTVI_VLM_MODEL_TO_USE=openai-compat |
該端點所公開的模型/部署名稱 |
設定
export BASE_URL="http://localhost:${RTVI_VLM_PORT:-8018}" # 主機端 RT-VLM 連接埠
export API_KEY="${NGC_CLI_API_KEY:-${RTVI_VLM_API_KEY:-}}" # 主機端 curl 指令使用的 Bearer 憑證
: "${API_KEY:?在呼叫需驗證的端點前,請先設定 NGC_CLI_API_KEY 或 RTVI_VLM_API_KEY}"
以下每個請求均使用Authorization: Bearer $API_KEY。狀態檢查端點
(/v1/health/*,/v1/ready,/v1/live,/v1/startup) 通常無需驗證即可運作。
使用前請進行初步測試:
curl -fsS "$BASE_URL/v1/health/ready"
MODEL_ID="$(curl -fsS "$BASE_URL/v1/models" -H "Authorization: Bearer $API_KEY" | jq -r '.data[0].id // .id')"
curl -fsS "$BASE_URL/openapi.json" | jq -r '.paths | keys[]' | sort
RTSP 範例串流防護
當任務或評估中提及RTSP_SAMPLE_URL 時,請將該精確的環境
變數視為必填輸入。 在探測或
註冊任何串流之前,請驗證該變數是否已設定且內容不為空;若缺失,則顯示明確的失敗訊息並停止執行。請
勿從 NvStreamer、VIOS、樣本資料封裝或任何其他
備用方案推導替代值,因為這會驗證與呼叫者所請求不同的串流。
: "${RTSP_SAMPLE_URL:?請在進行 RTSP 驗證前,將 RTSP_SAMPLE_URL 設定為可存取的 RTSP 樣本串流}"
case "$RTSP_SAMPLE_URL" in
rtsp://*) ;;
*) echo "RTSP_SAMPLE_URL 必須為 rtsp:// 網址,收到:$RTSP_SAMPLE_URL" >&2; exit 1 ;;
esac
if command -v ffprobe >/dev/null 2>&1; then
ffprobe -v error -rtsp_transport tcp \
-select_streams v:0 -show_entries stream=codec_type \
-of csv=p=0 "$RTSP_SAMPLE_URL" | grep -qx video
elif command -v gst-discoverer-1.0 >/dev/null 2>&1; then
gst-discoverer-1.0 "$RTSP_SAMPLE_URL" | grep -qi 'video'
else
echo "請在進行 RTSP 驗證前安裝 ffprobe 或 gst-discoverer-1.0。" >&2
exit 1
fi
快速入門 — 從本機影片擷取密集式字幕
# 1. 上傳影片,擷取其檔案 ID
FILE_ID=$(curl -fsS -X POST "$BASE_URL/v1/files" \
-H "Authorization: Bearer $API_KEY" \
-F "file=@/path/to/warehouse.mp4" \
-F "purpose=vision" \
-F "media_type=video" | jq -r '.id')
# 2. 產生字幕與警示(分塊回應的 SSE 串流)
curl -N -X POST "$BASE_URL/v1/generate_captions" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d "{
\"id\": \"$FILE_ID\",
\"prompt\": \"為這段倉庫影片的每個 10 秒片段撰寫簡潔精煉的字幕。\",
\"model\": \"$MODEL_ID\",
\"chunk_duration\": 10,
\"stream\": true
}"
API 介面
在呼叫可選的端點之前,請以即時 OpenAPI 作為權威來源:
curl -fsS "$BASE_URL/openapi.json" | jq -r '.paths | keys[]' | sort
VSS 3.2 的核心路徑如下:
POST /v1/files用於多部分媒體上傳;將回傳的檔案ID傳入 字幕生成功能,並在完成後刪除該檔案。POST /v1/generate_captions用於檔案或串流字幕生成。請使用GET /v1/models所返回的精確 模型 ID;諸如cosmos-reason2之類的別名僅為 後端選取器,並非請求模型 ID。POST /v1/streams/add、GET /v1/streams/get-stream-info以及DELETE /v1/streams/delete/{stream_id}用於 RTSP 生命週期管理。請從results[0].id中 解析串流 ID。POST /v1/chat/completions用於 OpenAI 相容的文字與多模態呼叫。 目前 26.05 版本針對純文字的/v1/completions會返回 HTTP 400 狀態碼;在驗證舊版行為時, 請將此視為預期行為。GET /v1/health/ready、/v1/models、/v1/assets/stats以及/v1/metrics用於服務探測。除非 OpenAPI 清單中列出,否則請勿假設/v1/license存在。
詳細的端點架構、回應格式、CV 風格的單一串流端點,
以及 26.05 相容性說明,請參閱
references/api-surface-26.05.md。
常見工作流程
- 儲存檔案字幕:透過
POST /v1/files上傳檔案,呼叫/v1/generate_captions並傳入回傳的檔案 ID,若要使用 SSE請設定 stream=true,最後刪除檔案以釋放儲存空間。 - RTSP 即時字幕生成:當呼叫方提供
RTSP_SAMPLE_URL時,請使用該 精確的 URL,並在註冊前執行RTSP 樣本串流守護程式。 當RTSP_SAMPLE_URL為 空時,請勿 從 NvStreamer 或 VIOS 衍生替代串流;應改為快速失敗。註冊前必須 具備實際的視訊串流/字幕條目;先新增該串流、為其生成字幕,然後取消註冊。 - 警示提示:包含一條確定性的
「偵測到異常:是/否」提示行。 Kafka 發佈為伺服器端設定,會附加於 HTTP 回應中,並 詳載於references/kafka-workflows.md 中。 - Kafka 驗證:主題名稱請以實際運作中的
vss-rtvi-vlm環境為準。 在完整的 VSS 警報即時設定檔中,請使用現有的 VSS Kafka 容器mdx-kafka進行 CLI 檢查及最終的事件消費者指令。若為 獨立驗證,請使用廣播${HOST_IP}:9092的仲介器;切勿 在未經使用者確認的情況下停止或替換現有的仲介器。
錯誤參考
常見原因:400 表示請求格式或模型 ID 無效;401/403 表示
承載者憑證遺失或錯誤;404 表示檔案/串流已被刪除或端點不受支援;
413 表示上傳檔案過大、422 表示架構驗證失敗、429 表示
並發過高、500 表示推論/執行時失敗,以及 503 表示啟動
仍在進行中。請檢查Docker 日誌 vss-rtvi-vlm以查明服務端故障原因。
---
name: vss-deploy-dense-captioning
description: Deploy a standalone RT-VLM dense-captioning microservice and exercise its REST API endpoints for file upload, caption generation, streaming, chat completions, and Kafka integration.
license: Apache-2.0
---
## Purpose
Stand up the RT-VLM dense-captioning microservice on its own and exercise every endpoint it exposes (file upload, generate_captions, stream add/delete, chat-completions, Kafka topics).
## Prerequisites
For standalone RT-VLM deployment:
- Docker, Docker Compose, NVIDIA Container Toolkit, and a visible GPU.
- NGC registry credentials in `$NGC_CLI_API_KEY` for `docker login nvcr.io`,
image pulls, and local NGC model/artifact downloads.
- `curl`, `jq`, and any writable working directory for the standalone compose copy.
For API calls against an existing service:
- Running RT-VLM service reachable at `$BASE_URL`.
- Bearer token in `$RTVI_VLM_API_KEY` or `$NGC_CLI_API_KEY`, depending on how the
service was configured.
For full VSS profile deployment:
- Use `../vss-deploy-profile/SKILL.md`; this skill does not deploy full VSS profiles.
## Instructions
Follow the routing tables and step-by-step workflows below. Each section that ends in *workflow*, *quick start*, or *flow* is intended to be executed top-to-bottom. Detailed reference material lives in `references/`; execute the documented workflows directly unless a future revision names a concrete helper.
## Examples
Worked end-to-end examples are kept under `evals/` (each `*.json` manifest contains a runnable scenario) and inline in the per-workflow `curl` blocks below. Run a Tier-3 evaluation with `nv-base validate <this-skill-dir> --agent-eval` to replay them.
## Limitations
- Requires either a standalone RT-VLM service deployed via this skill or an
existing RT-VLM service reachable from the caller.
- NGC-hosted models and NIMs may be subject to rate-limits, GPU memory requirements, and license restrictions.
- Concurrency, GPU memory, and storage limits depend on the host hardware and the profile's compose file.
- Keep `NGC_CLI_API_KEY`, `RTVI_VLM_API_KEY`, and `.env` files out of git and out of logs; do not echo credential values or include them in final responses.
- Docker group access and `sudo` are effectively root-level privileges. Use the non-interactive `sudo -n` guard in the deploy reference and stop for host-owner action when passwordless sudo is unavailable.
## Troubleshooting
- **Error**: REST call returns connection refused. **Cause**: target microservice not running. **Solution**: probe `/docs` or `/health`; redeploy via `vss-deploy-profile` or the matching `vss-deploy-*` skill.
- **Error**: HTTP 401/403 from NGC pulls. **Cause**: missing/expired `NGC_CLI_API_KEY`. **Solution**: `docker login nvcr.io` and re-export the key before retrying.
- **Error**: container OOM or model fails to load. **Cause**: insufficient GPU memory for the selected profile. **Solution**: switch to a smaller variant or free GPUs via `docker compose down`.
# Deploy and Use RT-VLM Dense Captioning (VSS 3.2)
RT-VLM is NVIDIA's real-time vision-language microservice: decode video (file or
RTSP), segment it into chunks, run a VLM (`cosmos-reason1`, `cosmos-reason2`, or any
OpenAI-compatible model), stream dense captions back over SSE/HTTP, and publish
captions, incident alerts, and errors to Kafka. Use this skill to deploy the
standalone RT-VLM service when a full VSS profile is not already running, then call
its `/v1/...` API for caption generation, file upload, live-stream management, health
checks, NIM-compatible chat completions, or Prometheus metrics. API reference:
<https://docs.nvidia.com/vss/latest/real-time-vlm-api.html>.
## Deployment Routing
If the user asks to deploy a full VSS profile, use
[`../vss-deploy-profile/SKILL.md`](../vss-deploy-profile/SKILL.md). That skill
owns profile routing, `generated.env`, `resolved.yml`, multi-service sizing, and
full-stack deploy/teardown.
If the user asks for standalone RT-VLM dense captioning, or no VSS profile is
already running, use the standalone RT-VLM flow in
[`references/deploy-rt-vlm-service.md`](references/deploy-rt-vlm-service.md)
before calling the API. This follows the same compose-centric pattern as
`vss-deploy-profile`: gather context, run preflights, work from a local copy,
dry-run with `docker compose config`, review, deploy, then wait for health.
## Standalone Deployment Flow
Always follow this sequence. Never skip the dry-run.
```bash
# 1. Copy deploy/docker/services/rtvi/rtvi-vlm/rtvi-vlm-docker-compose.yml
# into any writable standalone working directory.
# 2. Derive RTVI_VLM_IMAGE_TAG from that compose copy.
# 3. Strip the standalone-only dangling depends_on block from the copy.
# 4. Create a gitignored .env with the required RT-VLM values.
# 5. Prepare host bind paths such as $VSS_DATA_DIR/data_log/vst/clip_storage.
# Use `sudo -n` for ownership fixes; if passwordless sudo is unavailable,
# stop and ask the host owner to run the printed command manually.
# 6. docker compose --env-file .env -f rtvi-vlm-docker-compose.yml config --quiet
# 7. docker pull the exact RT-VLM image tag.
# 8. docker compose ... up -d rtvi-vlm, wait for ready, then smoke test.
```
Run preflights before any pull or `up`; stop and fix failures here before
debugging RT-VLM itself:
```bash
nvidia-smi --query-gpu=index,name --format=csv,noheader
nvidia-container-cli info
docker compose version
docker run --rm --gpus all nvidia/cuda:12.4.0-base-ubuntu22.04 nvidia-smi
```
For standalone single-file deployments, do not run the raw
`deploy/docker/services/rtvi/rtvi-vlm/rtvi-vlm-docker-compose.yml` directly: it
contains `depends_on` references to sibling VLM/NIM services that are only
defined in the full VSS/met-blueprints compose project. The standalone reference
shows how to copy the compose file, derive the current image tag from it, strip
the `depends_on` block, and validate the result before `up`.
For agent-driven validation, never let `sudo` prompt interactively. Before any
privileged ownership or Docker operation, use the non-interactive guard in
[`references/deploy-rt-vlm-service.md`](references/deploy-rt-vlm-service.md):
prefer plain `docker`; otherwise use `sudo -n docker`; if `sudo -n` fails, stop
with the exact manual command for the host owner instead of retrying with
interactive sudo or weakening permissions.
If `docker pull` fails with a containerd snapshotter/unpack error on Docker 28+,
apply the `/etc/docker/daemon.json` `containerd-snapshotter=false` fix in the
standalone reference before retrying.
Minimum standalone `.env` values:
| Host env var | Required when | Purpose |
|---|---|---|
| `NGC_CLI_API_KEY` | Standalone deploy path | NGC registry image pull and NGC model/artifact download |
| `RTVI_VLM_API_KEY` or `NGC_CLI_API_KEY` | Authenticated API calls | RT-VLM bearer auth after the service is running |
| `RTVI_VLM_PORT` | Always | Host API port mapped to container `8000` |
| `HOST_IP` | Always | Kafka bootstrap host (`${HOST_IP}:9092`) |
| `VSS_DATA_DIR` | Always | Required clip-storage bind mount |
| `RTVI_VLM_MODEL_TO_USE` | Always for standalone | Backend selector; use `cosmos-reason2` for the default local model or `openai-compat` for a remote/sibling endpoint |
| `RTVI_VLM_MODEL_PATH` | Local self-hosted model | Source-backed Cosmos Reason 2 path: `ngc:nim/nvidia/cosmos-reason2-8b:hf-1208` |
| `RTVI_VLM_ENDPOINT` | `RTVI_VLM_MODEL_TO_USE=openai-compat` | Remote/sibling OpenAI-compatible VLM endpoint |
| `VLM_NAME` | `RTVI_VLM_MODEL_TO_USE=openai-compat` | Model/deployment name exposed by that endpoint |
## Setup
```bash
export BASE_URL="http://localhost:${RTVI_VLM_PORT:-8018}" # host-side RT-VLM port
export API_KEY="${NGC_CLI_API_KEY:-${RTVI_VLM_API_KEY:-}}" # bearer token used by host-side curl commands
: "${API_KEY:?Set NGC_CLI_API_KEY or RTVI_VLM_API_KEY before calling authenticated endpoints}"
```
Every request below uses `Authorization: Bearer $API_KEY`. Health endpoints
(`/v1/health/*`, `/v1/ready`, `/v1/live`, `/v1/startup`) typically work without auth.
**Smoke test before use:**
```bash
curl -fsS "$BASE_URL/v1/health/ready"
MODEL_ID="$(curl -fsS "$BASE_URL/v1/models" -H "Authorization: Bearer $API_KEY" | jq -r '.data[0].id // .id')"
curl -fsS "$BASE_URL/openapi.json" | jq -r '.paths | keys[]' | sort
```
## RTSP Sample Stream Guard
When a task or eval names `RTSP_SAMPLE_URL`, treat that exact environment
variable as a required input. Verify it is set and non-empty before probing or
registering any stream; if it is missing, stop with a clear failure message. Do
not derive a substitute from NvStreamer, VIOS, sample-data bundles, or any other
fallback, because that validates a different stream than the caller requested.
```bash
: "${RTSP_SAMPLE_URL:?Set RTSP_SAMPLE_URL to a reachable RTSP sample stream before RTSP validation}"
case "$RTSP_SAMPLE_URL" in
rtsp://*) ;;
*) echo "RTSP_SAMPLE_URL must be an rtsp:// URL, got: $RTSP_SAMPLE_URL" >&2; exit 1 ;;
esac
if command -v ffprobe >/dev/null 2>&1; then
ffprobe -v error -rtsp_transport tcp \
-select_streams v:0 -show_entries stream=codec_type \
-of csv=p=0 "$RTSP_SAMPLE_URL" | grep -qx video
elif command -v gst-discoverer-1.0 >/dev/null 2>&1; then
gst-discoverer-1.0 "$RTSP_SAMPLE_URL" | grep -qi 'video'
else
echo "Install ffprobe or gst-discoverer-1.0 before RTSP validation." >&2
exit 1
fi
```
## Quick Start — dense captions from a local video
```bash
# 1. Upload the video, capture its file id
FILE_ID=$(curl -fsS -X POST "$BASE_URL/v1/files" \
-H "Authorization: Bearer $API_KEY" \
-F "file=@/path/to/warehouse.mp4" \
-F "purpose=vision" \
-F "media_type=video" | jq -r '.id')
# 2. Generate captions + alerts (SSE stream of chunked responses)
curl -N -X POST "$BASE_URL/v1/generate_captions" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d "{
\"id\": \"$FILE_ID\",
\"prompt\": \"Write a concise dense caption for each 10-second segment of this warehouse video.\",
\"model\": \"$MODEL_ID\",
\"chunk_duration\": 10,
\"stream\": true
}"
```
## API Surface
Use the live OpenAPI as the source of truth before calling optional endpoints:
```bash
curl -fsS "$BASE_URL/openapi.json" | jq -r '.paths | keys[]' | sort
```
Core paths for VSS 3.2 are:
- `POST /v1/files` for multipart media upload; pass the returned file `id` into
caption generation and delete the file when finished.
- `POST /v1/generate_captions` for file or stream captioning. Use the exact
model id returned by `GET /v1/models`; aliases such as `cosmos-reason2` are
backend selectors, not request model ids.
- `POST /v1/streams/add`, `GET /v1/streams/get-stream-info`, and
`DELETE /v1/streams/delete/{stream_id}` for RTSP lifecycle. Parse stream ids
from `results[0].id`.
- `POST /v1/chat/completions` for OpenAI-compatible text and multimodal calls.
Current 26.05 builds return HTTP 400 for text-only `/v1/completions`; treat
that as expected when validating legacy behavior.
- `GET /v1/health/ready`, `/v1/models`, `/v1/assets/stats`, and `/v1/metrics`
for service probes. Do not assume `/v1/license` exists unless OpenAPI lists it.
Detailed endpoint schemas, response shapes, CV-style singular stream endpoints,
and 26.05 compatibility notes live in
[`references/api-surface-26.05.md`](references/api-surface-26.05.md).
## Common Workflows
- Stored file captioning: upload with `POST /v1/files`, call
`/v1/generate_captions` with the returned file id, use `stream=true` for SSE,
then delete the file to release storage.
- RTSP live captioning: when the caller provides `RTSP_SAMPLE_URL`, use that
exact URL and run the **RTSP Sample Stream Guard** before registration. Do not
derive a replacement stream from NvStreamer or VIOS when `RTSP_SAMPLE_URL` is
empty; fail fast instead. Require an actual video stream/caps entry before
registration; add the stream, caption it, then unregister it.
- Alert prompts: include a deterministic `Anomaly Detected: Yes/No` line.
Kafka publication is server-side config, additive to HTTP responses, and
documented in [`references/kafka-workflows.md`](references/kafka-workflows.md).
- Kafka validation: trust the live `vss-rtvi-vlm` environment for topic names.
In a full VSS alerts real-time profile, use the existing VSS Kafka container
`mdx-kafka` for CLI checks and final incident-consumer commands. For
standalone validation, use a broker that advertises `${HOST_IP}:9092`; never
stop or replace a pre-existing broker without user confirmation.
## Error Reference
Common causes: 400 for invalid request shape or model id, 401/403 for missing
or wrong bearer token, 404 for deleted files/streams or unsupported endpoints,
413 for oversized uploads, 422 for schema validation, 429 for too much
concurrency, 500 for inference/runtime failures, and 503 while startup is still
in progress. Inspect `docker logs vss-rtvi-vlm` for service-side failures.
安裝 vss-deploy-dense-captioning
請將技能檔案下載並解壓縮至您的 .claude/skills/ 目錄中。
下載 ZIP複製儲存庫並將技能檔案複製到您的專案中。
git clone https://github.com/NVIDIA/skills/tree/main/skills/vss-deploy-dense-captioning # Copy SKILL.md to your .claude/skills/ directory
複製





首頁
