オプション
家家 Skill DevOps と CI/CD vss-deploy-dense-captioning

vss-deploy-dense-captioning

NVIDIA/skills NVIDIA/skills

スタンドアロンの RT-VLM 高密度キャプション生成マイクロサービスをデプロイし、ファイルのアップロード、キャプション生成、ストリーミング、チャット補完、および Kafka との統合に関する REST API エンドポイントを操作します。

...すべて拡張します
2
更新された時間 2026年9月27日

目的

RT-VLMの高密度キャプション生成マイクロサービスを単独で立ち上げて、公開されているすべてのエンドポイント(ファイルアップロード、generate_captions、ストリームの追加/削除、チャット補完、Kafkaトピック)をテストする。

前提条件

RT-VLMのスタンドアロン展開には以下が必要です:

  • Docker、Docker Compose、NVIDIA Container Toolkit、および利用可能なGPU。
  • docker login nvcr.io、 イメージのプル、およびローカルでの NGC モデル/アーティファクトのダウンロードを行うための、$NGC_CLI_API_KEYに設定された NGC レジストリの認証情報。
  • curl、jq、およびスタンドアロン版 Compose コピー用の書き込み可能な作業ディレクトリ。

既存のサービスに対する API 呼び出しの場合:

  • $BASE_URL からアクセス可能な RT-VLM サービスが稼働していること。
  • サービスの設定方法に応じて、$RTVI_VLM_API_KEYまたは$NGC_CLI_API_KEY にベアラー・トークンが格納されていること。 完全な VSS プロファイルのデプロイの場合:

完全な VSS プロファイルのデプロイの場合:

  • ../vss-deploy-profile/SKILL.md を使用してください。このスキルでは、完全な VSS プロファイルはデプロイされません。

手順

以下のルーティングテーブルおよびステップバイステップのワークフローに従ってください。「workflow」、「quick start」、または「flow」で終わる各セクションは、上から下へと順に実行することを想定しています。詳細なリファレンス資料はreferences/ ディレクトリにあります。将来の改訂で具体的なヘルパーが指定されない限り、文書化されたワークフローを直接実行してください。

例

エンドツーエンドで動作するサンプルは、evals/ディレクトリ(各*.jsonマニフェストには実行可能なシナリオが含まれています)および以下のワークフローごとのcurlブロック内に記載されています。これらを再現するには、nv-base validate --agent-evalを使用して Tier-3 評価を実行してください。

制限事項

  • このスキルを通じてデプロイされたスタンドアロンの RT-VLM サービス、または呼び出し元からアクセス可能な 既存の RT-VLM サービスのいずれかが必要です。
  • NGCでホストされているモデルやNIMには、レート制限、GPUメモリ要件、およびライセンス制限が適用される場合があります。
  • 同時実行数、GPU メモリ、およびストレージの制限は、ホストのハードウェアおよびプロファイルの compose ファイルによって異なります。
  • NGC_CLI_API_KEY、RTVI_VLM_API_KEY、および.envファイルは、git やログに含めないでください。また、認証情報の値をエコーしたり、最終的な応答に含めたりしないでください。
  • Docker グループへのアクセスおよびsudo は、事実上 root レベルの権限となります。パスワードレス sudo が利用できない場合は、デプロイのリファレンスで非対話型のsudo -nガードを使用し、ホスト所有者のアクションについては停止してください。

トラブルシューティング

  • エラー: REST 呼び出しで「接続拒否」が返されました。原因: 対象のマイクロサービスが実行されていません。解決策:/docsまたは/health をプローブし、vss-deploy-profileまたは対応するvss-deploy-*スキルを使用して再デプロイしてください。
  • エラー:NGCからのプルでHTTP 401/403が発生。原因:NGC_CLI_API_KEYの欠落または有効期限切れ。解決策:docker login nvcr.ioを実行し、キーを再エクスポートしてから再試行してください。
  • エラー:コンテナの OOM またはモデルの読み込みに失敗しました。原因:選択したプロファイルに対して GPU メモリが不足しています。解決策:より小さいバリアントに切り替えるか、`docker compose down` を使用して GPU を解放してください。

RT-VLM デンセ・キャプション(VSS 3.2)のデプロイと使用

RT-VLM は NVIDIA のリアルタイムビジョン・言語マイクロサービスです。動画(ファイルまたは RTSP)をデコードし、チャンクに分割して、VLM(cosmos-reason1、cosmos-reason2、または任意の OpenAI互換モデル)を実行し、SSE/HTTP経由でdenseキャプションをストリーミングし、 キャプション、インシデントアラート、エラーをKafkaにパブリッシュします。このスキルを使用して、 完全なVSSプロファイルがまだ実行されていない場合にスタンドアロンのRT-VLMサービスをデプロイし、その後、 その/v1/...APIを呼び出して、キャプション生成、ファイルアップロード、ライブストリーム管理、ヘルス チェック、NIM互換のチャット補完、またはPrometheusメトリクスを実行します。APIリファレンス: https://docs.nvidia.com/vss/latest/real-time-vlm-api.html。

デプロイのルーティング

ユーザーが完全なVSSプロファイルのデプロイを要求した場合は、 ../vss-deploy-profile/SKILL.md を使用してください。このスキルは、 プロファイルのルーティング、generated.env、resolved.yml、マルチサービスのサイジング、および フルスタックのデプロイ/ティアダウンを管理します。

ユーザーがスタンドアロンの RT-VLM 高密度キャプション処理を要求した場合、または VSS プロファイルが まだ実行されていない場合は、API を呼び出す前に references/deploy-rt-vlm-service.md にあるスタンドアロン RT-VLM フローを使用してください。 これは、 vss-deploy-profile と同じ Compose 中心のパターンに従います。つまり、コンテキストの収集、プリフライトの実行、ローカルコピーでの作業、 Docker Compose 設定によるドライラン、レビュー、デプロイ、そしてヘルスステータスの確認を待ちます。

スタンドアロン展開フロー

必ずこの順序に従ってください。ドライランをスキップしてはいけません。

# 1. deploy/docker/services/rtvi/rtvi-vlm/rtvi-vlm-docker-compose.yml
#    を、書き込み可能な任意のスタンドアロン用作業ディレクトリにコピーします。
# 2. その Compose ファイルのコピーから RTVI_VLM_IMAGE_TAG を導出します。
# 3. コピーから、スタンドアロン専用で不要な `depends_on` ブロックを削除する。
# 4. 必要な RT-VLM 値を含む、Git で無視される `.env` ファイルを作成する。
# 5. $VSS_DATA_DIR/data_log/vst/clip_storage などのホストのバインドパスを準備する。
#    所有権の修正には `sudo -n` を使用します。パスワード不要の sudo が利用できない場合は、
#    処理を中止し、ホストの所有者に表示されたコマンドを手動で実行するよう依頼してください。
# 6. docker compose --env-file .env -f rtvi-vlm-docker-compose.yml config --quiet
# 7. 正確な RT-VLM イメージタグを `docker pull` で取得します。
# 8. `docker compose ... up -d rtvi-vlm` を実行し、「ready」状態になるのを待ってから、スモークテストを行います。

pull やup を実行する前にプリフライトを実行し、RT-VLM 本体のデバッグに入る前に、 ここで失敗箇所を特定して修正する:

nvidia-smi --query-gpu=index,name --format=csv,noheader
nvidia-container-cli info
docker compose version
docker run --rm --gpus all nvidia/cuda:12.4.0-base-ubuntu22.04 nvidia-smi

スタンドアロンの単一ファイル展開の場合、 deploy/docker/services/rtvi/rtvi-vlm/rtvi-vlm-docker-compose.yml を直接実行しないでください。このファイルには、 VSS/met-blueprints の完全な Compose プロジェクト内でのみ定義されている、同階層の VLM/NIM サービスへの 完全な VSS/met-blueprints Compose プロジェクト内でのみ定義されています。スタンドアロンのリファレンスでは、 Compose ファイルのコピー方法、そこから現在のイメージタグを導出する方法、 `depends_on` ブロックの削除方法、および`up` を実行する前に結果を検証する方法を示しています。

エージェント主導の検証を行う場合、sudoプロンプトを対話モードで表示させてはいけません。特権所有権の取得や Docker 操作を行う前に、 references/deploy-rt-vlm-service.md に記載されている非対話型のガードを使用してください: plain docker を優先してください。そうでない場合は sudo -n docker を使用してください。sudo -n が失敗した場合は、 plaindocker を優先し、それが不可能な場合はsudo -n docker を使用してください。sudo -n が失敗した場合は、 対話型の sudo で再試行したり権限を弱めたりするのではなく、ホスト所有者向けの正確な手動コマンドを実行して 処理を中止してください。

Docker 28 以降でdocker pull がcontainerd のスナップショッター/アンパックエラーにより失敗した場合は、 再試行する前に、スタンドアロンリファレンス内の/etc/docker/daemon.json に containerd-snapshotter=falseを適用してください。

スタンドアロン環境における.envの最小設定値:

ホストの環境変数 必要な場合 目的
NGC_CLI_API_KEY スタンドアロン展開パス NGC レジストリのイメージ取得および NGC モデル/アーティファクトのダウンロード
RTVI_VLM_API_KEYまたはNGC_CLI_API_KEY 認証済みAPI呼び出し サービス実行後の RT-VLM ベアラー認証
RTVI_VLM_PORT 常に コンテナ8000にマッピングされたホスト API ポート
HOST_IP 常に Kafka ブートストラップホスト (${HOST_IP}:9092)
VSS_DATA_DIR 常に 必須のクリップストレージバインドマウント
RTVI_VLM_MODEL_TO_USE スタンドアロンの場合は常に バックエンドセレクタ。デフォルトのローカルモデルにはcosmos-reason2 を、リモートまたは同階層のエンドポイントにはopenai-compatを使用
RTVI_VLM_MODEL_PATH ローカルのセルフホスト型モデル ソースベースの Cosmos Reason 2 パス:ngc:nim/nvidia/cosmos-reason2-8b:hf-1208
RTVI_VLM_ENDPOINT RTVI_VLM_MODEL_TO_USE=openai-compat リモート/同階層の OpenAI 互換 VLM エンドポイント
VLM_NAME RTVI_VLM_MODEL_TO_USE=openai-compat そのエンドポイントによって公開されるモデル/デプロイメント名

セットアップ

export BASE_URL="http://localhost:${RTVI_VLM_PORT:-8018}"  # ホスト側の RT-VLM ポート
export API_KEY="${NGC_CLI_API_KEY:-${RTVI_VLM_API_KEY:-}}" # ホスト側の curl コマンドで使用されるベアラートークン
: "${API_KEY:?認証が必要なエンドポイントを呼び出す前に、NGC_CLI_API_KEY または RTVI_VLM_API_KEY を設定してください}"

以下のすべてのリクエストでは、Authorization: Bearer $API_KEY を使用します。ヘルスエンドポイント (/v1/health/*,/v1/ready,/v1/live,/v1/startup) は通常、認証なしで動作します。

使用前の動作確認:

curl -fsS "$BASE_URL/v1/health/ready"
MODEL_ID="$(curl -fsS "$BASE_URL/v1/models" -H "Authorization: Bearer $API_KEY" | jq -r '.data[0].id // .id')"
curl -fsS "$BASE_URL/openapi.json" | jq -r '.paths | keys[]' | sort

RTSP サンプル・ストリーム・ガード

タスクまたは評価でRTSP_SAMPLE_URL が指定されている場合、その環境変数の 値を必須の入力として扱います。 ストリームのプローブや 登録を行う前に、この変数が設定されており、空ではないことを確認してください。設定されていない場合は、明確な失敗メッセージを表示して処理を中止してください。 NvStreamer、VIOS、サンプルデータバンドル、またはその他の フォールバックから代替値を導出しないでください。そうすると、呼び出し元が要求したストリームとは異なるストリームが検証されてしまうためです。

: "${RTSP_SAMPLE_URL:?RTSP 検証の前に、RTSP_SAMPLE_URL を到達可能な RTSP サンプルストリームに設定してください}"
case "$RTSP_SAMPLE_URL" in
  rtsp://*) ;;
  *) echo "RTSP_SAMPLE_URL は rtsp:// で始まる URL でなければなりません。取得した値: $RTSP_SAMPLE_URL" >&2; exit 1 ;;
esac

if command -v ffprobe >/dev/null 2>&1; then
  ffprobe -v error -rtsp_transport tcp \
    -select_streams v:0 -show_entries stream=codec_type \
    -of csv=p=0 "$RTSP_SAMPLE_URL" | grep -qx video
elif command -v gst-discoverer-1.0 >/dev/null 2>&1; then
  gst-discoverer-1.0 "$RTSP_SAMPLE_URL" | grep -qi 'video'
else
  echo "RTSPの検証を行う前に、ffprobeまたはgst-discoverer-1.0をインストールしてください。" >&2
  exit 1
fi

クイックスタート — ローカル動画からの高密度キャプション

# 1. 動画をアップロードし、ファイル ID を取得する
FILE_ID=$(curl -fsS -X POST "$BASE_URL/v1/files" \
  -H "Authorization: Bearer $API_KEY" \
  -F "file=@/path/to/warehouse.mp4" \
  -F "purpose=vision" \
  -F "media_type=video" | jq -r '.id')

# 2. キャプションとアラートを生成する(チャンク化されたレスポンスの SSE ストリーム)
curl -N -X POST "$BASE_URL/v1/generate_captions" \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d "{
    \"id\": \"$FILE_ID\",
    \"prompt\": \"この倉庫の動画の10秒ごとのセグメントごとに、簡潔で要点をまとめたキャプションを作成してください。\",
    \"model\": \"$MODEL_ID\",
    \"chunk_duration\": 10,
    \"stream\": true
  }"

API インターフェース

オプションのエンドポイントを呼び出す前に、最新の OpenAPI を信頼できる情報源として使用してください:

curl -fsS "$BASE_URL/openapi.json" | jq -r '.paths | keys[]' | sort

VSS 3.2 の主要なパスは以下の通りです:

  • マルチパートメディアのアップロードにはPOST /v1/files を使用します。返されたファイルIDを キャプション生成に渡して、完了後にファイルを削除してください。
  • ファイルまたはストリームのキャプション生成にはPOST /v1/generate_captionsを使用します。GET /v1/models によって返される正確な モデル ID を使用してください。cosmos-reason2などのエイリアスは バックエンドセレクタであり、リクエストのモデル ID ではありません。
  • RTSPライフサイクル用のPOST /v1/streams/add、GET /v1/streams/get-stream-info、および DELETE /v1/streams/delete/{stream_id}。ストリームIDは results[0].id から取得してください。
  • OpenAI互換のテキストおよびマルチモーダル呼び出しには、POST /v1/chat/completions を使用します。 現在の 26.05 ビルドでは、テキストのみの/v1/completions に対して HTTP 400 が返されます。 レガシー動作を検証する際は、これを想定内として扱ってください。
  • サービスプローブには、 GET /v1/health/ready、/v1/models、/v1/assets/stats、および/v1/metrics を使用します。OpenAPIに明記されていない限り、/v1/licenseが存在すると仮定しないでください。

詳細なエンドポイントスキーマ、レスポンス形状、CVスタイルの単一ストリームエンドポイント、 および26.05互換性に関する注意事項は、 references/api-surface-26.05.mdに記載されています。

一般的なワークフロー

  • 保存済みファイルのキャプション作成:POST /v1/files でアップロードし、 返されたファイル ID を使用して/v1/generate_captions を呼び出し、SSE の場合はstream=true を指定し、 その後、ファイルを削除してストレージを解放します。
  • RTSPライブキャプション:呼び出し元がRTSP_SAMPLE_URLを指定した場合は、その URLをそのまま使用し、登録前にRTSPサンプルストリームガードを実行してください。RTSP_SAMPLE_URLが 空の場合、NvStreamerやVIOSから代替ストリームを 生成しないでください。代わりに、速やかに処理を中止してください。登録前に実際のビデオストリーム/キャプションのエントリを 必須とし、ストリームを追加してキャプションを付与した後、登録を解除してください。
  • アラートプロンプト:「異常が検出されました:はい/いいえ」という確定的な行を含める。 Kafkaへのパブリッシュはサーバーサイドの設定であり、HTTPレスポンスに追加されるもので、 references/kafka-workflows.mdにドキュメント化されている。
  • Kafkaの検証:トピック名については、稼働中のvss-rtvi-vlm環境を信頼する。 完全なVSSアラートリアルタイムプロファイルでは、CLIチェックおよび最終的なインシデントコンシューマーコマンド用に、既存のVSS Kafkaコンテナ mdx-kafkaを使用してください。スタンドアロンの 検証では、${HOST_IP}:9092 をアドバタイズするブローカーを使用してください。ユーザーの確認なしに、既存のブローカーを 停止または置き換えないでください。

エラー参照

一般的な原因:400 はリクエスト形式またはモデル ID が無効な場合、401/403 は ベアラートークンの欠落または誤り、404 はファイル/ストリームが削除されているか、サポートされていないエンドポイントの場合、 413:アップロードサイズ超過、422:スキーマ検証エラー、429:同時実行数が多すぎる 場合、500:推論/実行時の失敗、503:起動がまだ 進行中の場合。サービス側の障害については、Dockerのvss-rtvi-vlmログを確認してください。

GitHubで見る
---
name: vss-deploy-dense-captioning
description: Deploy a standalone RT-VLM dense-captioning microservice and exercise its REST API endpoints for file upload, caption generation, streaming, chat completions, and Kafka integration.
license: Apache-2.0
---
## Purpose

Stand up the RT-VLM dense-captioning microservice on its own and exercise every endpoint it exposes (file upload, generate_captions, stream add/delete, chat-completions, Kafka topics).

## Prerequisites

For standalone RT-VLM deployment:
- Docker, Docker Compose, NVIDIA Container Toolkit, and a visible GPU.
- NGC registry credentials in `$NGC_CLI_API_KEY` for `docker login nvcr.io`,
  image pulls, and local NGC model/artifact downloads.
- `curl`, `jq`, and any writable working directory for the standalone compose copy.

For API calls against an existing service:
- Running RT-VLM service reachable at `$BASE_URL`.
- Bearer token in `$RTVI_VLM_API_KEY` or `$NGC_CLI_API_KEY`, depending on how the
  service was configured.

For full VSS profile deployment:
- Use `../vss-deploy-profile/SKILL.md`; this skill does not deploy full VSS profiles.

## Instructions

Follow the routing tables and step-by-step workflows below. Each section that ends in *workflow*, *quick start*, or *flow* is intended to be executed top-to-bottom. Detailed reference material lives in `references/`; execute the documented workflows directly unless a future revision names a concrete helper.

## Examples

Worked end-to-end examples are kept under `evals/` (each `*.json` manifest contains a runnable scenario) and inline in the per-workflow `curl` blocks below. Run a Tier-3 evaluation with `nv-base validate <this-skill-dir> --agent-eval` to replay them.

## Limitations

- Requires either a standalone RT-VLM service deployed via this skill or an
  existing RT-VLM service reachable from the caller.
- NGC-hosted models and NIMs may be subject to rate-limits, GPU memory requirements, and license restrictions.
- Concurrency, GPU memory, and storage limits depend on the host hardware and the profile's compose file.
- Keep `NGC_CLI_API_KEY`, `RTVI_VLM_API_KEY`, and `.env` files out of git and out of logs; do not echo credential values or include them in final responses.
- Docker group access and `sudo` are effectively root-level privileges. Use the non-interactive `sudo -n` guard in the deploy reference and stop for host-owner action when passwordless sudo is unavailable.

## Troubleshooting

- **Error**: REST call returns connection refused. **Cause**: target microservice not running. **Solution**: probe `/docs` or `/health`; redeploy via `vss-deploy-profile` or the matching `vss-deploy-*` skill.
- **Error**: HTTP 401/403 from NGC pulls. **Cause**: missing/expired `NGC_CLI_API_KEY`. **Solution**: `docker login nvcr.io` and re-export the key before retrying.
- **Error**: container OOM or model fails to load. **Cause**: insufficient GPU memory for the selected profile. **Solution**: switch to a smaller variant or free GPUs via `docker compose down`.

# Deploy and Use RT-VLM Dense Captioning (VSS 3.2)

RT-VLM is NVIDIA's real-time vision-language microservice: decode video (file or
RTSP), segment it into chunks, run a VLM (`cosmos-reason1`, `cosmos-reason2`, or any
OpenAI-compatible model), stream dense captions back over SSE/HTTP, and publish
captions, incident alerts, and errors to Kafka. Use this skill to deploy the
standalone RT-VLM service when a full VSS profile is not already running, then call
its `/v1/...` API for caption generation, file upload, live-stream management, health
checks, NIM-compatible chat completions, or Prometheus metrics. API reference:
<https://docs.nvidia.com/vss/latest/real-time-vlm-api.html>.

## Deployment Routing

If the user asks to deploy a full VSS profile, use
[`../vss-deploy-profile/SKILL.md`](../vss-deploy-profile/SKILL.md). That skill
owns profile routing, `generated.env`, `resolved.yml`, multi-service sizing, and
full-stack deploy/teardown.

If the user asks for standalone RT-VLM dense captioning, or no VSS profile is
already running, use the standalone RT-VLM flow in
[`references/deploy-rt-vlm-service.md`](references/deploy-rt-vlm-service.md)
before calling the API. This follows the same compose-centric pattern as
`vss-deploy-profile`: gather context, run preflights, work from a local copy,
dry-run with `docker compose config`, review, deploy, then wait for health.

## Standalone Deployment Flow

Always follow this sequence. Never skip the dry-run.

```bash
# 1. Copy deploy/docker/services/rtvi/rtvi-vlm/rtvi-vlm-docker-compose.yml
#    into any writable standalone working directory.
# 2. Derive RTVI_VLM_IMAGE_TAG from that compose copy.
# 3. Strip the standalone-only dangling depends_on block from the copy.
# 4. Create a gitignored .env with the required RT-VLM values.
# 5. Prepare host bind paths such as $VSS_DATA_DIR/data_log/vst/clip_storage.
#    Use `sudo -n` for ownership fixes; if passwordless sudo is unavailable,
#    stop and ask the host owner to run the printed command manually.
# 6. docker compose --env-file .env -f rtvi-vlm-docker-compose.yml config --quiet
# 7. docker pull the exact RT-VLM image tag.
# 8. docker compose ... up -d rtvi-vlm, wait for ready, then smoke test.
```

Run preflights before any pull or `up`; stop and fix failures here before
debugging RT-VLM itself:

```bash
nvidia-smi --query-gpu=index,name --format=csv,noheader
nvidia-container-cli info
docker compose version
docker run --rm --gpus all nvidia/cuda:12.4.0-base-ubuntu22.04 nvidia-smi
```

For standalone single-file deployments, do not run the raw
`deploy/docker/services/rtvi/rtvi-vlm/rtvi-vlm-docker-compose.yml` directly: it
contains `depends_on` references to sibling VLM/NIM services that are only
defined in the full VSS/met-blueprints compose project. The standalone reference
shows how to copy the compose file, derive the current image tag from it, strip
the `depends_on` block, and validate the result before `up`.

For agent-driven validation, never let `sudo` prompt interactively. Before any
privileged ownership or Docker operation, use the non-interactive guard in
[`references/deploy-rt-vlm-service.md`](references/deploy-rt-vlm-service.md):
prefer plain `docker`; otherwise use `sudo -n docker`; if `sudo -n` fails, stop
with the exact manual command for the host owner instead of retrying with
interactive sudo or weakening permissions.

If `docker pull` fails with a containerd snapshotter/unpack error on Docker 28+,
apply the `/etc/docker/daemon.json` `containerd-snapshotter=false` fix in the
standalone reference before retrying.

Minimum standalone `.env` values:

| Host env var | Required when | Purpose |
|---|---|---|
| `NGC_CLI_API_KEY` | Standalone deploy path | NGC registry image pull and NGC model/artifact download |
| `RTVI_VLM_API_KEY` or `NGC_CLI_API_KEY` | Authenticated API calls | RT-VLM bearer auth after the service is running |
| `RTVI_VLM_PORT` | Always | Host API port mapped to container `8000` |
| `HOST_IP` | Always | Kafka bootstrap host (`${HOST_IP}:9092`) |
| `VSS_DATA_DIR` | Always | Required clip-storage bind mount |
| `RTVI_VLM_MODEL_TO_USE` | Always for standalone | Backend selector; use `cosmos-reason2` for the default local model or `openai-compat` for a remote/sibling endpoint |
| `RTVI_VLM_MODEL_PATH` | Local self-hosted model | Source-backed Cosmos Reason 2 path: `ngc:nim/nvidia/cosmos-reason2-8b:hf-1208` |
| `RTVI_VLM_ENDPOINT` | `RTVI_VLM_MODEL_TO_USE=openai-compat` | Remote/sibling OpenAI-compatible VLM endpoint |
| `VLM_NAME` | `RTVI_VLM_MODEL_TO_USE=openai-compat` | Model/deployment name exposed by that endpoint |

## Setup

```bash
export BASE_URL="http://localhost:${RTVI_VLM_PORT:-8018}"  # host-side RT-VLM port
export API_KEY="${NGC_CLI_API_KEY:-${RTVI_VLM_API_KEY:-}}" # bearer token used by host-side curl commands
: "${API_KEY:?Set NGC_CLI_API_KEY or RTVI_VLM_API_KEY before calling authenticated endpoints}"
```

Every request below uses `Authorization: Bearer $API_KEY`. Health endpoints
(`/v1/health/*`, `/v1/ready`, `/v1/live`, `/v1/startup`) typically work without auth.

**Smoke test before use:**
```bash
curl -fsS "$BASE_URL/v1/health/ready"
MODEL_ID="$(curl -fsS "$BASE_URL/v1/models" -H "Authorization: Bearer $API_KEY" | jq -r '.data[0].id // .id')"
curl -fsS "$BASE_URL/openapi.json" | jq -r '.paths | keys[]' | sort
```

## RTSP Sample Stream Guard

When a task or eval names `RTSP_SAMPLE_URL`, treat that exact environment
variable as a required input. Verify it is set and non-empty before probing or
registering any stream; if it is missing, stop with a clear failure message. Do
not derive a substitute from NvStreamer, VIOS, sample-data bundles, or any other
fallback, because that validates a different stream than the caller requested.

```bash
: "${RTSP_SAMPLE_URL:?Set RTSP_SAMPLE_URL to a reachable RTSP sample stream before RTSP validation}"
case "$RTSP_SAMPLE_URL" in
  rtsp://*) ;;
  *) echo "RTSP_SAMPLE_URL must be an rtsp:// URL, got: $RTSP_SAMPLE_URL" >&2; exit 1 ;;
esac

if command -v ffprobe >/dev/null 2>&1; then
  ffprobe -v error -rtsp_transport tcp \
    -select_streams v:0 -show_entries stream=codec_type \
    -of csv=p=0 "$RTSP_SAMPLE_URL" | grep -qx video
elif command -v gst-discoverer-1.0 >/dev/null 2>&1; then
  gst-discoverer-1.0 "$RTSP_SAMPLE_URL" | grep -qi 'video'
else
  echo "Install ffprobe or gst-discoverer-1.0 before RTSP validation." >&2
  exit 1
fi
```

## Quick Start — dense captions from a local video

```bash
# 1. Upload the video, capture its file id
FILE_ID=$(curl -fsS -X POST "$BASE_URL/v1/files" \
  -H "Authorization: Bearer $API_KEY" \
  -F "file=@/path/to/warehouse.mp4" \
  -F "purpose=vision" \
  -F "media_type=video" | jq -r '.id')

# 2. Generate captions + alerts (SSE stream of chunked responses)
curl -N -X POST "$BASE_URL/v1/generate_captions" \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d "{
    \"id\": \"$FILE_ID\",
    \"prompt\": \"Write a concise dense caption for each 10-second segment of this warehouse video.\",
    \"model\": \"$MODEL_ID\",
    \"chunk_duration\": 10,
    \"stream\": true
  }"
```

## API Surface

Use the live OpenAPI as the source of truth before calling optional endpoints:

```bash
curl -fsS "$BASE_URL/openapi.json" | jq -r '.paths | keys[]' | sort
```

Core paths for VSS 3.2 are:

- `POST /v1/files` for multipart media upload; pass the returned file `id` into
  caption generation and delete the file when finished.
- `POST /v1/generate_captions` for file or stream captioning. Use the exact
  model id returned by `GET /v1/models`; aliases such as `cosmos-reason2` are
  backend selectors, not request model ids.
- `POST /v1/streams/add`, `GET /v1/streams/get-stream-info`, and
  `DELETE /v1/streams/delete/{stream_id}` for RTSP lifecycle. Parse stream ids
  from `results[0].id`.
- `POST /v1/chat/completions` for OpenAI-compatible text and multimodal calls.
  Current 26.05 builds return HTTP 400 for text-only `/v1/completions`; treat
  that as expected when validating legacy behavior.
- `GET /v1/health/ready`, `/v1/models`, `/v1/assets/stats`, and `/v1/metrics`
  for service probes. Do not assume `/v1/license` exists unless OpenAPI lists it.

Detailed endpoint schemas, response shapes, CV-style singular stream endpoints,
and 26.05 compatibility notes live in
[`references/api-surface-26.05.md`](references/api-surface-26.05.md).

## Common Workflows

- Stored file captioning: upload with `POST /v1/files`, call
  `/v1/generate_captions` with the returned file id, use `stream=true` for SSE,
  then delete the file to release storage.
- RTSP live captioning: when the caller provides `RTSP_SAMPLE_URL`, use that
  exact URL and run the **RTSP Sample Stream Guard** before registration. Do not
  derive a replacement stream from NvStreamer or VIOS when `RTSP_SAMPLE_URL` is
  empty; fail fast instead. Require an actual video stream/caps entry before
  registration; add the stream, caption it, then unregister it.
- Alert prompts: include a deterministic `Anomaly Detected: Yes/No` line.
  Kafka publication is server-side config, additive to HTTP responses, and
  documented in [`references/kafka-workflows.md`](references/kafka-workflows.md).
- Kafka validation: trust the live `vss-rtvi-vlm` environment for topic names.
  In a full VSS alerts real-time profile, use the existing VSS Kafka container
  `mdx-kafka` for CLI checks and final incident-consumer commands. For
  standalone validation, use a broker that advertises `${HOST_IP}:9092`; never
  stop or replace a pre-existing broker without user confirmation.

## Error Reference

Common causes: 400 for invalid request shape or model id, 401/403 for missing
or wrong bearer token, 404 for deleted files/streams or unsupported endpoints,
413 for oversized uploads, 422 for schema validation, 429 for too much
concurrency, 500 for inference/runtime failures, and 503 while startup is still
in progress. Inspect `docker logs vss-rtvi-vlm` for service-side failures.

vss-deploy-dense-captioningをインストール

スキルファイルをダウンロードし、.claude/skills/ ディレクトリに解凍してください。

ZIPをダウンロード

リポジトリをクローンし、スキルファイルをプロジェクトにコピーしてください。

git clone https://github.com/NVIDIA/skills/tree/main/skills/vss-deploy-dense-captioning # Copy SKILL.md to your .claude/skills/ directory

コピー コピー
クイックセットアップ: スキルフォルダを .claude/skills/ にコピーしてください。 Claude が自動的にそのスキルを検出して使用します。
リポジトリ NVIDIA/skills

関連スキル

klingai-upgrade-migration
更新された時間 2026年7月3日
Verification &amp; Quality Assurance
更新された時間 2026年6月29日
base44-cli
更新された時間 2026年6月29日
Railway CLI Management
更新された時間 2026年7月2日
OR