vss-generate-video-report
NVIDIA/skills
VLMバックエンドへのルーティングによるクリップごとの分析、またはインシデント範囲レポート用の分析バックエンドへのルーティングを通じて、動画分析レポートを生成します。これには、デプロイメントプロファイルの検証およびURL書き換えが含まれます。
...すべて拡張しますレポート
2つのバックエンドのいずれかにルーティングすることで動画分析レポートを生成します。VSSエージェントのPOST /generate を経由することは決してありません。
| モード | バックエンド |
|---|---|
| A. ビデオクリップ | /vss-manage-video-io-storage→ クリップURL →VLMチャット/補完機能 |
| B. インシデント範囲 | /vss-query-analytics→ インシデント一覧 → ナラティブレポート |
リクエストが曖昧な場合(例:「 」といった、時間範囲やインシデントの文言が指定されていない場合)、デフォルトではモードAとなります。ユーザーがセンサーと時間範囲の両方を指定している場合にのみ、確認を行ってください。各モードにルーティングされるリクエストの表現については、以下の例を参照してください。
手順
- モードを選択します。記録されたクリップまたはセンサー映像が 1 つの場合はモード A、リクエストに時間範囲またはインシデント/アラートが指定されている場合はモード B を選択します(「例」と照合してください)。
- 「導入の前提条件」でそのモードの導入プロファイルを確認し、プローブに失敗した場合は
/vss-deploy-profileに引き渡します。 - そのモードの番号付き手順を実行します— 以下のモード Aまたはモード B。
- レポートに埋め込む前に、ユーザー向けのすべてのクリップ URL を
$VSS_PUBLIC_HOST:$VSS_PUBLIC_PORTという 1 行の形式(ブラウザで再生可能なクリップ URL)に書き換えます。 - レンダリングされたレポートのマークダウンをユーザーに返す。
評価者向けの契約書:
- モード A のトップタイトルは、必ず「
# ビデオ分析レポート」でなければなりません。 - モード B の最上部のタイトルは、必ず「
# インシデント範囲レポート」でなければなりません(「# インシデントレポート」やセンサー名のバリエーションは絶対に使用しないでください)。 - モード B には、テンプレートで指定された正確な行(レポート識別子、範囲、スコープ、インシデント総数、確認済み/却下/未確認)を含む
「## 基本情報」を必ず含める必要があります。
例
- 「この動画に関するレポートを生成する」/「…に関するレポート」
」 →モードA - 「warehouse_01.mp4 を分析する」/「アップロードされた動画の分析レポートを作成する」→モード A
- 「12:31Z から 12:32Z までのインシデントに関するレポート」 →モード B
- 「今日のアラートに関するレポート」 / 「
」 →モードB - 「アラートの概要」
~の間のからまで」の間のアラートを要約する」 →モードB
否定トリガー
リクエストが以下のいずれかに該当する場合は、このスキルを使用しないでください:
- レポートを明示的に求めていないクリップに関するその場限りの視覚的Q&A(「トラックの色は何色?」「00:12で何が起きている?」) →
/vss-ask-videoを使用してください。 - アーカイブ/意味的類似性による検索(「フォークリフトを探す」、「すべての動画で車間距離が詰められている場面を検索する」) →
/vss-search-archiveを使用してください。 - レポートの生成を必要としない、読み取り専用のインシデント/メトリクスの照会 →
/vss-query-analyticsを使用してください。 - デプロイ/テアダウン/プロファイルの変更(「アラートのデプロイ」、「プロファイルの切り替え」、「ベースの起動」) →
/vss-deploy-profileを使用してください。 - リアルタイムのアラート/ルール管理リクエスト →
/vss-manage-alertsを使用してください。
レポートを VSS-agentPOST /generate 経由でルーティングしないでください。
デプロイの前提条件
モード A には、VSSベースプロファイル(VST + VLM NIM)が必要です。 モード Bには、VSSアラートプロファイル(VA-MCP + Elasticsearch)が必要です。
プローブ:
# モード A — VST + VLM の到達可能性
curl -sf --max-time 5 "http://${HOST_IP}:30888/vst/api/v1/sensor/version" >/dev/null
# モード B — VA-MCP
curl -sf --max-time 5 "http://${HOST_IP}:9901/" >/dev/null
プローブが失敗した場合は、-p base(モード A)または-p alerts(モード B)を指定して/vss-deploy-profileに引き継ぎます。必ず事前にユーザーに確認の上、デプロイを行ってください。
クリップURL:VLM入力とブラウザレポートリンク
VSTは、エージェント内部の${HOST_IP}:30888というホスト:ポートを使用してクリップURLを返します。
ローカルまたはクラスタ内のVLMフレーム取得には、その元のURLをVIDEO_URLとして保持してください。
ブラウザで再生可能にするためだけに、VLM入力URLを書き換えないでください。
レンダリングされたレポートに表示されるURLに対してのみ、BROWSER_CLIP_URLを作成してください。
デプロイ層は、ブラウザ向けのホスト:ポートを$VSS_PUBLIC_HOST/
$VSS_PUBLIC_PORT(スキームは$VSS_PUBLIC_HTTP_PROTOCOL)としてエクスポートします。
すべてのプロファイルの.env ファイル(Brev またはベアメタル)で、レポートリンクの書き換えは次のようになります:
: "${VSS_PUBLIC_HOST:?クリップ URL の書き換え前に VSS_PUBLIC_HOST を設定してください}"
: "${VSS_PUBLIC_PORT:?クリップ URL の書き換え前に VSS_PUBLIC_PORT を設定してください}"
VSS_PUBLIC_HTTP_PROTOCOL="${VSS_PUBLIC_HTTP_PROTOCOL:-http}"
BROWSER_CLIP_URL=$(echo "$RAW_URL" | sed -E "s|^https?://[^/]+|${VSS_PUBLIC_HTTP_PROTOCOL}://${VSS_PUBLIC_HOST}:${VSS_PUBLIC_PORT}|")
必要なパブリックホスト値のいずれかが欠落している場合は、レポート向けのクリップ
リンクを省略し、ブラウザで再生可能なURLを生成できなかったことを明記する。ただし、
ローカルVLM分析パスをブロックしてはならない。レンダリングされたレポートに表示されるすべてのクリップURLに
この書き換えを適用する(モードAのステップ4のクリップURL行、モードBの
インシデントごとのクリップのサブ項目)。VLMがローカルまたはクラスタ内にある場合、モードAの
ステップ3にあるVLMのvideo_urlコンテンツブロックは、元の内部URLのままにする。
モード A — 録画済みビデオクリップに関するレポート
VSSlvsプロファイルがデプロイされている場合—curl -sf --max-time 5 "http://${HOST_IP}:38111/v1/ready" がHTTP 200 を返す —/vss-summarize-videoを実行して要約を生成し、 その出力をステップ4のレポートテンプレートに貼り付け、ステップ1~3(VLMダイレクトパス)をスキップします。/v1/readyのレスポンスが200でない場合のみ、ステップ1~3を実行してください。
ステップ 1 — クリップの URL を特定する
/vss-manage-video-io-storageに処理を引き継ぎ、以下の操作を行います:
センサーを一覧表示し、指定された
が存在することを確認します(存在しない場合は先にアップロードします)。ユーザーが
startTime/endTime を指定していない場合、記録された範囲の/storage/取得します。/timelines を クリップの URL をリクエストします:
curl -s "http://${HOST_IP}:30888/vst/api/v1/storage/file//url?startTime= &endTime= &container=mp4&disableAudio=true" | jq -r .videoUrl これにより、ローカルまたはクラスタ内のVLMがフレームを取得できる直接
のmp4URLが得られます。 これをVIDEO_URL(ステップ 3 で VLM が使用する)にバインドし、RAW_URL="$VIDEO_URL"を設定してから、レポートリンクの書き換えを適用してステップ 4 用のBROWSER_CLIP_URL を生成します。ユーザーのブラウザは$VIDEO_URLに直接アクセスできないためです。 モード A では、選択された VLM エンドポイントがVIDEO_URLを取得できる必要があります。 ローカルの NIM/RT-VLM 展開では通常可能ですが、リモートエンドポイントは一般的にlocalhost、プライベートHOST_IP、または VST 内部の URL を取得できません。 ライブのVLM_ENDPOINTがリモートである場合、/v1/modelsのリクエストが成功した後に失敗するチャットリクエストを送信するのではなく、 その到達可能性要件を明示してください。
ステップ 2 — VLM エンドポイントとモデルの解決
デプロイでは、2つのスタックのいずれかを介してVLMを提供している可能性があります。どちらもOpenAI互換のチャット/補完APIを公開しています。稼働している方を選択してください:
| バックエンド | 環境変数 | 一般的なホストエンドポイント | 以下の場合に選択されます |
|---|---|---|---|
| NIM Cosmos | VLM_BASE_URL、VLM_NAME、VLM_MODE、VLM_MODEL_TYPE |
${VLM_BASE_URL}/v1(環境変数には末尾の/v1を含めない。エージェントが自動的に付加する) |
VLM_MODEL_TYPE != rtvi かつ VLM_MODE∈ {local,local_shared,remote}かつ VLM_BASE_URLが空でない場合 |
| RT-VLM Cosmos | RTVI_VLM_BASE_URL、RTVI_VLM_MODEL_TO_USE、VLM_MODEL_TYPE |
${RTVI_VLM_BASE_URL}/v1— 設定されていない場合は、${HOST_IP}から導出されます(アラート用はhttp://${HOST_IP}:8018/v1、http://${HOST_IP}:30082/v1(基本データの場合) |
VLM_MODEL_TYPE = rtvi、またはVLM_MODE=none、またはVLM_BASE_URLが空の場合。また、ウェアハウスの唯一のパス |
実行中のエージェントコンテナからリアルタイムの値を読み取る — 推測しないでください:
docker exec vss-agent sh -lc '
for k in HOST_IP VLM_MODE VLM_MODEL_TYPE VLM_BASE_URL VLM_NAME RTVI_VLM_BASE_URL RTVI_VLM_MODEL_TO_USE; do
v="$(printenv "$k")"
[ -n "$v" ] && printf "%s=%s\n" "$k" "$v"
done
'
vss-agentの環境変数からRTVI_VLM_ENDPOINT を必須としない。いくつかのプロファイルではこれを設定しないため。
選択ルール:
if [ "${VLM_MODEL_TYPE:-}" = "rtvi" ]; then
VLM_BACKEND="rtvlm"
VLM_ENDPOINT="${RTVI_VLM_BASE_URL:+${RTVI_VLM_BASE_URL%/}/v1}"
[ -z "${VLM_ENDPOINT}" ] && VLM_ENDPOINT="http://${HOST_IP}:8018/v1" # アラートのデフォルト
VLM_MODEL="${RTVI_VLM_MODEL_TO_USE}"
elif [ -n "${VLM_BASE_URL}" ] && [ "${VLM_MODE}" != "none" ]; then
VLM_BACKEND="nim_cosmos"
VLM_ENDPOINT="${VLM_BASE_URL%/}/v1"
VLM_MODEL="${VLM_NAME}"
else
VLM_BACKEND="rtvlm"
VLM_ENDPOINT="${RTVI_VLM_BASE_URL:+${RTVI_VLM_BASE_URL%/}/v1}"
[ -z "${VLM_ENDPOINT}" ] && VLM_ENDPOINT="http://${HOST_IP}:30082/v1" # 基本のデフォルト
VLM_MODEL="${RTVI_VLM_MODEL_TO_USE}"
fi
チャットリクエストを送信する前に/v1/models をプローブして、選択したエンドポイントが稼働しており、モデルが読み込まれていることを確認します:
curl -sf --max-time 5 "${VLM_ENDPOINT}/models" | jq -r '.data[].id'
プローブが失敗した場合、または一覧に${VLM_MODEL} が含まれていない場合は、別のバックエンドに切り替えます(またはエラーを通知します。サーバー上に存在しないモデルを黙って選択することは絶対に避けてください)。
ステップ 3 — VLM を直接呼び出す
video_urlコンテンツブロックを含む OpenAI 互換のchat/completionsエンドポイントを使用します。これは、src/vss_agents/tools/video_understanding.pyでvideo_understanding が構築するペイロードの形状およびマルチモーダル設定と同じものです (_build_vlm_messagesおよび Cosmos のbase_vlm.bind(...)呼び出し)を使用します。
フレームサンプリングとビジュアルトークン(ピクセル)の割り当ては、アクティブなプロファイルのlive video_understanding設定と一致している必要があります。 mm_processor_kwargsおよびmedia_io_kwargsを送信して、直接呼び出しがエージェント内のvideo_understandingツールと同じフレームサンプリングとピクセルバジェットを使用するようにします。これらを省略すると、VLM が独自のデフォルト値を適用するため、出力はエージェントパスとは異なるものになります。
PROMPT='動画内で何が起こっているかを、各セグメントまたはイベントごとのタイムスタンプ(クリップ開始からの開始~終了時刻(秒単位))とともに詳細に記述してください。シーン、物体、人物、車両、および注目すべき行動について記述してください。'
# 推論はデフォルトでOFFです — ベースプロファイルの video_understanding 設定(`reasoning: false`)と一致します。
# video_understanding.py は、呼び出し元が上書きしない限り config.reasoning を使用するため、デフォルトでは推論を無効にします。
# ユーザーが明示的に推論を求めた場合にのみ、Cosmos Reason 2 の推論サフィックスを追加する
# (Cosmos Reason 2 以外の VLM の場合は省略)。推論がオフの場合、応答には `reasoning` ブロックは 含まれない 。
if [ "${REASONING:-false}" = "true" ]; then
PROMPT="${PROMPT}
以下の形式で質問に答えてください:
あなたの推論。
` `タグの直後に最終回答を記述してください 。"
fi
# ステップ3を単独で実行する場合、現在の環境/モデルから欠落しているバックエンドを導出します。
[ -z "${VLM_BACKEND:-}" ] && {
if [ "${VLM_MODEL_TYPE:-}" = "rtvi" ]; then
VLM_BACKEND="rtvlm"
elif [[ "${VLM_MODEL:-}" == nvidia/cosmos* ]]; then
VLM_BACKEND="nim_cosmos"
else
VLM_BACKEND="rtvlm"
fi
}
# マルチモーダル設定 — ハードコーディングされた候補ではなく、VSSエージェントの設定ファイルのパスから解決する。
CFG_JSON=$(
docker exec vss-agent python3 -c '
import json, os, yaml
p = os.getenv("VSS_AGENT_CONFIG_FILE")
if not p:
raise SystemExit("vss-agent で VSS_AGENT_CONFIG_FILE が設定されていません")
if not os.path.isabs(p):
p = os.path.join("/vss-agent", p.lstrip("./"))
with open(p, encoding="utf-8") as f:
cfg = yaml.safe_load(f) or {}
vu = (cfg.get("functions", {}) or {}).get("video_understanding", {}) or {}
print(json.dumps({
"max_fps": int(vu.get("max_fps", 2)),
"max_frames": int(vu.get("max_frames", 30)),
"min_pixels": int(vu.get("min_pixels", 3136)),
"max_pixels": int(vu.get("max_pixels", 8388608)),
}))
')
)
[ -n "${CFG_JSON}" ] || { echo "vss-agent から video_understanding の設定を読み込めませんでした"; exit 1; }
jq -e . >/dev/null <<< "${CFG_JSON}" || { echo "vss-agent からの設定 JSON が無効です"; exit 1; }
MAX_FPS="$(jq -r '.max_fps' <<< "${CFG_JSON}")"
MAX_FRAMES="$(jq -r '.max_frames' <<< "${CFG_JSON}")"
MIN_PIXELS="$(jq -r '.min_pixels' <<< "${CFG_JSON}")"
MAX_PIXELS="$(jq -r '.max_pixels' <<< "${CFG_JSON}")"
# num_frames = min(int(clip_seconds) * max_fps, max_frames)、min 1 — video_understanding.py と一致。
# clip_seconds (ステップ1のendTime - startTime) には小数部が含まれる可能性があるため、整数秒に切り捨て — bashの$((...))
# は整数のみを扱い、「15.0」や「1.5」ではエラーとなる。デフォルトは15秒 → MAX_FRAMESを上限とする。
CLIP_SECONDS=$(awk -v s="${CLIP_SECONDS:-15}" 'BEGIN{printf "%d", s}')
NUM_FRAMES=$(( CLIP_SECONDS * MAX_FPS ))
[ "$NUM_FRAMES" -gt "$MAX_FRAMES" ] && NUM_FRAMES=$MAX_FRAMES
[ "$NUM_FRAMES" -lt 1 ] && NUM_FRAMES=1
# Cosmos mm/media コマンドライン引数は、NIM Cosmos パスでのみ適用する。
# RT-VLM モードは独自のサーバーサイド前処理を使用するため、これらのコマンドライン引数を受け取るべきではない。
MM_KWARGS=""
if [ "${VLM_BACKEND}" = "nim_cosmos" ]; then
case "$VLM_MODEL" in
*cosmos-reason2*) MM_KWARGS=", \"mm_processor_kwargs\": {\"size\": {\"shortest_edge\": ${MIN_PIXELS}, \"longest_edge\": ${MAX_PIXELS}}}, \"media_io_kwargs\": {\"video\": {\"num_frames\": ${NUM_FRAMES}}}" ;;
*cosmos*) MM_KWARGS=", \"mm_processor_kwargs\": {\"videos_kwargs\": {\"min_pixels\": ${MIN_PIXELS}, \"max_pixels\": ${MAX_PIXELS}}}, \"media_io_kwargs\": {\"video\": {\"num_frames\": ${NUM_FRAMES}}}" ;;
*) MM_KWARGS="" ;;
esac
fi
curl -s --connect-timeout 5 --max-time 120 -X POST "${VLM_ENDPOINT}/chat/completions" \
-H "Content-Type: application/json" \
-d @- <<EOF | jq -r '.choices[0].message.content'
{
"model": $(jq -Rs . <<< "${VLM_MODEL}"),
"messages": [
{
"role": "user",
"content": [
{"type": "text", "text": $(jq -Rs . <<< "${PROMPT}")},
{"type": "video_url", "video_url": {"url": $(jq -Rs . <<< "${VIDEO_URL}")}}
]
}
],
"max_tokens": 1024,
"temperature": 0.0${MM_KWARGS}
}
EOF
kwargs ブロックはバックエンドに依存します。
nim_cosmosでは、Reason2 バリアント(nvidia/cosmos-reason2*)はmm_processor_kwargs.size{shortest_edge,longest_edge}を使用し、その他の NIM Cosmos バリアント(nvidia/cosmos*)はmm_processor_kwargs.videos_kwargs{min_pixels,max_pixels}; いずれもmedia_io_kwargs.video.num_framesを送信します。rtvlmでは、Cosmos kwargs は送信されません。
VLM が ブロック(Cosmos Reason 推論モード)を返した場合、レポート本文として 以降のテキストのみをレポート本文として保持します。
ステップ 4 — 動画分析レポートテンプレートの作成
assets/video-analysis-report.md をコピーし、すべてのプレースホルダーに入力して、レンダリングされたマークダウンをユーザーに返します。ソースアセットは変更しないでください。レンダリングする前に、BROWSER_CLIP_URL が設定されており、空ではないことを確認してから、 を「Clip URL」行内のその正確な値に置き換えてください。出力にプレースホルダーを残したり、入力済みのセルにテンプレートの指示を含めたり、生のHOST_IP:30888URL を使用したりしてはいけません。
モード B — 指定した期間のインシデントに関するレポート
ステップ 1 — 時間範囲および(オプションで)センサーを特定する
start_time/end_time はISO 8601 UTC 形式(YYYY-MM-DDTHH:MM:SS.sssZ)で指定する必要があります。「過去 1 時間」、「今日」などの相対的な表現は、現在のホストの時計に基づいて解決してください。- ユーザーがセンサー名を指定した場合は、
source+source_type=sensorとして取得します。指定がない場合は、両方を未設定のままにして、すべてのセンサーを対象としたクエリとします。
ステップ 2 —/vss-query-analytics経由でインシデントを取得する
以下の内容で/vss-query-analyticsに引き渡します(初期化 →tools/call):
{
"jsonrpc": "2.0",
"method": "tools/call",
"params": {
"name": "video_analytics__get_incidents",
"arguments": {
"source": "",
"source_type": "sensor",
"start_time": "",
"end_time": "",
"max_count": 100,
"includes": ["objectIds", "info"]
}
},
"id": 1
}
読み取り専用境界(必須):
- モード B は、厳密に読み取り専用の分析データ取得です。Elasticsearch/VA データへの書き込み、シード、バックフィル、または変更は絶対に行わないでください。
- 禁止される例:合成インシデントのインデックス登録、フィクスチャのペイロードを ES に再送信すること、レポート用に「データを利用可能にする」ために write/update/delete API を呼び出すこと。
- 要求された範囲/スコープに該当するインシデントが存在しない場合は、結果が空であるものとして処理してください(下記参照)。データを捏造してはなりません。
各インシデントについて、id、sensorId、timestamp、end、category、place.name、info.verdict、info.reasoning、objectIds、およびクリップURL(通常はinfo.clip_url、clip_url、またはレスポンスに含まれるクリップポインタフィールドのいずれか)を保持してください。レポートに貼り付ける前に、すべてのクリップ URL に$VSS_PUBLIC_HOST:$VSS_PUBLIC_PORTの書き換え(上記の「 ブラウザで再生可能なクリップ URL」を参照)を適用してください。生の値はHOST_IP:30888という URL であり、ユーザーのブラウザからはアクセスできません。
ステップ 3 — インシデント範囲レポートテンプレートの作成
assets/incident-range-report.md をコピーし、センサーごとに(センサーのスコープがない場合はカテゴリごとに)グループ分けし、判定結果を集計して、各インシデントをタイムスタンプ/カテゴリ/判定/理由とともにリストアップします。元のファイルは変更しないでください。 すべてのインシデントクリップ値は、ブラウザで再生可能なURLに書き換えられたものでなければなりません。インシデントにクリップURLが含まれていない場合は、クリップ行を省略してください。入力済みのセルにテンプレートの指示を含めないでください。
get_incidents が0 件の結果を返した場合は、処理を中止し、要求された範囲とスコープを明記した 1 行の空範囲ステートメントを正確に返すこと。インシデント範囲テンプレートを完全にレンダリングしたり、インシデントをでっち上げたり、テストデータを生成したり、モード A にフォールバックしたりしてはならない。
エラー処理
- プローブ、
curl、VLM呼び出し、または/vss-query-analyticsリクエストが失敗した場合は、ワークフローを停止し、失敗したエンドポイント、HTTPステータスまたはコマンドエラー、および次に取るべき有効な回復手順を報告してください。不完全または欠落したデータからレポートを捏造してはなりません。 - VLMの応答が空の場合、形式が不正な場合、または推論ブロックのみを含む場合は、その応答の問題を明示し、再試行前にモデルの準備状況やログを確認するよう提案してください。
- クリップのURLをパブリックホスト/ポートに書き換えることができない場合は、生成されたレポートからそれを省略し、ブラウザで再生可能なURLが生成できなかったことを明記してください。
- モード B の場合、オプションのインシデントフィールド(
info.reasoning、objectIds、クリップ URL)の欠落はレポート上の省略として扱いますが、ID、タイムスタンプ、またはカテゴリの欠落は、報告すべきデータ品質エラーとして扱ってください。
相互参照
/vss-manage-video-io-storage— モード A ステップ 1 のセンサーリスト、タイムライン、およびクリップ URL。/vss-query-analytics— モード B ステップ 2 におけるインシデントの取得(および判定/推論の補足)。/vss-ask-video— 単一のクリップに関するアドホックな VLM Q&A(構造化されたレポートではありません)。/vss-summarize-video—lvsプロファイルが展開されている場合、モード A が要約本文を生成するために使用します。レポートテンプレート(ステップ 4)は、ここで入力されます。
---
name: vss-generate-video-report
description: Generates video analysis reports by routing to a VLM backend for per-clip analysis or an analytics backend for incident-range reports, with deployment profile verification and URL rewriting.
license: Apache-2.0
---
# Report
Generate a video analysis report by routing to one of two backends — **never via** `POST /generate` on the VSS agent.
| Mode | Backend |
|---|---|
| **A. Video clip** | `/vss-manage-video-io-storage` → clip URL → **VLM chat/completions** |
| **B. Incident range** | `/vss-query-analytics` → incident list → narrative report |
If the request is ambiguous (e.g. "report on `<sensor>`" with no time range and no incident wording), default to **Mode A**. Ask only if the user mentions both a sensor and a time range. See **Examples** below for the request phrasings that route to each mode.
---
## Instructions
1. **Pick the mode** — Mode A for a single recorded clip/sensor video, Mode B when the request names a time range or incidents/alerts (match against *Examples*).
2. **Verify the deployment profile** for that mode under *Deployment prerequisite*; hand off to `/vss-deploy-profile` if its probe fails.
3. **Run that mode's numbered steps** — *Mode A* or *Mode B* below.
4. **Rewrite every user-facing clip URL** with the `$VSS_PUBLIC_HOST:$VSS_PUBLIC_PORT` one-liner (*Browser-playable clip URL*) before embedding it in the report.
5. **Return the rendered report markdown** to the user.
Output contract for evaluators:
- Mode A top title MUST be exactly `# Video Analysis Report`.
- Mode B top title MUST be exactly `# Incident Range Report` (never `# Incident Report` or sensor-named variants).
- Mode B MUST include `## Basic Information` with the exact required rows from the template (Report Identifier, Range, Scope, Total Incidents, Confirmed / Rejected / Unverified).
---
## Examples
- "Generate a report for this video" / "report on `<sensor-id>`" → **Mode A**
- "Analyze warehouse_01.mp4" / "create an analysis report on the uploaded video" → **Mode A**
- "Report on incidents from 12:31Z to 12:32Z" → **Mode B**
- "Report on alerts today" / "what incidents happened on `<sensor>` last hour" → **Mode B**
- "Summarize alerts on `<sensor>` between `<t1>` and `<t2>`" → **Mode B**
---
## Negative Triggers
Do **not** use this skill when the request is one of the following:
- Ad-hoc visual Q&A on a clip that do not ask explicitly for a report ("what color is the truck?", "what happens at 00:12?") → use `/vss-ask-video`.
- Archive/semantic similarity retrieval ("find forklifts", "search all videos for tailgating") → use `/vss-search-archive`.
- Read-only incident/metrics lookup without report rendering needs → use `/vss-query-analytics`.
- Deploy/teardown/profile changes ("deploy alerts", "switch profile", "bring up base") → use `/vss-deploy-profile`.
- Real-time alert/rule management requests → use `/vss-manage-alerts`.
Never route reports through VSS-agent `POST /generate`.
---
## Deployment prerequisite
**Mode A** needs the VSS **base** profile (VST + VLM NIM).
**Mode B** needs the VSS **alerts** profile (VA-MCP + Elasticsearch).
Probe:
```bash
# Mode A — VST + VLM reachability
curl -sf --max-time 5 "http://${HOST_IP}:30888/vst/api/v1/sensor/version" >/dev/null
# Mode B — VA-MCP
curl -sf --max-time 5 "http://${HOST_IP}:9901/" >/dev/null
```
If the probe fails, hand off to `/vss-deploy-profile` with `-p base` (Mode A) or `-p alerts` (Mode B). **Always** confirm the deploy with the user first.
---
## Clip URLs: VLM input vs browser report link
VST returns clip URLs using the agent-internal `${HOST_IP}:30888` host:port.
Keep that original URL as `VIDEO_URL` for local / in-cluster VLM frame pulls.
Do **not** rewrite the VLM input URL just to make it browser-playable.
Only create `BROWSER_CLIP_URL` for URLs shown in the rendered report. The
deploy layer exports the browser-facing host:port as `$VSS_PUBLIC_HOST` /
`$VSS_PUBLIC_PORT` (and scheme as `$VSS_PUBLIC_HTTP_PROTOCOL`) in every
profile `.env` — Brev or bare-metal — so the report-link rewrite is:
```bash
: "${VSS_PUBLIC_HOST:?Set VSS_PUBLIC_HOST before rewriting clip URLs}"
: "${VSS_PUBLIC_PORT:?Set VSS_PUBLIC_PORT before rewriting clip URLs}"
VSS_PUBLIC_HTTP_PROTOCOL="${VSS_PUBLIC_HTTP_PROTOCOL:-http}"
BROWSER_CLIP_URL=$(echo "$RAW_URL" | sed -E "s|^https?://[^/]+|${VSS_PUBLIC_HTTP_PROTOCOL}://${VSS_PUBLIC_HOST}:${VSS_PUBLIC_PORT}|")
```
If either required public host value is missing, omit the report-facing clip
link and call out that a browser-playable URL could not be produced; do not
block the local VLM analysis path. Apply the rewrite to **every clip URL
surfaced in the rendered report** (Mode A Step 4 Clip URL row; Mode B
per-incident clip sub-bullet). Leave the VLM `video_url` content block in Mode A
Step 3 on the original internal URL when the VLM is local / in-cluster.
---
## Mode A — Report on a recorded video clip
**If the VSS `lvs` profile is deployed** — `curl -sf --max-time 5 "http://${HOST_IP}:38111/v1/ready"` returns HTTP 200 — run `/vss-summarize-video` to produce the summary, then paste its output into the report template in Step 4 and skip Steps 1–3 (the VLM-direct path). Run Steps 1–3 only when `/v1/ready` is non-200.
### Step 1 — Resolve the clip URL
Hand off to `/vss-manage-video-io-storage` to:
1. List sensors and confirm the named `<sensor-id>` exists (upload first if not).
2. Fetch `/storage/<streamId>/timelines` for the recorded range when the user did not supply `startTime` / `endTime`.
3. Request a clip URL:
```bash
curl -s "http://${HOST_IP}:30888/vst/api/v1/storage/file/<streamId>/url?startTime=<startTime>&endTime=<endTime>&container=mp4&disableAudio=true" | jq -r .videoUrl
```
That gives a direct `mp4` URL that the local / in-cluster VLM can pull frames from. Bind it to `VIDEO_URL` (used by the VLM in Step 3) and set `RAW_URL="$VIDEO_URL"` before applying the report-link rewrite to produce `BROWSER_CLIP_URL` for Step 4 — the user's browser cannot reach `$VIDEO_URL` directly.
Mode A requires the selected VLM endpoint to be able to fetch `VIDEO_URL`.
Local NIM/RT-VLM deployments normally can; remote endpoints generally cannot
fetch `localhost`, private `HOST_IP`, or VST-internal URLs. If the live
`VLM_ENDPOINT` is remote, surface that reachability requirement instead of
making a chat request that will fail after `/v1/models` succeeds.
### Step 2 — Resolve VLM endpoint and model
The deploy may serve the VLM through either of two stacks. Both expose an OpenAI-compatible `chat/completions` API — pick whichever is live:
| Backend | Env vars | Typical host endpoint | Picked when |
|---|---|---|---|
| **NIM Cosmos** | `VLM_BASE_URL`, `VLM_NAME`, `VLM_MODE`, `VLM_MODEL_TYPE` | `${VLM_BASE_URL}/v1` (no trailing `/v1` on the env var; the agent appends it) | `VLM_MODEL_TYPE != rtvi` **and** `VLM_MODE` ∈ {`local`, `local_shared`, `remote`} **and** `VLM_BASE_URL` is non-empty |
| **RT-VLM Cosmos** | `RTVI_VLM_BASE_URL`, `RTVI_VLM_MODEL_TO_USE`, `VLM_MODEL_TYPE` | `${RTVI_VLM_BASE_URL}/v1` — if unset, derive from `${HOST_IP}` (`http://${HOST_IP}:8018/v1` for alerts, `http://${HOST_IP}:30082/v1` for base) | `VLM_MODEL_TYPE = rtvi`, or `VLM_MODE=none`, or `VLM_BASE_URL` empty; also the only path for `warehouse` |
Read the live values off the running agent container — do not guess:
```bash
docker exec vss-agent sh -lc '
for k in HOST_IP VLM_MODE VLM_MODEL_TYPE VLM_BASE_URL VLM_NAME RTVI_VLM_BASE_URL RTVI_VLM_MODEL_TO_USE; do
v="$(printenv "$k")"
[ -n "$v" ] && printf "%s=%s\n" "$k" "$v"
done
'
```
Do not require `RTVI_VLM_ENDPOINT` from `vss-agent` env; several profiles do not inject it.
Selection rule:
```bash
if [ "${VLM_MODEL_TYPE:-}" = "rtvi" ]; then
VLM_BACKEND="rtvlm"
VLM_ENDPOINT="${RTVI_VLM_BASE_URL:+${RTVI_VLM_BASE_URL%/}/v1}"
[ -z "${VLM_ENDPOINT}" ] && VLM_ENDPOINT="http://${HOST_IP}:8018/v1" # alerts default
VLM_MODEL="${RTVI_VLM_MODEL_TO_USE}"
elif [ -n "${VLM_BASE_URL}" ] && [ "${VLM_MODE}" != "none" ]; then
VLM_BACKEND="nim_cosmos"
VLM_ENDPOINT="${VLM_BASE_URL%/}/v1"
VLM_MODEL="${VLM_NAME}"
else
VLM_BACKEND="rtvlm"
VLM_ENDPOINT="${RTVI_VLM_BASE_URL:+${RTVI_VLM_BASE_URL%/}/v1}"
[ -z "${VLM_ENDPOINT}" ] && VLM_ENDPOINT="http://${HOST_IP}:30082/v1" # base default
VLM_MODEL="${RTVI_VLM_MODEL_TO_USE}"
fi
```
Probe `/v1/models` before sending a chat request to confirm the chosen endpoint is alive and the model is loaded:
```bash
curl -sf --max-time 5 "${VLM_ENDPOINT}/models" | jq -r '.data[].id'
```
If the probe fails or the listed ids don't include `${VLM_MODEL}`, fall back to the other backend (or surface the error — never silently pick a model that isn't on the server).
### Step 3 — Call the VLM directly
Use the OpenAI-compatible `chat/completions` endpoint with a `video_url` content block — the same payload shape **and multimodal settings** `video_understanding` builds in `src/vss_agents/tools/video_understanding.py` (`_build_vlm_messages` + the Cosmos `base_vlm.bind(...)` call).
The frame sampling and visual-token (pixel) budget must mirror the **live** `video_understanding` settings for the active profile. **Send `mm_processor_kwargs` and `media_io_kwargs`** so the direct call uses the same frame sampling and pixel budget as the in-agent `video_understanding` tool — omitting them lets the VLM apply its own defaults, so the output diverges from the agent path.
```bash
PROMPT='Describe in detail what happens in the video, with timestamps (start–end in seconds from clip start) for each segment or event. Cover scenes, objects, people, vehicles, and notable actions.'
# Reasoning is OFF by default — matches the base-profile video_understanding config (`reasoning: false`).
# video_understanding.py uses config.reasoning unless the caller overrides it, so default to non-reasoning.
# Append the Cosmos Reason 2 reasoning suffix ONLY when the user explicitly asks for reasoning
# (drop it for non-cosmos-reason2 VLMs). With reasoning off, the response has no <think> block.
if [ "${REASONING:-false}" = "true" ]; then
PROMPT="${PROMPT}
Answer the question using the following format:
<think>
Your reasoning.
</think>
Write your final answer immediately after the </think> tag."
fi
# If Step 3 is run standalone, derive missing backend from current env/model.
[ -z "${VLM_BACKEND:-}" ] && {
if [ "${VLM_MODEL_TYPE:-}" = "rtvi" ]; then
VLM_BACKEND="rtvlm"
elif [[ "${VLM_MODEL:-}" == nvidia/cosmos* ]]; then
VLM_BACKEND="nim_cosmos"
else
VLM_BACKEND="rtvlm"
fi
}
# Multimodal settings — resolve from the live agent config file path, not hardcoded candidates.
CFG_JSON=$(
docker exec vss-agent python3 -c '
import json, os, yaml
p = os.getenv("VSS_AGENT_CONFIG_FILE")
if not p:
raise SystemExit("VSS_AGENT_CONFIG_FILE is not set in vss-agent")
if not os.path.isabs(p):
p = os.path.join("/vss-agent", p.lstrip("./"))
with open(p, encoding="utf-8") as f:
cfg = yaml.safe_load(f) or {}
vu = (cfg.get("functions", {}) or {}).get("video_understanding", {}) or {}
print(json.dumps({
"max_fps": int(vu.get("max_fps", 2)),
"max_frames": int(vu.get("max_frames", 30)),
"min_pixels": int(vu.get("min_pixels", 3136)),
"max_pixels": int(vu.get("max_pixels", 8388608)),
}))
')
)
[ -n "${CFG_JSON}" ] || { echo "Failed to read video_understanding config from vss-agent"; exit 1; }
jq -e . >/dev/null <<< "${CFG_JSON}" || { echo "Invalid config JSON from vss-agent"; exit 1; }
MAX_FPS="$(jq -r '.max_fps' <<< "${CFG_JSON}")"
MAX_FRAMES="$(jq -r '.max_frames' <<< "${CFG_JSON}")"
MIN_PIXELS="$(jq -r '.min_pixels' <<< "${CFG_JSON}")"
MAX_PIXELS="$(jq -r '.max_pixels' <<< "${CFG_JSON}")"
# num_frames = min(int(clip_seconds) * max_fps, max_frames), min 1 — matches video_understanding.py.
# clip_seconds (Step 1 endTime-startTime) may be fractional; truncate to integer seconds — bash $((...))
# is integer-only and errors on "15.0"/"1.5". Default 15s -> caps at MAX_FRAMES.
CLIP_SECONDS=$(awk -v s="${CLIP_SECONDS:-15}" 'BEGIN{printf "%d", s}')
NUM_FRAMES=$(( CLIP_SECONDS * MAX_FPS ))
[ "$NUM_FRAMES" -gt "$MAX_FRAMES" ] && NUM_FRAMES=$MAX_FRAMES
[ "$NUM_FRAMES" -lt 1 ] && NUM_FRAMES=1
# Only apply Cosmos mm/media kwargs on the NIM Cosmos path.
# RT-VLM mode uses its own server-side preprocessing and should not receive these kwargs.
MM_KWARGS=""
if [ "${VLM_BACKEND}" = "nim_cosmos" ]; then
case "$VLM_MODEL" in
*cosmos-reason2*) MM_KWARGS=", \"mm_processor_kwargs\": {\"size\": {\"shortest_edge\": ${MIN_PIXELS}, \"longest_edge\": ${MAX_PIXELS}}}, \"media_io_kwargs\": {\"video\": {\"num_frames\": ${NUM_FRAMES}}}" ;;
*cosmos*) MM_KWARGS=", \"mm_processor_kwargs\": {\"videos_kwargs\": {\"min_pixels\": ${MIN_PIXELS}, \"max_pixels\": ${MAX_PIXELS}}}, \"media_io_kwargs\": {\"video\": {\"num_frames\": ${NUM_FRAMES}}}" ;;
*) MM_KWARGS="" ;;
esac
fi
curl -s --connect-timeout 5 --max-time 120 -X POST "${VLM_ENDPOINT}/chat/completions" \
-H "Content-Type: application/json" \
-d @- <<EOF | jq -r '.choices[0].message.content'
{
"model": $(jq -Rs . <<< "${VLM_MODEL}"),
"messages": [
{
"role": "user",
"content": [
{"type": "text", "text": $(jq -Rs . <<< "${PROMPT}")},
{"type": "video_url", "video_url": {"url": $(jq -Rs . <<< "${VIDEO_URL}")}}
]
}
],
"max_tokens": 1024,
"temperature": 0.0${MM_KWARGS}
}
EOF
```
> The kwargs block is backend-aware: on `nim_cosmos`, Reason2 variants (`nvidia/cosmos-reason2*`) use `mm_processor_kwargs.size{shortest_edge,longest_edge}` and other NIM Cosmos variants (`nvidia/cosmos*`) use `mm_processor_kwargs.videos_kwargs{min_pixels,max_pixels}`; both also send `media_io_kwargs.video.num_frames`. On `rtvlm`, no Cosmos kwargs are sent.
If the VLM returns a `<think>…</think>` block (Cosmos Reason reasoning mode), keep only the text after `</think>` as the report body.
### Step 4 — Fill the Video Analysis Report template
Copy [`assets/video-analysis-report.md`](assets/video-analysis-report.md), fill every placeholder, and return the rendered markdown to the user. Keep the source asset unchanged. Before rendering, verify `BROWSER_CLIP_URL` is set and non-empty, then replace `<BROWSER_CLIP_URL>` with that exact value in the `Clip URL` row. Never leave the placeholder in the output, never include template instructions in a filled cell, and never use the raw `HOST_IP:30888` URL.
---
## Mode B — Report on incidents in a time range
### Step 1 — Resolve the time range and (optionally) sensor
- `start_time` / `end_time` must be ISO 8601 UTC (`YYYY-MM-DDTHH:MM:SS.sssZ`). Resolve relative phrases ("last hour", "today") against the current host clock.
- If the user names a sensor, capture it as `source` + `source_type=sensor`. Otherwise leave both unset for an all-sensors query.
### Step 2 — Fetch incidents via `/vss-query-analytics`
Hand off to `/vss-query-analytics` (initialize → `tools/call`) with:
```json
{
"jsonrpc": "2.0",
"method": "tools/call",
"params": {
"name": "video_analytics__get_incidents",
"arguments": {
"source": "<sensor-id-or-omit>",
"source_type": "sensor",
"start_time": "<ISO>",
"end_time": "<ISO>",
"max_count": 100,
"includes": ["objectIds", "info"]
}
},
"id": 1
}
```
Read-only boundary (mandatory):
- Mode B is strictly read-only analytics retrieval. Never write, seed, backfill, or mutate Elasticsearch/VA data.
- Forbidden examples: indexing synthetic incidents, replaying fixture payloads into ES, calling write/update/delete APIs to "make data available" for the report.
- If no incidents exist for the requested range/scope, handle as empty results (see below); do not fabricate data.
For each incident keep: `id`, `sensorId`, `timestamp`, `end`, `category`, `place.name`, `info.verdict`, `info.reasoning`, `objectIds`, and the clip URL (commonly `info.clip_url`, `clip_url`, or whichever clip-pointer field the response carries). **Apply the `$VSS_PUBLIC_HOST:$VSS_PUBLIC_PORT` rewrite (see *Browser-playable clip URL* above) to every clip URL before pasting it into the report** — the raw value is a `HOST_IP:30888` URL the user's browser cannot reach.
### Step 3 — Fill the Incident Range Report template
Copy [`assets/incident-range-report.md`](assets/incident-range-report.md), then group by sensor (or by category if no sensor scope), tally verdicts, and list each incident with timestamp / category / verdict / reasoning. Keep the source asset unchanged. Every incident clip value must be a rewritten browser-playable URL; omit the clip line when the incident carries no clip URL. Never include template instructions in a filled cell.
If `get_incidents` returns zero results, STOP and return exactly a one-line empty-range statement naming the requested range and scope. Do not render the full Incident Range template, do not invent incidents, do not seed test data, and do not fall back to Mode A.
---
## Error Handling
- If a probe, `curl`, VLM call, or `/vss-query-analytics` request fails, stop the workflow and report the failing endpoint, HTTP status or command error, and the next useful recovery step. Do not fabricate a report from partial or missing data.
- If the VLM response is empty, malformed, or contains only a reasoning block, surface that response problem and suggest checking model readiness/logs before retrying.
- If a clip URL cannot be rewritten to the public host/port, omit it from the rendered report and call out that the browser-playable URL could not be produced.
- For Mode B, treat missing optional incident fields (`info.reasoning`, `objectIds`, clip URL) as omissions in the report, but treat missing `id`, `timestamp`, or `category` as a data-quality error that should be reported.
---
## Cross-Reference
- **`/vss-manage-video-io-storage`** — sensor list, timelines, and clip URL for Mode A Step 1.
- **`/vss-query-analytics`** — incident retrieval (and verdict / reasoning enrichment) for Mode B Step 2.
- **`/vss-ask-video`** — ad-hoc VLM Q&A on a single clip (not a structured report).
- **`/vss-summarize-video`** — used by Mode A to produce the summary body when the `lvs` profile is deployed; the report template (Step 4) is still filled here.
すべてのファイル
8件のファイルvss-generate-video-reportをインストール
スキルファイルをダウンロードし、.claude/skills/ ディレクトリに解凍してください。
ZIPをダウンロードリポジトリをクローンし、スキルファイルをプロジェクトにコピーしてください。
git clone https://github.com/NVIDIA/skills/tree/main/skills/vss-generate-video-report # Copy SKILL.md to your .claude/skills/ directory
コピー





家
