vss-generate-video-report
NVIDIA/skills
透過將資料路由至 VLM 後端進行逐片段分析,或路由至分析後端以產生事件範圍報告,來生成影片分析報告,並包含部署設定檔驗證及 URL 重寫功能。
...展開全部報告
透過將請求路由至兩個後端之一來產生影片分析報告——切勿透過VSS 代理伺服器的POST /generate方法進行。
| 模式 | 後端 |
|---|---|
| A. 影片片段 | /vss-manage-video-io-storage→ 片段 URL →VLM 聊天/自動完成功能 |
| B. 事件範圍 | /vss-query-analytics→ 事件清單 → 敘述性報告 |
若請求內容含糊不清(例如「關於 」且未指定時間範圍及事件描述),則預設採用模式 A。僅當使用者同時提及感測器與時間範圍時才進行詢問。請參閱下方的範例,了解會導向各模式的請求措辭。
操作說明
- 選擇模式— 若為單一錄製片段/感測器影片則選模式A,若請求中指定了時間範圍或事件/警報則選模式 B(請參照範例進行比對)。
- 在「部署先決條件」下驗證該模式的部署設定檔;若其探測失敗,則交由
/vss-deploy-profile處理。 - 執行該模式的編號步驟— 參見下方的模式 A或模式 B。
- 在將片段嵌入報告之前,將每個面向使用者的片段 URL 改寫為
$VSS_PUBLIC_HOST:$VSS_PUBLIC_PORT單行格式(可於瀏覽器播放的片段 URL)。 - 將渲染後的報告 Markdown 內容回傳給使用者。
供評估人員參考的輸出規範:
- 模式 A 的頂部標題必須精確為
# 影片分析報告。 - 模式 B 的頂部標題必須精確為
# 事件範圍報告(絕不能是# 事件報告或包含感測器名稱的變體)。 - 模式 B 必須包含
## 基本資訊,且須包含範本中確切要求的行項目(報告識別碼、範圍、範圍、總事件數、已確認/已駁回/未經核實)。
範例
- 「為此影片生成報告」/「關於
” →模式 A - 「分析 warehouse_01.mp4」/「針對上傳的影片建立分析報告」→模式 A
- 「關於 12:31Z 至 12:32Z 期間的事件報告」→模式 B
- 「彙整今日的警示報告」/「過去
」 →模式 B - 「彙總
期間與" → 模式 B" →模式 B
負面觸發條件
當請求屬於以下任一情況時,請勿使用此技能:
- 針對影片片段進行即興視覺問答,且未明確要求報告(「卡車是什麼顏色?」、「00:12 發生了什麼事?」)→ 請使用
/vss-ask-video。 - 檔案庫/語義相似度檢索(「尋找堆高機」、「搜尋所有影片中的跟車過近情境」)→ 請使用
/vss-search-archive。 - 無需生成報告的唯讀事件/指標查詢 → 請使用
/vss-query-analytics。 - 部署/拆除/設定檔變更(「部署警示」、「切換設定檔」、「啟動基礎環境」)→ 請使用
/vss-deploy-profile。 - 即時警報/規則管理請求 → 請使用
/vss-manage-alerts。
切勿透過 VSS-agentPOST /generate 傳送報告。
部署先決條件
模式 A需要 VSS基礎配置檔(VST + VLM NIM)。 模式 B需要 VSS警示配置檔(VA-MCP + Elasticsearch)。
測試:
# 模式 A — VST + VLM 可達性
curl -sf --max-time 5 "http://${HOST_IP}:30888/vst/api/v1/sensor/version" >/dev/null
# 模式 B — VA-MCP
curl -sf --max-time 5 "http://${HOST_IP}:9901/" >/dev/null
若探測失敗,請透過-p base(模式 A)或-p alerts(模式 B)將任務移交至/vss-deploy-profile。務必先與使用者確認部署事宜。
片段 URL:VLM 輸入與瀏覽器報告連結
VST 會使用代理程式內部${HOST_IP}:30888主機:埠 來回傳片段 URL。
請將該原始 URL 保留為VIDEO_URL,供本地端/叢集內的 VLM 幀擷取使用。
切勿僅為了使 VLM 輸入 URL 能在瀏覽器中播放而重新撰寫該 URL。
僅針對渲染後報告中顯示的 URL 建立BROWSER_CLIP_URL。
部署層會將對外公開的主機:埠號以$VSS_PUBLIC_HOST/
$VSS_PUBLIC_PORT(以及方案為$VSS_PUBLIC_HTTP_PROTOCOL)的形式,
因此報告連結的重寫規則為:
: "${VSS_PUBLIC_HOST:?請在重寫片段 URL 之前設定 VSS_PUBLIC_HOST}"
: "${VSS_PUBLIC_PORT:?請在重寫片段 URL 之前設定 VSS_PUBLIC_PORT}"
VSS_PUBLIC_HTTP_PROTOCOL="${VSS_PUBLIC_HTTP_PROTOCOL:-http}"
BROWSER_CLIP_URL=$(echo "$RAW_URL" | sed -E "s|^https?://[^/]+|${VSS_PUBLIC_HTTP_PROTOCOL}://${VSS_PUBLIC_HOST}:${VSS_PUBLIC_PORT}|")
若任一必需的公開主機值缺失,則省略報告中的片段
連結,並標註無法產生可在瀏覽器中播放的 URL;切勿
阻擋本機 VLM 分析路徑。 將此重寫規則套用至渲染報告中呈現的每個片段 URL
(模式 A 第 4 步的「片段 URL」行;模式 B
的「每起事件片段」子項目)。當 VLM 位於本地端或叢集內時,請保留模式 A
第 3 步中 VLM 的video_url內容區塊,並維持原始內部 URL。
模式 A — 針對錄製的影片片段產生報告
若已部署 VSSlvs設定檔—執行 `curl -sf --max-time 5 "http://${HOST_IP}:38111/v1/ready"`並返回 HTTP 200 狀態碼 — 請執行`/vss-summarize-video` 產生摘要, 接著將其輸出內容貼入步驟 4 的報告範本中,並跳過步驟 1–3(VLM 直接路徑)。僅當/v1/ready返回非 200 狀態時,才執行步驟 1–3。
步驟 1 — 解析片段 URL
交由/vss-manage-video-io-storage執行以下操作:
列出感測器並確認指定
是否存在(若不存在,請先上傳)。當使用者未提供
startTime/endTime時,從/storage/擷取所錄製區間的資料。/timelines 請求片段網址:
curl -s "http://${HOST_IP}:30888/vst/api/v1/storage/file//url?startTime= &endTime= &container=mp4&disableAudio=true" | jq -r .videoUrl 這會產生一個直接的
mp4網址,讓本機/叢集內的 VLM 能從中擷取畫面。 將其綁定至VIDEO_URL(於步驟 3 中由 VLM 使用),並在套用報告連結重寫以產生步驟 4 所需的BROWSER_CLIP_URL之前,設定RAW_URL="$VIDEO_URL"— 因為使用者的瀏覽器無法直接存取$VIDEO_URL。 模式 A 要求所選的 VLM 端點能夠取得VIDEO_URL。 本地的 NIM/RT-VLM 部署通常可以;遠端端點一般無法 取得localhost、私有HOST_IP或 VST 內部 URL。 若即時VLM_ENDPOINT為遠端,應明確提出此可達性要求,而非 發出聊天請求——該請求在/v1/models成功後仍會失敗。
步驟 2 — 解析 VLM 端點與模型
該部署可透過以下兩種架構之一提供 VLM 服務。兩者均提供與 OpenAI 相容的聊天/補全API — 請選擇其中一個已上線的:
| 後端 | 環境變數 | 典型主機端點 | 何時選用 |
|---|---|---|---|
| NIM Cosmos | VLM_BASE_URL、VLM_NAME、VLM_MODE、VLM_MODEL_TYPE |
${VLM_BASE_URL}/v1(環境變數末尾不帶/v1;由代理程式自動追加) |
VLM_MODEL_TYPE ≠ rtvi 且 VLM_MODE∈ {local,local_shared,remote}且 VLM_BASE_URL不為空 |
| RT-VLM Cosmos | RTVI_VLM_BASE_URL、RTVI_VLM_MODEL_TO_USE、VLM_MODEL_TYPE |
${RTVI_VLM_BASE_URL}/v1— 若未設定,則從${HOST_IP}推導(警示用網址為http://${HOST_IP}:8018/v1, http://${HOST_IP}:30082/v1適用於基礎資料) |
VLM_MODEL_TYPE = rtvi,或VLM_MODE=none,或VLM_BASE_URL為空;此亦為資料倉儲的唯一路徑 |
從正在運行的代理容器中讀取即時數值 — 切勿憑空推測:
docker exec vss-agent sh -lc '
for k in HOST_IP VLM_MODE VLM_MODEL_TYPE VLM_BASE_URL VLM_NAME RTVI_VLM_BASE_URL RTVI_VLM_MODEL_TO_USE; do
v="$(printenv "$k")"
[ -n "$v" ] && printf "%s=%s\n" "$k" "$v"
done
'
無需從vss-agent環境變數中取得RTVI_VLM_ENDPOINT;因部分設定檔並未注入此變數。
篩選規則:
if [ "${VLM_MODEL_TYPE:-}" = "rtvi" ]; then
VLM_BACKEND="rtvlm"
VLM_ENDPOINT="${RTVI_VLM_BASE_URL:+${RTVI_VLM_BASE_URL%/}/v1}"
[ -z "${VLM_ENDPOINT}" ] && VLM_ENDPOINT="http://${HOST_IP}:8018/v1" # 預設告警
VLM_MODEL="${RTVI_VLM_MODEL_TO_USE}"
elif [ -n "${VLM_BASE_URL}" ] && [ "${VLM_MODE}" != "none" ]; then
VLM_BACKEND="nim_cosmos"
VLM_ENDPOINT="${VLM_BASE_URL%/}/v1"
VLM_MODEL="${VLM_NAME}"
else
VLM_BACKEND="rtvlm"
VLM_ENDPOINT="${RTVI_VLM_BASE_URL:+${RTVI_VLM_BASE_URL%/}/v1}"
[ -z "${VLM_ENDPOINT}" ] && VLM_ENDPOINT="http://${HOST_IP}:30082/v1" # 預設值
VLM_MODEL="${RTVI_VLM_MODEL_TO_USE}"
fi
在發送聊天請求前,先對/v1/models進行探測,以確認所選端點正常運作且模型已載入:
curl -sf --max-time 5 "${VLM_ENDPOINT}/models" | jq -r '.data[].id'
若探測失敗,或清單中的 ID 不包含${VLM_MODEL},則切換至其他後端(或顯示錯誤訊息 — 絕不默默選用伺服器上不存在的模型)。
步驟 3 — 直接呼叫 VLM
使用相容於 OpenAI 的chat/completions端點,並傳入包含video_url內容區塊的請求 — 其載荷結構與多模態設定應與 src/vss_agents/tools/video_understanding.py中video_understanding模組所建構的相同 (_build_vlm_messages加上 Cosmos 的base_vlm.bind(...)呼叫)。
幀採樣與視覺標記(像素)預算必須與當前配置檔的即時 video_understanding設定一致。請傳送mm_processor_kwargs和media_io_kwargs,以便直接呼叫時能採用與代理程式內建video_understanding工具相同的幀採樣和像素預算 — 若省略這些參數,VLM 將套用其預設值,導致輸出結果與代理程式路徑產生差異。
PROMPT='請詳細描述影片中的情境,並為每個片段或事件標註時間戳記(以影片起始時間為基準的起訖秒數)。內容應涵蓋場景、物體、人物、車輛及值得注意的動作。'
# 預設關閉推理功能 — 與基礎配置檔 `video_understanding` 的設定(`reasoning: false`)一致。
# `video_understanding.py` 會使用 `config.reasoning`,除非呼叫方覆寫該設定,因此預設為不進行推理。
# 僅當使用者明確要求推理時,才附加 Cosmos Reason 2 的推理後綴
# (對於非 Cosmos Reason 2 的 VLM,則省略此後綴)。若推理功能關閉,回應中將不包含 `reasoning ` 區塊。
if [ "${REASONING:-false}" = "true" ]; then
PROMPT="${PROMPT}
請使用以下格式回答問題:
您的推理過程。
請在 標籤後立即寫下您的最終答案 。"
fi
# 若第 3 步驟以獨立模式執行,則從當前環境/模型推導出缺失的後端。
[ -z "${VLM_BACKEND:-}" ] && {
if [ "${VLM_MODEL_TYPE:-}" = "rtvi" ]; then
VLM_BACKEND="rtvlm"
elif [[ "${VLM_MODEL:-}" == nvidia/cosmos* ]]; then
VLM_BACKEND="nim_cosmos"
else
VLM_BACKEND="rtvlm"
fi
}
# 多模態設定 — 從 VSS 代理程式設定檔路徑解析,而非使用硬編碼的候選值。
CFG_JSON=$(
docker exec vss-agent python3 -c '
import json, os, yaml
p = os.getenv("VSS_AGENT_CONFIG_FILE")
if not p:
raise SystemExit("vss-agent 中未設定 VSS_AGENT_CONFIG_FILE")
if not os.path.isabs(p):
p = os.path.join("/vss-agent", p.lstrip("./"))
with open(p, encoding="utf-8") as f:
cfg = yaml.safe_load(f) or {}
vu = (cfg.get("functions", {}) or {}).get("video_understanding", {}) or {}
print(json.dumps({
"max_fps": int(vu.get("max_fps", 2)),
"max_frames": int(vu.get("max_frames", 30)),
"min_pixels": int(vu.get("min_pixels", 3136)),
"max_pixels": int(vu.get("max_pixels", 8388608)),
}))
')
)
[ -n "${CFG_JSON}" ] || { echo "無法從 vss-agent 讀取 video_understanding 設定檔"; exit 1; }
jq -e . >/dev/null <<< "${CFG_JSON}" || { echo "來自 vss-agent 的 JSON 設定檔無效"; exit 1; }
MAX_FPS="$(jq -r '.max_fps' <<< "${CFG_JSON}")"
MAX_FRAMES="$(jq -r '.max_frames' <<< "${CFG_JSON}")"
MIN_PIXELS="$(jq -r '.min_pixels' <<< "${CFG_JSON}")"
MAX_PIXELS="$(jq -r '.max_pixels' <<< "${CFG_JSON}")"
# num_frames = min(int(clip_seconds) * max_fps, max_frames),最小值為 1 — 與 video_understanding.py 一致。
# clip_seconds(步驟 1 的 endTime - startTime)可能包含小數;截斷為整數秒 — bash $((...))
# 僅接受整數,若為 "15.0"/"1.5" 會報錯。預設 15 秒 → 上限為 MAX_FRAMES。
CLIP_SECONDS=$(awk -v s="${CLIP_SECONDS:-15}" 'BEGIN{printf "%d", s}')
NUM_FRAMES=$(( CLIP_SECONDS * MAX_FPS ))
[ "$NUM_FRAMES" -gt "$MAX_FRAMES" ] && NUM_FRAMES=$MAX_FRAMES
[ "$NUM_FRAMES" -lt 1 ] && NUM_FRAMES=1
# 僅在 NIM Cosmos 路徑上套用 Cosmos mm/media 命令列參數。
# RT-VLM 模式使用其專屬的伺服器端預處理,不應接收這些命令列參數。
MM_KWARGS=""
if [ "${VLM_BACKEND}" = "nim_cosmos" ]; then
case "$VLM_MODEL" in
*cosmos-reason2*) MM_KWARGS=", \"mm_processor_kwargs\": {\"size\": {\"shortest_edge\": ${MIN_PIXELS}, \"longest_edge\": ${MAX_PIXELS}}}, \"media_io_kwargs\": {\"video\": {\"num_frames\": ${NUM_FRAMES}}}" ;;
*cosmos*) MM_KWARGS=", \"mm_processor_kwargs\": {\"videos_kwargs\": {\"min_pixels\": ${MIN_PIXELS}, \"max_pixels\": ${MAX_PIXELS}}}, \"media_io_kwargs\": {\"video\": {\"num_frames\": ${NUM_FRAMES}}}" ;;
*) MM_KWARGS="" ;;
esac
fi
curl -s --connect-timeout 5 --max-time 120 -X POST "${VLM_ENDPOINT}/chat/completions" \
-H "Content-Type: application/json" \
-d @- <<EOF | jq -r '.choices[0].message.content'
{
"model": $(jq -Rs . <<< "${VLM_MODEL}"),
"messages": [
{
"role": "user",
"content": [
{"type": "text", "text": $(jq -Rs . <<< "${PROMPT}")},
{"type": "video_url", "video_url": {"url": $(jq -Rs . <<< "${VIDEO_URL}")}}
]
}
],
"max_tokens": 1024,
"temperature": 0.0${MM_KWARGS}
}
EOF
kwargs 區塊會根據後端進行調整:在
nim_cosmos上,Reason2 變體(nvidia/cosmos-reason2*)使用mm_processor_kwargs.size{shortest_edge,longest_edge},而其他 NIM Cosmos 變體(nvidia/cosmos*)則使用mm_processor_kwargs.videos_kwargs{min_pixels,max_pixels};兩者均會傳送media_io_kwargs.video.num_frames。在rtvlm上,不會傳送任何 Cosmos 參數。
若 VLM 傳回一個 區塊(Cosmos Reason 推理模式),則僅保留 之後的文字作為報告正文。
步驟 4 — 填寫影片分析報告範本
複製assets/video-analysis-report.md,填寫所有佔位符,並將渲染後的 Markdown 內容回傳給使用者。原始檔案請保持不變。渲染前,請確認BROWSER_CLIP_URL已設定且非空,然後將 該列中的「Clip URL」以該精確值替換。輸出時絕不可保留佔位符,絕不可在已填入內容的儲存格中包含範本說明,且絕不可使用原始的HOST_IP:30888URL。
模式 B — 針對特定時間區間的事件建立報告
步驟 1 — 解析時間範圍及(可選)感測器
start_time/end_time必須符合 ISO 8601 UTC 格式(YYYY-MM-DDTHH:MM:SS.sssZ)。將相對時間表述(如「過去一小時」、「今天」)轉換為當前主機時鐘的時間。- 若使用者指定感測器名稱,請將其擷取為
source+source_type=sensor。否則,請將兩者均設為未設定,以執行所有感測器的查詢。
步驟 2 — 透過/vss-query-analytics擷取事件
將處理權交給/vss-query-analytics(初始化 →tools/call),參數如下:
{
"jsonrpc": "2.0",
"method": "tools/call",
"params": {
"name": "video_analytics__get_incidents",
"arguments": {
"source": "",
"source_type": "sensor",
"start_time": "",
"end_time": "",
"max_count": 100,
"includes": ["objectIds", "info"]
}
},
"id": 1
}
唯讀邊界(強制要求):
- 模式 B 嚴格限於唯讀分析資料擷取。切勿對 Elasticsearch/VA 資料進行寫入、初始化、回填或變更。
- 禁止的範例:將模擬事件進行索引、將測試資料載入 Elasticsearch、呼叫寫入/更新/刪除 API 以「讓資料可供報告使用」。
- 若請求的範圍/範圍內不存在任何事件,請將其視為空結果處理(參見下文);切勿捏造資料。
針對每個事件,應保留:id、sensorId、timestamp、end、category、place.name、info.verdict、info.reasoning、objectIds 以及片段 URL(通常為info.clip_url、clip_url,或回應中攜帶的任何片段指標欄位)。在將每個片段 URL 貼入報告前,請先套用$VSS_PUBLIC_HOST:$VSS_PUBLIC_PORT重寫規則(參見上文「可在瀏覽器中播放的片段 URL」)——原始值為HOST_IP:30888格式的 URL,用戶的瀏覽器無法存取此網址。
步驟 3 — 填寫「事件範圍報告」範本
複製assets/incident-range-report.md,然後按感測器分組(若無感測器範圍,則按類別分組),統計判定結果,並列出每個事件的時戳/類別/判定結果/理由。請勿修改原始檔案。 每個事件片段值必須是經重寫後、瀏覽器可播放的 URL;若事件未包含片段 URL,則省略該片段行。切勿在已填寫的儲存格中包含範本說明。
若get_incidents返回零個結果,請立即停止並回傳精確一行、僅包含所請求範圍與範圍的空範圍聲明。請勿渲染完整的「事件範圍」範本、請勿捏造事件、請勿植入測試資料,亦請勿回退至模式 A。
錯誤處理
- 若探針、
curl、VLM 呼叫或/vss-query-analytics請求失敗,請停止工作流程,並回報失敗的端點、HTTP 狀態碼或命令錯誤,以及下一項有用的恢復步驟。切勿根據不完整或缺失的資料捏造報告。 - 若 VLM 回應為空、格式錯誤,或僅包含推理區塊,請明確指出該回應問題,並建議在重試前檢查模型就緒狀態/日誌。
- 若無法將片段 URL 重寫為公開主機/埠號,請將其從生成的報告中省略,並特別註明無法產生可在瀏覽器中播放的 URL。
- 針對模式 B,若缺少可選的事件欄位(如
info.reasoning、objectIds、片段 URL),應視為報告中的遺漏項目;但若缺少ID、時間戳記或類別,則應視為需通報的資料品質錯誤。
參見
/vss-manage-video-io-storage— 模式 A 第 1 步驟的感測器清單、時間軸及片段 URL。/vss-query-analytics— 模式 B 步驟 2 的事件檢索(以及判定/推論增益)。/vss-ask-video— 針對單一片段的即席 VLM 問答(非結構化報告)。/vss-summarize-video— 當lvs設定檔部署時,由模式 A 用於生成摘要正文;報告範本(步驟 4)仍在此處填入。
---
name: vss-generate-video-report
description: Generates video analysis reports by routing to a VLM backend for per-clip analysis or an analytics backend for incident-range reports, with deployment profile verification and URL rewriting.
license: Apache-2.0
---
# Report
Generate a video analysis report by routing to one of two backends — **never via** `POST /generate` on the VSS agent.
| Mode | Backend |
|---|---|
| **A. Video clip** | `/vss-manage-video-io-storage` → clip URL → **VLM chat/completions** |
| **B. Incident range** | `/vss-query-analytics` → incident list → narrative report |
If the request is ambiguous (e.g. "report on `<sensor>`" with no time range and no incident wording), default to **Mode A**. Ask only if the user mentions both a sensor and a time range. See **Examples** below for the request phrasings that route to each mode.
---
## Instructions
1. **Pick the mode** — Mode A for a single recorded clip/sensor video, Mode B when the request names a time range or incidents/alerts (match against *Examples*).
2. **Verify the deployment profile** for that mode under *Deployment prerequisite*; hand off to `/vss-deploy-profile` if its probe fails.
3. **Run that mode's numbered steps** — *Mode A* or *Mode B* below.
4. **Rewrite every user-facing clip URL** with the `$VSS_PUBLIC_HOST:$VSS_PUBLIC_PORT` one-liner (*Browser-playable clip URL*) before embedding it in the report.
5. **Return the rendered report markdown** to the user.
Output contract for evaluators:
- Mode A top title MUST be exactly `# Video Analysis Report`.
- Mode B top title MUST be exactly `# Incident Range Report` (never `# Incident Report` or sensor-named variants).
- Mode B MUST include `## Basic Information` with the exact required rows from the template (Report Identifier, Range, Scope, Total Incidents, Confirmed / Rejected / Unverified).
---
## Examples
- "Generate a report for this video" / "report on `<sensor-id>`" → **Mode A**
- "Analyze warehouse_01.mp4" / "create an analysis report on the uploaded video" → **Mode A**
- "Report on incidents from 12:31Z to 12:32Z" → **Mode B**
- "Report on alerts today" / "what incidents happened on `<sensor>` last hour" → **Mode B**
- "Summarize alerts on `<sensor>` between `<t1>` and `<t2>`" → **Mode B**
---
## Negative Triggers
Do **not** use this skill when the request is one of the following:
- Ad-hoc visual Q&A on a clip that do not ask explicitly for a report ("what color is the truck?", "what happens at 00:12?") → use `/vss-ask-video`.
- Archive/semantic similarity retrieval ("find forklifts", "search all videos for tailgating") → use `/vss-search-archive`.
- Read-only incident/metrics lookup without report rendering needs → use `/vss-query-analytics`.
- Deploy/teardown/profile changes ("deploy alerts", "switch profile", "bring up base") → use `/vss-deploy-profile`.
- Real-time alert/rule management requests → use `/vss-manage-alerts`.
Never route reports through VSS-agent `POST /generate`.
---
## Deployment prerequisite
**Mode A** needs the VSS **base** profile (VST + VLM NIM).
**Mode B** needs the VSS **alerts** profile (VA-MCP + Elasticsearch).
Probe:
```bash
# Mode A — VST + VLM reachability
curl -sf --max-time 5 "http://${HOST_IP}:30888/vst/api/v1/sensor/version" >/dev/null
# Mode B — VA-MCP
curl -sf --max-time 5 "http://${HOST_IP}:9901/" >/dev/null
```
If the probe fails, hand off to `/vss-deploy-profile` with `-p base` (Mode A) or `-p alerts` (Mode B). **Always** confirm the deploy with the user first.
---
## Clip URLs: VLM input vs browser report link
VST returns clip URLs using the agent-internal `${HOST_IP}:30888` host:port.
Keep that original URL as `VIDEO_URL` for local / in-cluster VLM frame pulls.
Do **not** rewrite the VLM input URL just to make it browser-playable.
Only create `BROWSER_CLIP_URL` for URLs shown in the rendered report. The
deploy layer exports the browser-facing host:port as `$VSS_PUBLIC_HOST` /
`$VSS_PUBLIC_PORT` (and scheme as `$VSS_PUBLIC_HTTP_PROTOCOL`) in every
profile `.env` — Brev or bare-metal — so the report-link rewrite is:
```bash
: "${VSS_PUBLIC_HOST:?Set VSS_PUBLIC_HOST before rewriting clip URLs}"
: "${VSS_PUBLIC_PORT:?Set VSS_PUBLIC_PORT before rewriting clip URLs}"
VSS_PUBLIC_HTTP_PROTOCOL="${VSS_PUBLIC_HTTP_PROTOCOL:-http}"
BROWSER_CLIP_URL=$(echo "$RAW_URL" | sed -E "s|^https?://[^/]+|${VSS_PUBLIC_HTTP_PROTOCOL}://${VSS_PUBLIC_HOST}:${VSS_PUBLIC_PORT}|")
```
If either required public host value is missing, omit the report-facing clip
link and call out that a browser-playable URL could not be produced; do not
block the local VLM analysis path. Apply the rewrite to **every clip URL
surfaced in the rendered report** (Mode A Step 4 Clip URL row; Mode B
per-incident clip sub-bullet). Leave the VLM `video_url` content block in Mode A
Step 3 on the original internal URL when the VLM is local / in-cluster.
---
## Mode A — Report on a recorded video clip
**If the VSS `lvs` profile is deployed** — `curl -sf --max-time 5 "http://${HOST_IP}:38111/v1/ready"` returns HTTP 200 — run `/vss-summarize-video` to produce the summary, then paste its output into the report template in Step 4 and skip Steps 1–3 (the VLM-direct path). Run Steps 1–3 only when `/v1/ready` is non-200.
### Step 1 — Resolve the clip URL
Hand off to `/vss-manage-video-io-storage` to:
1. List sensors and confirm the named `<sensor-id>` exists (upload first if not).
2. Fetch `/storage/<streamId>/timelines` for the recorded range when the user did not supply `startTime` / `endTime`.
3. Request a clip URL:
```bash
curl -s "http://${HOST_IP}:30888/vst/api/v1/storage/file/<streamId>/url?startTime=<startTime>&endTime=<endTime>&container=mp4&disableAudio=true" | jq -r .videoUrl
```
That gives a direct `mp4` URL that the local / in-cluster VLM can pull frames from. Bind it to `VIDEO_URL` (used by the VLM in Step 3) and set `RAW_URL="$VIDEO_URL"` before applying the report-link rewrite to produce `BROWSER_CLIP_URL` for Step 4 — the user's browser cannot reach `$VIDEO_URL` directly.
Mode A requires the selected VLM endpoint to be able to fetch `VIDEO_URL`.
Local NIM/RT-VLM deployments normally can; remote endpoints generally cannot
fetch `localhost`, private `HOST_IP`, or VST-internal URLs. If the live
`VLM_ENDPOINT` is remote, surface that reachability requirement instead of
making a chat request that will fail after `/v1/models` succeeds.
### Step 2 — Resolve VLM endpoint and model
The deploy may serve the VLM through either of two stacks. Both expose an OpenAI-compatible `chat/completions` API — pick whichever is live:
| Backend | Env vars | Typical host endpoint | Picked when |
|---|---|---|---|
| **NIM Cosmos** | `VLM_BASE_URL`, `VLM_NAME`, `VLM_MODE`, `VLM_MODEL_TYPE` | `${VLM_BASE_URL}/v1` (no trailing `/v1` on the env var; the agent appends it) | `VLM_MODEL_TYPE != rtvi` **and** `VLM_MODE` ∈ {`local`, `local_shared`, `remote`} **and** `VLM_BASE_URL` is non-empty |
| **RT-VLM Cosmos** | `RTVI_VLM_BASE_URL`, `RTVI_VLM_MODEL_TO_USE`, `VLM_MODEL_TYPE` | `${RTVI_VLM_BASE_URL}/v1` — if unset, derive from `${HOST_IP}` (`http://${HOST_IP}:8018/v1` for alerts, `http://${HOST_IP}:30082/v1` for base) | `VLM_MODEL_TYPE = rtvi`, or `VLM_MODE=none`, or `VLM_BASE_URL` empty; also the only path for `warehouse` |
Read the live values off the running agent container — do not guess:
```bash
docker exec vss-agent sh -lc '
for k in HOST_IP VLM_MODE VLM_MODEL_TYPE VLM_BASE_URL VLM_NAME RTVI_VLM_BASE_URL RTVI_VLM_MODEL_TO_USE; do
v="$(printenv "$k")"
[ -n "$v" ] && printf "%s=%s\n" "$k" "$v"
done
'
```
Do not require `RTVI_VLM_ENDPOINT` from `vss-agent` env; several profiles do not inject it.
Selection rule:
```bash
if [ "${VLM_MODEL_TYPE:-}" = "rtvi" ]; then
VLM_BACKEND="rtvlm"
VLM_ENDPOINT="${RTVI_VLM_BASE_URL:+${RTVI_VLM_BASE_URL%/}/v1}"
[ -z "${VLM_ENDPOINT}" ] && VLM_ENDPOINT="http://${HOST_IP}:8018/v1" # alerts default
VLM_MODEL="${RTVI_VLM_MODEL_TO_USE}"
elif [ -n "${VLM_BASE_URL}" ] && [ "${VLM_MODE}" != "none" ]; then
VLM_BACKEND="nim_cosmos"
VLM_ENDPOINT="${VLM_BASE_URL%/}/v1"
VLM_MODEL="${VLM_NAME}"
else
VLM_BACKEND="rtvlm"
VLM_ENDPOINT="${RTVI_VLM_BASE_URL:+${RTVI_VLM_BASE_URL%/}/v1}"
[ -z "${VLM_ENDPOINT}" ] && VLM_ENDPOINT="http://${HOST_IP}:30082/v1" # base default
VLM_MODEL="${RTVI_VLM_MODEL_TO_USE}"
fi
```
Probe `/v1/models` before sending a chat request to confirm the chosen endpoint is alive and the model is loaded:
```bash
curl -sf --max-time 5 "${VLM_ENDPOINT}/models" | jq -r '.data[].id'
```
If the probe fails or the listed ids don't include `${VLM_MODEL}`, fall back to the other backend (or surface the error — never silently pick a model that isn't on the server).
### Step 3 — Call the VLM directly
Use the OpenAI-compatible `chat/completions` endpoint with a `video_url` content block — the same payload shape **and multimodal settings** `video_understanding` builds in `src/vss_agents/tools/video_understanding.py` (`_build_vlm_messages` + the Cosmos `base_vlm.bind(...)` call).
The frame sampling and visual-token (pixel) budget must mirror the **live** `video_understanding` settings for the active profile. **Send `mm_processor_kwargs` and `media_io_kwargs`** so the direct call uses the same frame sampling and pixel budget as the in-agent `video_understanding` tool — omitting them lets the VLM apply its own defaults, so the output diverges from the agent path.
```bash
PROMPT='Describe in detail what happens in the video, with timestamps (start–end in seconds from clip start) for each segment or event. Cover scenes, objects, people, vehicles, and notable actions.'
# Reasoning is OFF by default — matches the base-profile video_understanding config (`reasoning: false`).
# video_understanding.py uses config.reasoning unless the caller overrides it, so default to non-reasoning.
# Append the Cosmos Reason 2 reasoning suffix ONLY when the user explicitly asks for reasoning
# (drop it for non-cosmos-reason2 VLMs). With reasoning off, the response has no <think> block.
if [ "${REASONING:-false}" = "true" ]; then
PROMPT="${PROMPT}
Answer the question using the following format:
<think>
Your reasoning.
</think>
Write your final answer immediately after the </think> tag."
fi
# If Step 3 is run standalone, derive missing backend from current env/model.
[ -z "${VLM_BACKEND:-}" ] && {
if [ "${VLM_MODEL_TYPE:-}" = "rtvi" ]; then
VLM_BACKEND="rtvlm"
elif [[ "${VLM_MODEL:-}" == nvidia/cosmos* ]]; then
VLM_BACKEND="nim_cosmos"
else
VLM_BACKEND="rtvlm"
fi
}
# Multimodal settings — resolve from the live agent config file path, not hardcoded candidates.
CFG_JSON=$(
docker exec vss-agent python3 -c '
import json, os, yaml
p = os.getenv("VSS_AGENT_CONFIG_FILE")
if not p:
raise SystemExit("VSS_AGENT_CONFIG_FILE is not set in vss-agent")
if not os.path.isabs(p):
p = os.path.join("/vss-agent", p.lstrip("./"))
with open(p, encoding="utf-8") as f:
cfg = yaml.safe_load(f) or {}
vu = (cfg.get("functions", {}) or {}).get("video_understanding", {}) or {}
print(json.dumps({
"max_fps": int(vu.get("max_fps", 2)),
"max_frames": int(vu.get("max_frames", 30)),
"min_pixels": int(vu.get("min_pixels", 3136)),
"max_pixels": int(vu.get("max_pixels", 8388608)),
}))
')
)
[ -n "${CFG_JSON}" ] || { echo "Failed to read video_understanding config from vss-agent"; exit 1; }
jq -e . >/dev/null <<< "${CFG_JSON}" || { echo "Invalid config JSON from vss-agent"; exit 1; }
MAX_FPS="$(jq -r '.max_fps' <<< "${CFG_JSON}")"
MAX_FRAMES="$(jq -r '.max_frames' <<< "${CFG_JSON}")"
MIN_PIXELS="$(jq -r '.min_pixels' <<< "${CFG_JSON}")"
MAX_PIXELS="$(jq -r '.max_pixels' <<< "${CFG_JSON}")"
# num_frames = min(int(clip_seconds) * max_fps, max_frames), min 1 — matches video_understanding.py.
# clip_seconds (Step 1 endTime-startTime) may be fractional; truncate to integer seconds — bash $((...))
# is integer-only and errors on "15.0"/"1.5". Default 15s -> caps at MAX_FRAMES.
CLIP_SECONDS=$(awk -v s="${CLIP_SECONDS:-15}" 'BEGIN{printf "%d", s}')
NUM_FRAMES=$(( CLIP_SECONDS * MAX_FPS ))
[ "$NUM_FRAMES" -gt "$MAX_FRAMES" ] && NUM_FRAMES=$MAX_FRAMES
[ "$NUM_FRAMES" -lt 1 ] && NUM_FRAMES=1
# Only apply Cosmos mm/media kwargs on the NIM Cosmos path.
# RT-VLM mode uses its own server-side preprocessing and should not receive these kwargs.
MM_KWARGS=""
if [ "${VLM_BACKEND}" = "nim_cosmos" ]; then
case "$VLM_MODEL" in
*cosmos-reason2*) MM_KWARGS=", \"mm_processor_kwargs\": {\"size\": {\"shortest_edge\": ${MIN_PIXELS}, \"longest_edge\": ${MAX_PIXELS}}}, \"media_io_kwargs\": {\"video\": {\"num_frames\": ${NUM_FRAMES}}}" ;;
*cosmos*) MM_KWARGS=", \"mm_processor_kwargs\": {\"videos_kwargs\": {\"min_pixels\": ${MIN_PIXELS}, \"max_pixels\": ${MAX_PIXELS}}}, \"media_io_kwargs\": {\"video\": {\"num_frames\": ${NUM_FRAMES}}}" ;;
*) MM_KWARGS="" ;;
esac
fi
curl -s --connect-timeout 5 --max-time 120 -X POST "${VLM_ENDPOINT}/chat/completions" \
-H "Content-Type: application/json" \
-d @- <<EOF | jq -r '.choices[0].message.content'
{
"model": $(jq -Rs . <<< "${VLM_MODEL}"),
"messages": [
{
"role": "user",
"content": [
{"type": "text", "text": $(jq -Rs . <<< "${PROMPT}")},
{"type": "video_url", "video_url": {"url": $(jq -Rs . <<< "${VIDEO_URL}")}}
]
}
],
"max_tokens": 1024,
"temperature": 0.0${MM_KWARGS}
}
EOF
```
> The kwargs block is backend-aware: on `nim_cosmos`, Reason2 variants (`nvidia/cosmos-reason2*`) use `mm_processor_kwargs.size{shortest_edge,longest_edge}` and other NIM Cosmos variants (`nvidia/cosmos*`) use `mm_processor_kwargs.videos_kwargs{min_pixels,max_pixels}`; both also send `media_io_kwargs.video.num_frames`. On `rtvlm`, no Cosmos kwargs are sent.
If the VLM returns a `<think>…</think>` block (Cosmos Reason reasoning mode), keep only the text after `</think>` as the report body.
### Step 4 — Fill the Video Analysis Report template
Copy [`assets/video-analysis-report.md`](assets/video-analysis-report.md), fill every placeholder, and return the rendered markdown to the user. Keep the source asset unchanged. Before rendering, verify `BROWSER_CLIP_URL` is set and non-empty, then replace `<BROWSER_CLIP_URL>` with that exact value in the `Clip URL` row. Never leave the placeholder in the output, never include template instructions in a filled cell, and never use the raw `HOST_IP:30888` URL.
---
## Mode B — Report on incidents in a time range
### Step 1 — Resolve the time range and (optionally) sensor
- `start_time` / `end_time` must be ISO 8601 UTC (`YYYY-MM-DDTHH:MM:SS.sssZ`). Resolve relative phrases ("last hour", "today") against the current host clock.
- If the user names a sensor, capture it as `source` + `source_type=sensor`. Otherwise leave both unset for an all-sensors query.
### Step 2 — Fetch incidents via `/vss-query-analytics`
Hand off to `/vss-query-analytics` (initialize → `tools/call`) with:
```json
{
"jsonrpc": "2.0",
"method": "tools/call",
"params": {
"name": "video_analytics__get_incidents",
"arguments": {
"source": "<sensor-id-or-omit>",
"source_type": "sensor",
"start_time": "<ISO>",
"end_time": "<ISO>",
"max_count": 100,
"includes": ["objectIds", "info"]
}
},
"id": 1
}
```
Read-only boundary (mandatory):
- Mode B is strictly read-only analytics retrieval. Never write, seed, backfill, or mutate Elasticsearch/VA data.
- Forbidden examples: indexing synthetic incidents, replaying fixture payloads into ES, calling write/update/delete APIs to "make data available" for the report.
- If no incidents exist for the requested range/scope, handle as empty results (see below); do not fabricate data.
For each incident keep: `id`, `sensorId`, `timestamp`, `end`, `category`, `place.name`, `info.verdict`, `info.reasoning`, `objectIds`, and the clip URL (commonly `info.clip_url`, `clip_url`, or whichever clip-pointer field the response carries). **Apply the `$VSS_PUBLIC_HOST:$VSS_PUBLIC_PORT` rewrite (see *Browser-playable clip URL* above) to every clip URL before pasting it into the report** — the raw value is a `HOST_IP:30888` URL the user's browser cannot reach.
### Step 3 — Fill the Incident Range Report template
Copy [`assets/incident-range-report.md`](assets/incident-range-report.md), then group by sensor (or by category if no sensor scope), tally verdicts, and list each incident with timestamp / category / verdict / reasoning. Keep the source asset unchanged. Every incident clip value must be a rewritten browser-playable URL; omit the clip line when the incident carries no clip URL. Never include template instructions in a filled cell.
If `get_incidents` returns zero results, STOP and return exactly a one-line empty-range statement naming the requested range and scope. Do not render the full Incident Range template, do not invent incidents, do not seed test data, and do not fall back to Mode A.
---
## Error Handling
- If a probe, `curl`, VLM call, or `/vss-query-analytics` request fails, stop the workflow and report the failing endpoint, HTTP status or command error, and the next useful recovery step. Do not fabricate a report from partial or missing data.
- If the VLM response is empty, malformed, or contains only a reasoning block, surface that response problem and suggest checking model readiness/logs before retrying.
- If a clip URL cannot be rewritten to the public host/port, omit it from the rendered report and call out that the browser-playable URL could not be produced.
- For Mode B, treat missing optional incident fields (`info.reasoning`, `objectIds`, clip URL) as omissions in the report, but treat missing `id`, `timestamp`, or `category` as a data-quality error that should be reported.
---
## Cross-Reference
- **`/vss-manage-video-io-storage`** — sensor list, timelines, and clip URL for Mode A Step 1.
- **`/vss-query-analytics`** — incident retrieval (and verdict / reasoning enrichment) for Mode B Step 2.
- **`/vss-ask-video`** — ad-hoc VLM Q&A on a single clip (not a structured report).
- **`/vss-summarize-video`** — used by Mode A to produce the summary body when the `lvs` profile is deployed; the report template (Step 4) is still filled here.
所有檔案
8 個檔案安裝 vss-generate-video-report
請下載並將技能檔案解壓縮至您的 .claude/skills/ 目錄中。
下載 ZIP複製儲存庫並將技能檔案複製到您的專案中。
git clone https://github.com/NVIDIA/skills/tree/main/skills/vss-generate-video-report # Copy SKILL.md to your .claude/skills/ directory
複製





首頁
