选项
首页首页 Skill 数据库管理 vss-generate-video-report

vss-generate-video-report

NVIDIA/skills NVIDIA/skills

通过将数据路由至 VLM 后端进行逐片段分析,或路由至分析后端生成事件范围报告,从而生成视频分析报告,同时支持部署配置文件验证和 URL 重写。

...展开全部
2
更新时间 2026-09-27

报告

通过将请求路由至两个后端之一来生成视频分析报告——切勿通过VSS 代理上的POST /generate方法进行操作。

模式 后端
A. 视频片段 /vss-manage-video-io-storage→ 视频片段 URL →VLM 聊天/自动补全
B. 事件范围 /vss-query-analytics→ 事件列表 → 叙述性报告

如果请求含糊不清(例如“生成 ”且未指定时间范围及事件描述),则默认采用模式A。仅当用户同时提及传感器和时间范围时才进行确认。请参阅下文示例,了解会路由至各模式的请求表述方式。

说明

  1. 选择模式——若请求涉及单个录制片段/传感器视频,则选模式A;若请求指明了时间范围或事件/警报,则选模式 B(请对照示例进行匹配)。
  2. 在“部署先决条件”下验证该模式的部署配置文件;若其探针测试失败,则移交至/vss-deploy-profile。
  3. 执行该模式的编号步骤——如下文的模式 A或模式 B。
  4. 在将视频片段嵌入报告之前,将所有面向用户的视频片段 URL 重写为$VSS_PUBLIC_HOST:$VSS_PUBLIC_PORT这一行格式(可在浏览器中播放的视频片段 URL)。
  5. 将渲染后的报告 Markdown 内容返回给用户。

面向评估人员的输出规范:

  • 模式 A 的顶部标题必须精确为# 视频分析报告。
  • 模式 B 的顶部标题必须精确为# 事件范围报告(绝不能是# 事件报告或以传感器命名的变体)。
  • 模式 B 必须包含## 基本信息,且必须包含模板中要求的精确行(报告标识符、范围、范围、总事件数、已确认/已拒绝/未核实)。

示例

  • “为该视频生成报告” / “关于 ” →模式 A
  • “分析 warehouse_01.mp4” / “针对上传的视频创建分析报告” →模式 A
  • “报告 12:31Z 至 12:32Z 期间的事件” →模式 B
  • “生成今日警报报告” / “过去 过去一小时内发生了哪些事件” →模式 B
  • “汇总 期间 和" → 模式B ”期间的警报” →模式 B

否定触发条件

当请求属于以下情况之一时,请勿使用此技能:

  • 针对视频片段的即兴视觉问答,且未明确要求生成报告(例如“卡车是什么颜色的?”、“00:12 发生了什么?”)→ 请使用/vss-ask-video。
  • 存档/语义相似性检索(“查找叉车”、“搜索所有视频中的跟车过近情况”)→ 请使用/vss-search-archive。
  • 无需生成报告的只读事件/指标查询 → 请使用/vss-query-analytics。
  • 部署/拆除/配置文件更改(“部署警报”、“切换配置文件”、“启动基础环境”)→ 请使用/vss-deploy-profile。
  • 实时警报/规则管理请求 → 使用/vss-manage-alerts。

切勿通过 VSS-agentPOST /generate 路由报告。

部署先决条件

模式 A需要 VSS基础配置文件(VST + VLM NIM)。 模式 B需要 VSS告警配置文件(VA-MCP + Elasticsearch)。

测试:

# 模式 A — VST + VLM 可达性
curl -sf --max-time 5 "http://${HOST_IP}:30888/vst/api/v1/sensor/version" >/dev/null

# 模式 B — VA-MCP
curl -sf --max-time 5 "http://${HOST_IP}:9901/" >/dev/null

如果探测失败,请使用-p base(模式 A)或-p alerts(模式 B)参数将任务移交至/vss-deploy-profile。请务必先与用户确认部署事宜。

片段 URL:VLM 输入与浏览器报告链接

VST 使用代理内部的${HOST_IP}:30888主机:端口返回剪辑 URL。 请保留该原始 URL 作为VIDEO_URL,用于本地或集群内的 VLM 帧提取。 请勿仅为了使其可在浏览器中播放而重写 VLM 输入 URL。

仅为渲染报告中显示的 URL 创建BROWSER_CLIP_URL。 部署层将面向浏览器的主机:端口导出为$VSS_PUBLIC_HOST/ $VSS_PUBLIC_PORT(协议格式为$VSS_PUBLIC_HTTP_PROTOCOL)的形式, 在每个配置文件.env中(无论是 Brev 还是裸机环境),因此报告链接的重写规则为:

: "${VSS_PUBLIC_HOST:?请在重写片段 URL 之前设置 VSS_PUBLIC_HOST}"
: "${VSS_PUBLIC_PORT:?请在重写片段 URL 之前设置 VSS_PUBLIC_PORT}"
VSS_PUBLIC_HTTP_PROTOCOL="${VSS_PUBLIC_HTTP_PROTOCOL:-http}"
BROWSER_CLIP_URL=$(echo "$RAW_URL" | sed -E "s|^https?://[^/]+|${VSS_PUBLIC_HTTP_PROTOCOL}://${VSS_PUBLIC_HOST}:${VSS_PUBLIC_PORT}|")

如果任一必需的公共主机值缺失,则省略面向报告的片段 链接,并指出无法生成可在浏览器中播放的 URL;不要 阻断本地 VLM 分析路径。对渲染报告中显示的每个片段 URL 应用重写(模式 A 第 4 步的“片段 URL”行;模式 B 的“按事件”片段子项目)。当 VLM 为本地或位于集群内时, 请将模式 A 第 3 步中的 VLMvideo_url内容块保留为原始内部 URL。

模式 A — 关于录制视频片段的报告

如果已部署 VSSlvs配置文件—执行 curl -sf --max-time 5 "http://${HOST_IP}:38111/v1/ready"返回 HTTP 200 状态码 — 运行/vss-summarize-video生成摘要, 然后将输出粘贴到第 4 步的报告模板中,并跳过第 1–3 步(VLM 直接路径)。仅当/v1/ready返回非 200 状态码时才执行第 1–3 步。

步骤 1 — 解析片段 URL

将任务交由/vss-manage-video-io-storage处理:

  1. 列出传感器并确认指定 是否存在(若不存在,请先上传)。

  2. 当用户未提供startTime/endTime 时,获取/storage//timelines中的录制范围。

  3. 请求片段 URL:

    curl -s "http://${HOST_IP}:30888/vst/api/v1/storage/file//url?startTime=&endTime=&container=mp4&disableAudio=true" | jq -r .videoUrl
    

    这将生成一个直接的mp4链接,本地/集群内的 VLM 可从中提取帧。 将其绑定到VIDEO_URL(在第 3 步中由 VLM 使用),并在应用报告链接重写以生成第 4 步所需的BROWSER_CLIP_URL之前,设置RAW_URL="$VIDEO_URL"—— 用户的浏览器无法直接访问$VIDEO_URL。 模式 A 要求所选的 VLM 端点能够获取VIDEO_URL。 本地 NIM/RT-VLM 部署通常可以;远程端点通常无法 获取localhost、私有HOST_IP 或 VST 内部 URL。 如果实时 VLM_ENDPOINT是远程的,请明确该可达性要求,而不是 发起一个聊天请求——该请求在/v1/models成功后仍会失败。

步骤 2 — 解析 VLM 端点和模型

该部署可能通过以下两种架构之一提供 VLM 服务。两者均提供与 OpenAI 兼容的聊天/补全API —— 请选择其中已上线的任意一种:

后端 环境变量 典型主机端点 在何时选择
NIM Cosmos VLM_BASE_URL、VLM_NAME、VLM_MODE、VLM_MODEL_TYPE ${VLM_BASE_URL}/v1(环境变量中不带尾随的/v1;由代理自动追加) VLM_MODEL_TYPE ≠ rtvi 且 VLM_MODE∈ {local,local_shared,remote}且 VLM_BASE_URL不为空
RT-VLM Cosmos RTVI_VLM_BASE_URL、RTVI_VLM_MODEL_TO_USE、VLM_MODEL_TYPE ${RTVI_VLM_BASE_URL}/v1— 若未设置,则从${HOST_IP}推导(警报地址为http://${HOST_IP}:8018/v1, http://${HOST_IP}:30082/v1用于基础数据) VLM_MODEL_TYPE = rtvi,或VLM_MODE=none,或VLM_BASE_URL为空;同时也是数据仓库的唯一路径

从正在运行的代理容器中读取实时值——切勿猜测:

docker exec vss-agent sh -lc '
for k in HOST_IP VLM_MODE VLM_MODEL_TYPE VLM_BASE_URL VLM_NAME RTVI_VLM_BASE_URL RTVI_VLM_MODEL_TO_USE; do
  v="$(printenv "$k")"
  [ -n "$v" ] && printf "%s=%s\n" "$k" "$v"
done
'

无需从vss-agent环境变量中获取RTVI_VLM_ENDPOINT;部分配置文件不会注入该变量。

选择规则:

if [ "${VLM_MODEL_TYPE:-}" = "rtvi" ]; then
  VLM_BACKEND="rtvlm"
  VLM_ENDPOINT="${RTVI_VLM_BASE_URL:+${RTVI_VLM_BASE_URL%/}/v1}"
  [ -z "${VLM_ENDPOINT}" ] && VLM_ENDPOINT="http://${HOST_IP}:8018/v1"   # 警报默认值
  VLM_MODEL="${RTVI_VLM_MODEL_TO_USE}"
elif [ -n "${VLM_BASE_URL}" ] && [ "${VLM_MODE}" != "none" ]; then
  VLM_BACKEND="nim_cosmos"
  VLM_ENDPOINT="${VLM_BASE_URL%/}/v1"
  VLM_MODEL="${VLM_NAME}"
else
  VLM_BACKEND="rtvlm"
  VLM_ENDPOINT="${RTVI_VLM_BASE_URL:+${RTVI_VLM_BASE_URL%/}/v1}"
  [ -z "${VLM_ENDPOINT}" ] && VLM_ENDPOINT="http://${HOST_IP}:30082/v1"  # 默认基础设置
  VLM_MODEL="${RTVI_VLM_MODEL_TO_USE}"
fi

在发送聊天请求前,先探测/v1/models以确认所选端点处于活动状态且模型已加载:

curl -sf --max-time 5 "${VLM_ENDPOINT}/models" | jq -r '.data[].id'

如果探测失败,或者列出的 ID 中不包含${VLM_MODEL},则回退到另一个后端(或报错——切勿在服务器上不存在模型的情况下默认选用该模型)。

步骤 3 — 直接调用 VLM

使用兼容 OpenAI 的chat/completions端点,并包含video_url内容块——其有效负载结构和多模态设置与 src/vss_agents/tools/video_understanding.py中video_understanding构建时采用的完全一致 (_build_vlm_messages以及 Cosmos 的base_vlm.bind(...)调用)。

帧采样和视觉令牌(像素)预算必须与当前配置文件的实时 video_understanding设置保持一致。发送mm_processor_kwargs和media_io_kwargs,以便直接调用采用与代理内部video_understanding工具相同的帧采样和像素配额——省略这些参数将使 VLM 应用其自身的默认值,导致输出结果与代理路径产生偏差。

PROMPT='详细描述视频中的事件,并为每个片段或事件提供时间戳(以秒为单位,从片段开始算起的起止时间)。涵盖场景、物体、人物、车辆以及值得注意的动作。'

# 推理功能默认关闭——与基础配置文件 video_understanding 的配置(`reasoning: false`)一致。
# video_understanding.py 会使用 config.reasoning,除非调用方进行覆盖,因此默认不启用推理。
# 仅当用户明确要求提供推理时,才附加 Cosmos Reason 2 推理后缀
# (对于非 Cosmos Reason 2 类型的视觉语言模型,则省略该后缀)。当推理功能关闭时,响应中不包含 `reasoning ` 块。
if [ "${REASONING:-false}" = "true" ]; then
PROMPT="${PROMPT}

请按以下格式回答问题:


您的推理过程。


请在  标签后立即写下最终答案 。"
fi

# 如果单独运行第 3 步,则根据当前环境/模型推导缺失的后端。
[ -z "${VLM_BACKEND:-}" ] && {
  if [ "${VLM_MODEL_TYPE:-}" = "rtvi" ]; then
    VLM_BACKEND="rtvlm"
  elif [[ "${VLM_MODEL:-}" == nvidia/cosmos* ]]; then
    VLM_BACKEND="nim_cosmos"
  else
    VLM_BACKEND="rtvlm"
  fi
}

# 多模态设置 — 从实时代理配置文件路径中解析,而非使用硬编码的候选项。
CFG_JSON=$(
docker exec vss-agent python3 -c '
import json, os, yaml
p = os.getenv("VSS_AGENT_CONFIG_FILE")
if not p:
    raise SystemExit("vss-agent 中未设置 VSS_AGENT_CONFIG_FILE")
if not os.path.isabs(p):
    p = os.path.join("/vss-agent", p.lstrip("./"))
with open(p, encoding="utf-8") as f:
    cfg = yaml.safe_load(f) or {}
vu = (cfg.get("functions", {}) or {}).get("video_understanding", {}) or {}
print(json.dumps({
    "max_fps": int(vu.get("max_fps", 2)),
    "max_frames": int(vu.get("max_frames", 30)),
    "min_pixels": int(vu.get("min_pixels", 3136)),
    "max_pixels": int(vu.get("max_pixels", 8388608)),
}))
')
)
[ -n "${CFG_JSON}" ] || { echo "无法从 vss-agent 读取 video_understanding 配置"; exit 1; }
jq -e . >/dev/null <<< "${CFG_JSON}" || { echo "来自 vss-agent 的配置 JSON 无效"; exit 1; }
MAX_FPS="$(jq -r '.max_fps' <<< "${CFG_JSON}")"
MAX_FRAMES="$(jq -r '.max_frames' <<< "${CFG_JSON}")"
MIN_PIXELS="$(jq -r '.min_pixels' <<< "${CFG_JSON}")"
MAX_PIXELS="$(jq -r '.max_pixels' <<< "${CFG_JSON}")"

# num_frames = min(int(clip_seconds) * max_fps, max_frames),下限为 1 — 与 video_understanding.py 保持一致。
# clip_seconds(步骤 1 的 endTime - startTime)可能包含小数部分;截断为整数秒 — bash $((...))
# 仅支持整数,遇到 "15.0"/"1.5" 会报错。默认 15 秒 -> 上限为 MAX_FRAMES。
CLIP_SECONDS=$(awk -v s="${CLIP_SECONDS:-15}" 'BEGIN{printf "%d", s}')
NUM_FRAMES=$(( CLIP_SECONDS * MAX_FPS ))
[ "$NUM_FRAMES" -gt "$MAX_FRAMES" ] && NUM_FRAMES=$MAX_FRAMES
[ "$NUM_FRAMES" -lt 1 ] && NUM_FRAMES=1

# 仅在 NIM Cosmos 路径上应用 Cosmos mm/media 关键参数。
# RT-VLM 模式使用其自身的服务器端预处理,不应接收这些关键参数。
MM_KWARGS=""
if [ "${VLM_BACKEND}" = "nim_cosmos" ]; then
  case "$VLM_MODEL" in
    *cosmos-reason2*) MM_KWARGS=", \"mm_processor_kwargs\": {\"size\": {\"shortest_edge\": ${MIN_PIXELS}, \"longest_edge\": ${MAX_PIXELS}}}, \"media_io_kwargs\": {\"video\": {\"num_frames\": ${NUM_FRAMES}}}" ;;
    *cosmos*)         MM_KWARGS=", \"mm_processor_kwargs\": {\"videos_kwargs\": {\"min_pixels\": ${MIN_PIXELS}, \"max_pixels\": ${MAX_PIXELS}}}, \"media_io_kwargs\": {\"video\": {\"num_frames\": ${NUM_FRAMES}}}" ;;
    *)                      MM_KWARGS="" ;;
  esac
fi

curl -s --connect-timeout 5 --max-time 120 -X POST "${VLM_ENDPOINT}/chat/completions" \
  -H "Content-Type: application/json" \
  -d @- <<EOF | jq -r '.choices[0].message.content'
{
  "model": $(jq -Rs . <<< "${VLM_MODEL}"),
  "messages": [
    {
      "role": "user",
      "content": [
        {"type": "text", "text": $(jq -Rs . <<< "${PROMPT}")},
        {"type": "video_url", "video_url": {"url": $(jq -Rs . <<< "${VIDEO_URL}")}}
      ]
    }
  ],
  "max_tokens": 1024,
  "temperature": 0.0${MM_KWARGS}
}
EOF

kwargs 代码块会根据后端进行适配:在nim_cosmos 上,Reason2 变体(nvidia/cosmos-reason2*)使用mm_processor_kwargs.size{shortest_edge,longest_edge},而其他 NIM Cosmos 变体(nvidia/cosmos*)使用mm_processor_kwargs.videos_kwargs{min_pixels,max_pixels};两者均会发送media_io_kwargs.video.num_frames。在rtvlm 上,不发送任何 Cosmos 命令行参数。

如果 VLM 返回一个 … 块(Cosmos Reason推理模式),则仅保留 后的文本作为报告正文。

步骤 4 — 填写视频分析报告模板

复制assets/video-analysis-report.md 文件,填写所有占位符,并将渲染后的 Markdown 内容返回给用户。保持源文件不变。在渲染之前,请确认BROWSER_CLIP_URL已设置且不为空,然后在“Clip URL”行中将 “Clip URL”行中的占位符替换为该确切值。输出中绝不保留占位符,已填写的单元格中绝不包含模板说明,且绝不使用原始的HOST_IP:30888URL。

模式 B — 生成特定时间范围内的事件报告

步骤 1 — 解析时间范围和(可选)传感器

  • start_time/end_time必须采用 ISO 8601 UTC 格式(YYYY-MM-DDTHH:MM:SS.sssZ)。将相对表述(如“过去一小时”、“今天”)转换为当前主机时钟的时间。
  • 如果用户指定了传感器,请将其捕获为source+source_type=sensor。否则,请将两者均留空,以执行所有传感器的查询。

步骤 2 — 通过/vss-query-analytics获取事件

将请求交由/vss-query-analytics处理(初始化 →tools/call),参数如下:

{
  "jsonrpc": "2.0",
  "method": "tools/call",
  "params": {
    "name": "video_analytics__get_incidents",
    "arguments": {
      "source": "",
      "source_type": "sensor",
      "start_time": "",
      "end_time": "",
      "max_count": 100,
      "includes": ["objectIds", "info"]
    }
  },
  "id": 1
}

只读边界(必选):

  • 模式 B 严格限定为只读分析数据检索。切勿对 Elasticsearch/VA 数据进行写入、初始化、回填或修改操作。
  • 禁止操作示例:将合成事件写入索引、将测试数据重放回 Elasticsearch、调用写入/更新/删除 API 以“为报告提供数据”。
  • 如果请求的范围/范围中不存在任何事件,则将其视为空结果(见下文);切勿伪造数据。

对于每个事件,请保留:id、sensorId、timestamp、end、category、place.name、info.verdict、info.reasoning、objectIds 以及片段 URL(通常为info.clip_url、clip_url 或响应中携带的任何片段指针字段)。在将每个片段 URL 粘贴到报告中之前,请对其应用$VSS_PUBLIC_HOST:$VSS_PUBLIC_PORT重写规则(参见上文“可在浏览器中播放的片段 URL”)——原始值是一个HOST_IP:30888格式的 URL,用户的浏览器无法访问该地址。

步骤 3 — 填写“事件范围报告”模板

复制assets/incident-range-report.md,然后按传感器分组(如果没有传感器范围,则按类别分组),统计判定结果,并列出每个事件的时间戳、类别、判定结果及理由。保持源文件不变。 每个事件片段值必须是经过重写、可由浏览器播放的 URL;若事件不包含片段 URL,则省略该行。切勿在已填写的单元格中包含模板说明。

如果get_incidents返回零条结果,请停止操作,并返回一条精确的空范围声明,其中需明确指明所请求的范围和范围。请勿渲染完整的“事件范围”模板,请勿虚构事件,请勿插入测试数据,也请勿回退到模式 A。

错误处理

  • 如果探针、curl、VLM 调用或/vss-query-analytics请求失败,请停止工作流,并报告失败的端点、HTTP 状态或命令错误,以及下一个有用的恢复步骤。请勿根据不完整或缺失的数据编造报告。
  • 如果 VLM 响应为空、格式不正确或仅包含推理块,请指出该响应问题,并建议在重试前检查模型就绪性/日志。
  • 如果无法将片段 URL 重写为公共主机/端口,请将其从生成的报告中省略,并注明无法生成可在浏览器中播放的 URL。
  • 对于模式 B,将缺失的可选事件字段(info.reasoning、objectIds、片段 URL)视为报告中的遗漏项,但将缺失的ID、时间戳或类别视为应上报的数据质量错误。

交叉引用

  • /vss-manage-video-io-storage— 模式 A 步骤 1 的传感器列表、时间线和片段 URL。
  • /vss-query-analytics— 模式 B 步骤 2 的事件检索(以及裁决/推理丰富)。
  • /vss-ask-video— 针对单个片段的临时 VLM 问答(非结构化报告)。
  • /vss-summarize-video— 模式 A 在部署lvs配置文件时用于生成摘要正文;报告模板(第 4 步)仍在此处填充。
在 GitHub 上查看
---
name: vss-generate-video-report
description: Generates video analysis reports by routing to a VLM backend for per-clip analysis or an analytics backend for incident-range reports, with deployment profile verification and URL rewriting.
license: Apache-2.0
---

# Report

Generate a video analysis report by routing to one of two backends — **never via** `POST /generate` on the VSS agent.

| Mode | Backend |
|---|---|
| **A. Video clip** | `/vss-manage-video-io-storage` → clip URL → **VLM chat/completions** |
| **B. Incident range** | `/vss-query-analytics` → incident list → narrative report |

If the request is ambiguous (e.g. "report on `<sensor>`" with no time range and no incident wording), default to **Mode A**. Ask only if the user mentions both a sensor and a time range. See **Examples** below for the request phrasings that route to each mode.

---

## Instructions

1. **Pick the mode** — Mode A for a single recorded clip/sensor video, Mode B when the request names a time range or incidents/alerts (match against *Examples*).
2. **Verify the deployment profile** for that mode under *Deployment prerequisite*; hand off to `/vss-deploy-profile` if its probe fails.
3. **Run that mode's numbered steps** — *Mode A* or *Mode B* below.
4. **Rewrite every user-facing clip URL** with the `$VSS_PUBLIC_HOST:$VSS_PUBLIC_PORT` one-liner (*Browser-playable clip URL*) before embedding it in the report.
5. **Return the rendered report markdown** to the user.

Output contract for evaluators:
- Mode A top title MUST be exactly `# Video Analysis Report`.
- Mode B top title MUST be exactly `# Incident Range Report` (never `# Incident Report` or sensor-named variants).
- Mode B MUST include `## Basic Information` with the exact required rows from the template (Report Identifier, Range, Scope, Total Incidents, Confirmed / Rejected / Unverified).

---

## Examples

- "Generate a report for this video" / "report on `<sensor-id>`" → **Mode A**
- "Analyze warehouse_01.mp4" / "create an analysis report on the uploaded video" → **Mode A**
- "Report on incidents from 12:31Z to 12:32Z" → **Mode B**
- "Report on alerts today" / "what incidents happened on `<sensor>` last hour" → **Mode B**
- "Summarize alerts on `<sensor>` between `<t1>` and `<t2>`" → **Mode B**

---

## Negative Triggers

Do **not** use this skill when the request is one of the following:

- Ad-hoc visual Q&A on a clip that do not ask explicitly for a report ("what color is the truck?", "what happens at 00:12?") → use `/vss-ask-video`.
- Archive/semantic similarity retrieval ("find forklifts", "search all videos for tailgating") → use `/vss-search-archive`.
- Read-only incident/metrics lookup without report rendering needs → use `/vss-query-analytics`.
- Deploy/teardown/profile changes ("deploy alerts", "switch profile", "bring up base") → use `/vss-deploy-profile`.
- Real-time alert/rule management requests → use `/vss-manage-alerts`.

Never route reports through VSS-agent `POST /generate`.

---

## Deployment prerequisite

**Mode A** needs the VSS **base** profile (VST + VLM NIM).
**Mode B** needs the VSS **alerts** profile (VA-MCP + Elasticsearch).

Probe:

```bash
# Mode A — VST + VLM reachability
curl -sf --max-time 5 "http://${HOST_IP}:30888/vst/api/v1/sensor/version" >/dev/null

# Mode B — VA-MCP
curl -sf --max-time 5 "http://${HOST_IP}:9901/" >/dev/null
```

If the probe fails, hand off to `/vss-deploy-profile` with `-p base` (Mode A) or `-p alerts` (Mode B). **Always** confirm the deploy with the user first.

---

## Clip URLs: VLM input vs browser report link

VST returns clip URLs using the agent-internal `${HOST_IP}:30888` host:port.
Keep that original URL as `VIDEO_URL` for local / in-cluster VLM frame pulls.
Do **not** rewrite the VLM input URL just to make it browser-playable.

Only create `BROWSER_CLIP_URL` for URLs shown in the rendered report. The
deploy layer exports the browser-facing host:port as `$VSS_PUBLIC_HOST` /
`$VSS_PUBLIC_PORT` (and scheme as `$VSS_PUBLIC_HTTP_PROTOCOL`) in every
profile `.env` — Brev or bare-metal — so the report-link rewrite is:

```bash
: "${VSS_PUBLIC_HOST:?Set VSS_PUBLIC_HOST before rewriting clip URLs}"
: "${VSS_PUBLIC_PORT:?Set VSS_PUBLIC_PORT before rewriting clip URLs}"
VSS_PUBLIC_HTTP_PROTOCOL="${VSS_PUBLIC_HTTP_PROTOCOL:-http}"
BROWSER_CLIP_URL=$(echo "$RAW_URL" | sed -E "s|^https?://[^/]+|${VSS_PUBLIC_HTTP_PROTOCOL}://${VSS_PUBLIC_HOST}:${VSS_PUBLIC_PORT}|")
```

If either required public host value is missing, omit the report-facing clip
link and call out that a browser-playable URL could not be produced; do not
block the local VLM analysis path. Apply the rewrite to **every clip URL
surfaced in the rendered report** (Mode A Step 4 Clip URL row; Mode B
per-incident clip sub-bullet). Leave the VLM `video_url` content block in Mode A
Step 3 on the original internal URL when the VLM is local / in-cluster.

---

## Mode A — Report on a recorded video clip

**If the VSS `lvs` profile is deployed** — `curl -sf --max-time 5 "http://${HOST_IP}:38111/v1/ready"` returns HTTP 200 — run `/vss-summarize-video` to produce the summary, then paste its output into the report template in Step 4 and skip Steps 1–3 (the VLM-direct path). Run Steps 1–3 only when `/v1/ready` is non-200.

### Step 1 — Resolve the clip URL

Hand off to `/vss-manage-video-io-storage` to:

1. List sensors and confirm the named `<sensor-id>` exists (upload first if not).
2. Fetch `/storage/<streamId>/timelines` for the recorded range when the user did not supply `startTime` / `endTime`.
3. Request a clip URL:

   ```bash
   curl -s "http://${HOST_IP}:30888/vst/api/v1/storage/file/<streamId>/url?startTime=<startTime>&endTime=<endTime>&container=mp4&disableAudio=true" | jq -r .videoUrl
   ```

   That gives a direct `mp4` URL that the local / in-cluster VLM can pull frames from. Bind it to `VIDEO_URL` (used by the VLM in Step 3) and set `RAW_URL="$VIDEO_URL"` before applying the report-link rewrite to produce `BROWSER_CLIP_URL` for Step 4 — the user's browser cannot reach `$VIDEO_URL` directly.
   Mode A requires the selected VLM endpoint to be able to fetch `VIDEO_URL`.
   Local NIM/RT-VLM deployments normally can; remote endpoints generally cannot
   fetch `localhost`, private `HOST_IP`, or VST-internal URLs. If the live
   `VLM_ENDPOINT` is remote, surface that reachability requirement instead of
   making a chat request that will fail after `/v1/models` succeeds.

### Step 2 — Resolve VLM endpoint and model

The deploy may serve the VLM through either of two stacks. Both expose an OpenAI-compatible `chat/completions` API — pick whichever is live:

| Backend | Env vars | Typical host endpoint | Picked when |
|---|---|---|---|
| **NIM Cosmos** | `VLM_BASE_URL`, `VLM_NAME`, `VLM_MODE`, `VLM_MODEL_TYPE` | `${VLM_BASE_URL}/v1` (no trailing `/v1` on the env var; the agent appends it) | `VLM_MODEL_TYPE != rtvi` **and** `VLM_MODE` ∈ {`local`, `local_shared`, `remote`} **and** `VLM_BASE_URL` is non-empty |
| **RT-VLM Cosmos** | `RTVI_VLM_BASE_URL`, `RTVI_VLM_MODEL_TO_USE`, `VLM_MODEL_TYPE` | `${RTVI_VLM_BASE_URL}/v1` — if unset, derive from `${HOST_IP}` (`http://${HOST_IP}:8018/v1` for alerts, `http://${HOST_IP}:30082/v1` for base) | `VLM_MODEL_TYPE = rtvi`, or `VLM_MODE=none`, or `VLM_BASE_URL` empty; also the only path for `warehouse` |

Read the live values off the running agent container — do not guess:

```bash
docker exec vss-agent sh -lc '
for k in HOST_IP VLM_MODE VLM_MODEL_TYPE VLM_BASE_URL VLM_NAME RTVI_VLM_BASE_URL RTVI_VLM_MODEL_TO_USE; do
  v="$(printenv "$k")"
  [ -n "$v" ] && printf "%s=%s\n" "$k" "$v"
done
'
```

Do not require `RTVI_VLM_ENDPOINT` from `vss-agent` env; several profiles do not inject it.

Selection rule:

```bash
if [ "${VLM_MODEL_TYPE:-}" = "rtvi" ]; then
  VLM_BACKEND="rtvlm"
  VLM_ENDPOINT="${RTVI_VLM_BASE_URL:+${RTVI_VLM_BASE_URL%/}/v1}"
  [ -z "${VLM_ENDPOINT}" ] && VLM_ENDPOINT="http://${HOST_IP}:8018/v1"   # alerts default
  VLM_MODEL="${RTVI_VLM_MODEL_TO_USE}"
elif [ -n "${VLM_BASE_URL}" ] && [ "${VLM_MODE}" != "none" ]; then
  VLM_BACKEND="nim_cosmos"
  VLM_ENDPOINT="${VLM_BASE_URL%/}/v1"
  VLM_MODEL="${VLM_NAME}"
else
  VLM_BACKEND="rtvlm"
  VLM_ENDPOINT="${RTVI_VLM_BASE_URL:+${RTVI_VLM_BASE_URL%/}/v1}"
  [ -z "${VLM_ENDPOINT}" ] && VLM_ENDPOINT="http://${HOST_IP}:30082/v1"  # base default
  VLM_MODEL="${RTVI_VLM_MODEL_TO_USE}"
fi
```

Probe `/v1/models` before sending a chat request to confirm the chosen endpoint is alive and the model is loaded:

```bash
curl -sf --max-time 5 "${VLM_ENDPOINT}/models" | jq -r '.data[].id'
```

If the probe fails or the listed ids don't include `${VLM_MODEL}`, fall back to the other backend (or surface the error — never silently pick a model that isn't on the server).

### Step 3 — Call the VLM directly

Use the OpenAI-compatible `chat/completions` endpoint with a `video_url` content block — the same payload shape **and multimodal settings** `video_understanding` builds in `src/vss_agents/tools/video_understanding.py` (`_build_vlm_messages` + the Cosmos `base_vlm.bind(...)` call).

The frame sampling and visual-token (pixel) budget must mirror the **live** `video_understanding` settings for the active profile. **Send `mm_processor_kwargs` and `media_io_kwargs`** so the direct call uses the same frame sampling and pixel budget as the in-agent `video_understanding` tool — omitting them lets the VLM apply its own defaults, so the output diverges from the agent path.

```bash
PROMPT='Describe in detail what happens in the video, with timestamps (start–end in seconds from clip start) for each segment or event. Cover scenes, objects, people, vehicles, and notable actions.'

# Reasoning is OFF by default — matches the base-profile video_understanding config (`reasoning: false`).
# video_understanding.py uses config.reasoning unless the caller overrides it, so default to non-reasoning.
# Append the Cosmos Reason 2 reasoning suffix ONLY when the user explicitly asks for reasoning
# (drop it for non-cosmos-reason2 VLMs). With reasoning off, the response has no <think> block.
if [ "${REASONING:-false}" = "true" ]; then
PROMPT="${PROMPT}

Answer the question using the following format:

<think>
Your reasoning.
</think>

Write your final answer immediately after the </think> tag."
fi

# If Step 3 is run standalone, derive missing backend from current env/model.
[ -z "${VLM_BACKEND:-}" ] && {
  if [ "${VLM_MODEL_TYPE:-}" = "rtvi" ]; then
    VLM_BACKEND="rtvlm"
  elif [[ "${VLM_MODEL:-}" == nvidia/cosmos* ]]; then
    VLM_BACKEND="nim_cosmos"
  else
    VLM_BACKEND="rtvlm"
  fi
}

# Multimodal settings — resolve from the live agent config file path, not hardcoded candidates.
CFG_JSON=$(
docker exec vss-agent python3 -c '
import json, os, yaml
p = os.getenv("VSS_AGENT_CONFIG_FILE")
if not p:
    raise SystemExit("VSS_AGENT_CONFIG_FILE is not set in vss-agent")
if not os.path.isabs(p):
    p = os.path.join("/vss-agent", p.lstrip("./"))
with open(p, encoding="utf-8") as f:
    cfg = yaml.safe_load(f) or {}
vu = (cfg.get("functions", {}) or {}).get("video_understanding", {}) or {}
print(json.dumps({
    "max_fps": int(vu.get("max_fps", 2)),
    "max_frames": int(vu.get("max_frames", 30)),
    "min_pixels": int(vu.get("min_pixels", 3136)),
    "max_pixels": int(vu.get("max_pixels", 8388608)),
}))
')
)
[ -n "${CFG_JSON}" ] || { echo "Failed to read video_understanding config from vss-agent"; exit 1; }
jq -e . >/dev/null <<< "${CFG_JSON}" || { echo "Invalid config JSON from vss-agent"; exit 1; }
MAX_FPS="$(jq -r '.max_fps' <<< "${CFG_JSON}")"
MAX_FRAMES="$(jq -r '.max_frames' <<< "${CFG_JSON}")"
MIN_PIXELS="$(jq -r '.min_pixels' <<< "${CFG_JSON}")"
MAX_PIXELS="$(jq -r '.max_pixels' <<< "${CFG_JSON}")"

# num_frames = min(int(clip_seconds) * max_fps, max_frames), min 1 — matches video_understanding.py.
# clip_seconds (Step 1 endTime-startTime) may be fractional; truncate to integer seconds — bash $((...))
# is integer-only and errors on "15.0"/"1.5". Default 15s -> caps at MAX_FRAMES.
CLIP_SECONDS=$(awk -v s="${CLIP_SECONDS:-15}" 'BEGIN{printf "%d", s}')
NUM_FRAMES=$(( CLIP_SECONDS * MAX_FPS ))
[ "$NUM_FRAMES" -gt "$MAX_FRAMES" ] && NUM_FRAMES=$MAX_FRAMES
[ "$NUM_FRAMES" -lt 1 ] && NUM_FRAMES=1

# Only apply Cosmos mm/media kwargs on the NIM Cosmos path.
# RT-VLM mode uses its own server-side preprocessing and should not receive these kwargs.
MM_KWARGS=""
if [ "${VLM_BACKEND}" = "nim_cosmos" ]; then
  case "$VLM_MODEL" in
    *cosmos-reason2*) MM_KWARGS=", \"mm_processor_kwargs\": {\"size\": {\"shortest_edge\": ${MIN_PIXELS}, \"longest_edge\": ${MAX_PIXELS}}}, \"media_io_kwargs\": {\"video\": {\"num_frames\": ${NUM_FRAMES}}}" ;;
    *cosmos*)         MM_KWARGS=", \"mm_processor_kwargs\": {\"videos_kwargs\": {\"min_pixels\": ${MIN_PIXELS}, \"max_pixels\": ${MAX_PIXELS}}}, \"media_io_kwargs\": {\"video\": {\"num_frames\": ${NUM_FRAMES}}}" ;;
    *)                      MM_KWARGS="" ;;
  esac
fi

curl -s --connect-timeout 5 --max-time 120 -X POST "${VLM_ENDPOINT}/chat/completions" \
  -H "Content-Type: application/json" \
  -d @- <<EOF | jq -r '.choices[0].message.content'
{
  "model": $(jq -Rs . <<< "${VLM_MODEL}"),
  "messages": [
    {
      "role": "user",
      "content": [
        {"type": "text", "text": $(jq -Rs . <<< "${PROMPT}")},
        {"type": "video_url", "video_url": {"url": $(jq -Rs . <<< "${VIDEO_URL}")}}
      ]
    }
  ],
  "max_tokens": 1024,
  "temperature": 0.0${MM_KWARGS}
}
EOF
```

> The kwargs block is backend-aware: on `nim_cosmos`, Reason2 variants (`nvidia/cosmos-reason2*`) use `mm_processor_kwargs.size{shortest_edge,longest_edge}` and other NIM Cosmos variants (`nvidia/cosmos*`) use `mm_processor_kwargs.videos_kwargs{min_pixels,max_pixels}`; both also send `media_io_kwargs.video.num_frames`. On `rtvlm`, no Cosmos kwargs are sent.

If the VLM returns a `<think>…</think>` block (Cosmos Reason reasoning mode), keep only the text after `</think>` as the report body.

### Step 4 — Fill the Video Analysis Report template

Copy [`assets/video-analysis-report.md`](assets/video-analysis-report.md), fill every placeholder, and return the rendered markdown to the user. Keep the source asset unchanged. Before rendering, verify `BROWSER_CLIP_URL` is set and non-empty, then replace `<BROWSER_CLIP_URL>` with that exact value in the `Clip URL` row. Never leave the placeholder in the output, never include template instructions in a filled cell, and never use the raw `HOST_IP:30888` URL.

---

## Mode B — Report on incidents in a time range

### Step 1 — Resolve the time range and (optionally) sensor

- `start_time` / `end_time` must be ISO 8601 UTC (`YYYY-MM-DDTHH:MM:SS.sssZ`). Resolve relative phrases ("last hour", "today") against the current host clock.
- If the user names a sensor, capture it as `source` + `source_type=sensor`. Otherwise leave both unset for an all-sensors query.

### Step 2 — Fetch incidents via `/vss-query-analytics`

Hand off to `/vss-query-analytics` (initialize → `tools/call`) with:

```json
{
  "jsonrpc": "2.0",
  "method": "tools/call",
  "params": {
    "name": "video_analytics__get_incidents",
    "arguments": {
      "source": "<sensor-id-or-omit>",
      "source_type": "sensor",
      "start_time": "<ISO>",
      "end_time": "<ISO>",
      "max_count": 100,
      "includes": ["objectIds", "info"]
    }
  },
  "id": 1
}
```

Read-only boundary (mandatory):
- Mode B is strictly read-only analytics retrieval. Never write, seed, backfill, or mutate Elasticsearch/VA data.
- Forbidden examples: indexing synthetic incidents, replaying fixture payloads into ES, calling write/update/delete APIs to "make data available" for the report.
- If no incidents exist for the requested range/scope, handle as empty results (see below); do not fabricate data.

For each incident keep: `id`, `sensorId`, `timestamp`, `end`, `category`, `place.name`, `info.verdict`, `info.reasoning`, `objectIds`, and the clip URL (commonly `info.clip_url`, `clip_url`, or whichever clip-pointer field the response carries). **Apply the `$VSS_PUBLIC_HOST:$VSS_PUBLIC_PORT` rewrite (see *Browser-playable clip URL* above) to every clip URL before pasting it into the report** — the raw value is a `HOST_IP:30888` URL the user's browser cannot reach.

### Step 3 — Fill the Incident Range Report template

Copy [`assets/incident-range-report.md`](assets/incident-range-report.md), then group by sensor (or by category if no sensor scope), tally verdicts, and list each incident with timestamp / category / verdict / reasoning. Keep the source asset unchanged. Every incident clip value must be a rewritten browser-playable URL; omit the clip line when the incident carries no clip URL. Never include template instructions in a filled cell.

If `get_incidents` returns zero results, STOP and return exactly a one-line empty-range statement naming the requested range and scope. Do not render the full Incident Range template, do not invent incidents, do not seed test data, and do not fall back to Mode A.

---

## Error Handling

- If a probe, `curl`, VLM call, or `/vss-query-analytics` request fails, stop the workflow and report the failing endpoint, HTTP status or command error, and the next useful recovery step. Do not fabricate a report from partial or missing data.
- If the VLM response is empty, malformed, or contains only a reasoning block, surface that response problem and suggest checking model readiness/logs before retrying.
- If a clip URL cannot be rewritten to the public host/port, omit it from the rendered report and call out that the browser-playable URL could not be produced.
- For Mode B, treat missing optional incident fields (`info.reasoning`, `objectIds`, clip URL) as omissions in the report, but treat missing `id`, `timestamp`, or `category` as a data-quality error that should be reported.

---

## Cross-Reference

- **`/vss-manage-video-io-storage`** — sensor list, timelines, and clip URL for Mode A Step 1.
- **`/vss-query-analytics`** — incident retrieval (and verdict / reasoning enrichment) for Mode B Step 2.
- **`/vss-ask-video`** — ad-hoc VLM Q&A on a single clip (not a structured report).
- **`/vss-summarize-video`** — used by Mode A to produce the summary body when the `lvs` profile is deployed; the report template (Step 4) is still filled here.

安装 vss-generate-video-report

下载技能文件并将其解压到 .claude/skills/ 目录下。

下载ZIP

克隆仓库并复制技能文件到您的项目中。

git clone https://github.com/NVIDIA/skills/tree/main/skills/vss-generate-video-report # Copy SKILL.md to your .claude/skills/ directory

复制 复制
快速设置: 将技能文件夹复制到 .claude/skills/ Claude 将自动检测并使用该技能
仓库 NVIDIA/skills

相关技能

microservices-patterns
更新时间 2026-06-29
jpa-patterns
更新时间 2026-06-30
fabric-lakehouse
更新时间 2026-06-30
prisma-expert
更新时间 2026-06-29
OR