选项
首页首页 Skill 安全 jetson-memory-audit

jetson-memory-audit

NVIDIA/skills NVIDIA/skills

测量 Jetson 的 DRAM 和 NvMap 使用情况,捕获基准测试前的数据和基准测试后的数据,并利用实时审计数据验证内存回收情况。

...展开全部
15
更新时间 2026-09-24

Jetson 内存审计

针对 Jetson 的只读内存(ROM)快照,以及用于验证已释放内存是否确实显示为空闲(而非被缓存)的drop_caches验证循环辅助工具。

目的

测量当前 Jetson 的内存消耗情况,捕获变更前后的基准数据,并验证用户批准的更改是否确实回收了 DRAM。使用实时设备数据,而非基于容器大小、模型大小或通用进程内存的估算值。

重要提示:停止 vLLM / sglang 后内存似乎被占用(JetPack 7.2 以下版本 / L4T r39.0 以下版本)

这是 JetPack 7.2 之前或 L4T r39.0 之前 Jetson 版本中最常见的内存混淆问题。

在停止 vLLM、sglang 或 Ollama 服务器(或任何 CUDA 工作负载)后,即使进程已结束,free -h或tegrastats显示的可用内存也可能无法恢复。nvidia-smi还可能显示误导性的较低 GPU 可用内存值。

根本原因:在 CUDA 上下文退出后,Thor RM(资源管理器)会将其释放的 sysmem 页保留在自身的池中。在 Jetson 等采用统一内存架构(UMA)的设备上,cudaMemGetInfo会读取 RM 池的状态,并报告的可用内存远少于新进程实际可用的内存。

解决方法(适用于 7.2 之前的 JetPack 或 r39.0 之前的 L4T):

sudo sync && sudo sysctl -w vm.drop_caches=3

请在主机上运行此命令,而非在容器内。关键操作是sudo sysctl -w vm.drop_caches=3;务必在该命令前立即执行 sudo sync,以确保在释放可回收的页面/dentry/inode 缓存之前,先将脏数据刷新掉。 执行后,free -h和tegrastats将显示真实的可用内存。

对于受影响的版本,当用户反映:

  • “停止 vLLM/sglang 后内存未释放”
  • “为什么我的容器退出后,tegrastats 仍显示高内存占用?”
  • “明明什么都没在运行,却出现了 OOM 错误”
  • “昨天内存还好好的,现在却满了”

在 JetPack 7.2 以下或 L4T r39.0 以下的版本中,当 CUDA 工作负载退出后内存似乎卡住时,drop_caches是可靠的解决方法;在新版本中,仅当观察到相同症状且用户同意时才使用该方法。

何时使用

  • “这台 Jetson 当前占用了多少内存?是什么在占用这些内存?”
  • “我已禁用GUI/停止vLLM/退出容器——内存真的释放了吗?”
  • “为什么在停止工作负载后,free -h命令仍显示可用内存不足?”
  • 在应用jetson-headless-mode或其他与内存相关的更改之前,将其作为基准进行测量;并在更改后再次测量,以计算实际的内存变化量。

先决条件

  • 请在 Jetson 主机上运行,或在沙箱/容器中运行(该环境需能访问主机的/proc、/etc/nv_tegra_release、tegrastats 及进程数据)。
  • NvMap debugfs 的读取操作可能需要 root 权限。若无法获取,请报告 GPU 内存归属受限,而非进行猜测。
  • drop_caches.sh需要 root 权限或无需密码的sudo -n;仅在用户明确授权清除缓存后方可运行。

可用脚本

脚本 用途 参数
scripts/audit.sh 从jetson-diagnostic/scripts/snapshot.sh生成用于内存审计工作流的 JSON 快照。 无参数。
scripts/drop_caches.sh 清空可回收的页面/dentry/inode 缓存,并打印内存清空前后的变化量。 --mode 1|2|3,--quiet。

如果您的代理运行时支持run_script,请使用它来运行scripts/audit.sh或scripts/drop_caches.sh,并汇总返回的输出。否则,请从存储库根目录使用bash运行这些脚本。

操作说明

对于“当前占用了多少内存?”这类问题,请运行scripts/audit.sh并仅报告 JSON 快照中的数值。

报告指南

请勿仅打印或提及辅助工具的路径。应先调用该辅助工具,再汇总返回的数据。

  • 对于“当前内存使用量是多少”的提示,请运行scripts/audit.sh,并引用mem_total_gb、memory_kb.available,以及排在首位的procrank_top进程或nvmap.top_clients资源消耗者。
  • 对于 GUI/桌面内存提示,请运行scripts/audit.sh,并报告default_systemd_target以及candidate_services中任何显示管理器(gdm3、gdm、lightdm、sddm 或display-manager)。请勿禁用任何服务;请将问题移交至jetson-headless-mode以制定解决方案。
  • 对于明确授权在工作负载停止后清除缓存的提示,请运行scripts/drop_caches.sh(默认等同于sudo sync && sudo sysctl -w vm.drop_caches=3),并报告其执行前后的空闲空间、可用空间和缓存空间的变化量。 如果无法使用 root 权限,请说明必须在主机上使用 sudo 命令运行该操作。

如果您的代理运行时无法执行相对于此技能目录的辅助脚本,请使用 AgentSkills{baseDir}占位符解析脚本路径:

{baseDir}/scripts/audit.sh
{baseDir}/scripts/drop_caches.sh

请勿将 jetson-memory-audit 作为工具名称,除非运行时明确将技能注册为可调用的工具;代理技能通常由指令和文件组成,而非直接的工具函数。

针对代理的沙箱注意事项:看到此技能文件并不保证能访问 Jetson 主机内存数据。如果/proc/device-tree/model、/etc/nv_tegra_release、tegrastats、/sys/kernel/debug/nvmap 或主机进程数据缺失,则应提示沙箱缺乏对 Jetson 主机的可见性,并要求用户在 Jetson 主机上运行,或使用可访问主机的沙箱配置文件重新启动。 请勿捏造内存总量、可用内存、PSS、NvMap 或回收增量。

对于“此更改释放了多少内存?”之类的问题,请使用更改前后的增量值。请勿根据容器大小、镜像大小、RSS 或单个更改后的快照来估算释放的内存。

  1. 变更前,运行scripts/audit.sh并保存 JSON 基线。
  2. 执行用户批准的变更(停止容器、切换模式、应用调优建议等)。
  3. 在 JetPack 7.2 以下版本、L4T r39.0 以下版本中,或者当在新版本中观察到相同的内存卡死症状时,请在主机上(而非容器内部)清空可回收页面缓存,以便释放的页面显示为“空闲”而非“缓存”:
    sudo sync && sudo sysctl -w vm.drop_caches=3
    
    
  4. 重新运行scripts/audit.sh,并比较更改前后的memory_kb.available 值——该差值即为实际回收的内存量。

如果用户已经进行了更改且没有基线数据,则说明仅凭当前快照无法确定确切的释放量。请立即捕获新的基线数据,以便后续更改时能够进行测量。

以实时审计数据作为权威来源。 内存总量、可用内存、NvMap 总量、PSS 值、显示管理器状态以及节省量差异必须来自实际设备上的scripts/audit.sh、free -h 或tegrastats。如果某个数值未出现在这些输出中,请勿进行猜测。

audit.sh的输出规范

{
  "sku": "orin-nano",
  "variant": "orin-nano-8gb",
  "mem_total_gb": 8,
  "l4t_version": "36.4.0",
  "product_model": "nvidia jetson orin nano developer kit",
  "memory_kb": { "total": 8123456, "available": 4123456, "free": 1023456, "cached": 1234567, "swap_total": 0, "swap_free": 0 },
  "default_systemd_target": "graphical.target",
  "candidate_services": { "gdm3": { "active": "active", "enabled": "enabled" } },
  "tegrastats_sample": "内存 4011/8138MB (lfb 8x4MB) ...",
  "nvmap": { "readable": false, "total_kb": 0, "top_clients": [] },
  "procrank_top": [ { "pid": 4321, "pss_kb": 4000000, "cmd": "vllm" } ]
}

限制

  • 要精确计算内存释放增量,需要一个“之前”快照、用户批准的更改、在适当情况下进行缓存清空,以及一个“之后”快照。
  • NvMap 的归因依赖于对主机可见的 debugfs 访问;如果不可用,则报告有限的 GPU 内存归因,而不是进行猜测。
  • 沙箱/容器运行时可能无法访问主机的/proc、tegrastats、systemd 或 NvMap 数据,除非运行时将其公开。

错误处理

  • 如果scripts/audit.sh无法访问主机 Jetson 数据,请报告可见性缺失,并要求在 Jetson 主机上或可在主机可见的沙箱中重新运行。
  • 如果scripts/drop_caches.sh缺乏 root 权限或无需密码的sudo -n 权限,请报告缓存清除必须在获得 sudo 批准的主机上运行。
  • 如果不存在之前的快照,则说明仅凭当前状态无法恢复确切的回收量,并为下一次变更捕获新的基线。

安全性

只读模式。drop_caches操作是非破坏性的(内核仅会释放那些在资源压力下本可回收的页面;会先运行sync 命令以保留未写入磁盘的数据)。

移交至

  • jetson-headless-mode—— 对于仍通过graphical.target 启动的系统,这是用户空间层面最大的优化收益。
  • jetson-inference-mem-tune—— 当模型服务器是 NvMap / PSS 的最大消耗者时。
  • 如果运行时更改无法达到目标,应报告进一步回收超出了该技能的范围,而不是建议进行不安全的启动时编辑。
在 GitHub 上查看
---
name: jetson-memory-audit
description: Measure Jetson DRAM and NvMap usage, capture before/after baselines, and verify memory reclamation with live audit data.
license: Apache-2.0
---

# Jetson Memory Audit

Read-only memory-focused snapshot for a Jetson, plus the `drop_caches` verify-loop helper that confirms freed memory actually shows up as free instead of cached.

## Purpose

Measure current Jetson memory consumers, capture before/after baselines, and verify whether user-approved changes actually reclaimed DRAM. Use live device data rather than estimates from container size, model size, or generic process memory.

## CRITICAL: Memory appears stuck after stopping vLLM / sglang (JetPack below 7.2 / L4T below r39.0)

**This is the most common memory confusion on Jetson releases before JetPack 7.2 or before L4T r39.0.**

After you stop a vLLM, sglang, or Ollama server (or any CUDA workload), the memory shown as free by `free -h` or `tegrastats` may not recover — even though the process is gone. `nvidia-smi` may also show misleadingly low free GPU memory.

**Root cause:** The Thor RM (resource manager) holds freed sysmem pages in its own pool after a CUDA context exits. On Unified Memory Architecture (UMA) devices like Jetson, `cudaMemGetInfo` reads RM pool state and reports far less free memory than is actually available to a new process.

**Workaround (for JetPack below 7.2 or L4T below r39.0):**
```bash
sudo sync && sudo sysctl -w vm.drop_caches=3
```
Run this on the **host**, not inside a container. The important operation is `sudo sysctl -w vm.drop_caches=3`; keep `sudo sync` immediately before it so dirty data is flushed before reclaimable page/dentry/inode caches are dropped. After running it, `free -h` and `tegrastats` will reflect the true available memory.

For affected releases, recommend this command when a user says:
- "Memory didn't free after I stopped vLLM/sglang"
- "Why does tegrastats still show high usage after my container exited?"
- "OOM even though nothing is running"
- "Memory was fine yesterday but now it's full"

On JetPack below 7.2 or L4T below r39.0, `drop_caches` is the reliable workaround when memory appears stuck after a CUDA workload exits; on newer releases, use it only if the same symptom is observed and the user approves.

## When to use

- "How much memory is in use on this Jetson? What's holding it?"
- "I disabled the GUI / stopped vLLM / quit my container — did memory actually free?"
- "Why does `free -h` still show low free memory after I stopped my workload?"
- As the **baseline** before applying `jetson-headless-mode` or other memory-related changes, and again **after** to compute the actual delta.

## Prerequisites

- Run on the Jetson host, or in a sandbox/container with host-visible `/proc`, `/etc/nv_tegra_release`, `tegrastats`, and process data.
- NvMap debugfs reads may require root. If unavailable, report that GPU memory attribution is limited rather than guessing.
- `drop_caches.sh` requires root or passwordless `sudo -n`; run it only after the user explicitly authorizes cache dropping.

## Available Scripts

| Script | Purpose | Arguments |
|--------|---------|-----------|
| `scripts/audit.sh` | Emits a JSON snapshot from `jetson-diagnostic/scripts/snapshot.sh` for memory audit workflows. | No arguments. |
| `scripts/drop_caches.sh` | Flushes reclaimable page/dentry/inode caches and prints before/after memory deltas. | `--mode 1\|2\|3`, `--quiet`. |

If your agent runtime supports `run_script`, use it to run `scripts/audit.sh` or `scripts/drop_caches.sh` and summarize the returned output. Otherwise run the scripts with `bash` from the repository root.

## Instructions

For "how much memory is in use right now?" questions, run `scripts/audit.sh` and report only values from the JSON snapshot.

## Reporting guidance

Do not only print or mention the path to a helper. Invoke the helper and then summarize the returned data.

- For "how much memory is in use" prompts, run `scripts/audit.sh` and quote `mem_total_gb`, `memory_kb.available`, and the leading `procrank_top` process or `nvmap.top_clients` consumer.
- For GUI/desktop memory prompts, run `scripts/audit.sh` and report `default_systemd_target` plus any display manager in `candidate_services` (`gdm3`, `gdm`, `lightdm`, `sddm`, or `display-manager`). Do not disable anything; hand off to `jetson-headless-mode` for a plan.
- For prompts that explicitly authorize cache dropping after a stopped workload, run `scripts/drop_caches.sh` (equivalent to `sudo sync && sudo sysctl -w vm.drop_caches=3` by default) and report its before/after free, available, and cached deltas. If root is unavailable, explain that it must be run on the host with sudo.

If your agent runtime does not execute helper scripts relative to this skill directory, resolve script paths with the AgentSkills `{baseDir}` placeholder:

```bash
{baseDir}/scripts/audit.sh
{baseDir}/scripts/drop_caches.sh
```

Do not call `jetson-memory-audit` as a tool name unless the runtime explicitly registers skills as callable tools; Agent Skills are normally instructions plus files, not direct tool functions.

Sandbox note for agents: seeing this skill file does not guarantee access to Jetson host memory data. If `/proc/device-tree/model`, `/etc/nv_tegra_release`, `tegrastats`, `/sys/kernel/debug/nvmap`, or host process data are missing inside a NemoClaw/OpenClaw sandbox, say the sandbox lacks Jetson host visibility and ask the user to run on the Jetson host or relaunch with a host-visible sandbox profile. Do not fabricate memory totals, available memory, PSS, NvMap, or reclamation deltas.

For "how much memory did this change free?" questions, use a before/after delta. Do not estimate freed memory from container size, image size, RSS, or a single post-change snapshot.

1. Before the change, run `scripts/audit.sh` and save the JSON baseline.
2. Make the user-approved change (stop the container, switch mode, apply a tuning recommendation, etc.).
3. On JetPack below 7.2 / L4T below r39.0, or when the same stuck-memory symptom is observed on a newer release, flush reclaimable page cache on the **host** (not inside a container) so freed pages show up as free instead of cached:
   ```bash
   sudo sync && sudo sysctl -w vm.drop_caches=3
   ```
4. Re-run `scripts/audit.sh` and compare `memory_kb.available` before vs after — that delta is the real reclamation.

If the user already made the change and no baseline exists, say that the exact freed amount cannot be recovered from the current snapshot alone. Capture a new baseline now so the next change can be measured.

Use live audit data as the source of truth. Memory totals, available memory, NvMap totals, PSS values, display-manager state, and savings deltas must come from `scripts/audit.sh`, `free -h`, or `tegrastats` on the actual device. If a number is not present in those outputs, do not guess it.

## Output contract for `audit.sh`

```json
{
  "sku": "orin-nano",
  "variant": "orin-nano-8gb",
  "mem_total_gb": 8,
  "l4t_version": "36.4.0",
  "product_model": "nvidia jetson orin nano developer kit",
  "memory_kb": { "total": 8123456, "available": 4123456, "free": 1023456, "cached": 1234567, "swap_total": 0, "swap_free": 0 },
  "default_systemd_target": "graphical.target",
  "candidate_services": { "gdm3": { "active": "active", "enabled": "enabled" } },
  "tegrastats_sample": "RAM 4011/8138MB (lfb 8x4MB) ...",
  "nvmap": { "readable": false, "total_kb": 0, "top_clients": [] },
  "procrank_top": [ { "pid": 4321, "pss_kb": 4000000, "cmd": "vllm" } ]
}
```

## Limitations

- Exact freed-memory deltas require a before snapshot, the user-approved change, cache flush when appropriate, and an after snapshot.
- NvMap attribution depends on host-visible debugfs access; if it is unavailable, report limited GPU memory attribution instead of guessing.
- Sandbox/container runs may not see host `/proc`, `tegrastats`, systemd, or NvMap data unless the runtime exposes them.

## Error handling

- If `scripts/audit.sh` cannot access host Jetson data, report the missing visibility and ask to rerun on the Jetson host or in a host-visible sandbox.
- If `scripts/drop_caches.sh` lacks root or passwordless `sudo -n`, report that cache dropping must be run on the host with sudo approval.
- If no before snapshot exists, say the exact reclaimed amount cannot be recovered from the current state alone and capture a new baseline for the next change.

## Safety

Read-only. `drop_caches` is non-destructive (kernel only releases pages it could reclaim under pressure anyway; `sync` runs first to preserve dirty data).

## Hand off to

- `jetson-headless-mode` — biggest single user-space win on systems still booting `graphical.target`.
- `jetson-inference-mem-tune` — when a model server is the top NvMap / PSS consumer.
- If runtime changes cannot hit the target, report that further reclamation is outside this skill's scope rather than suggesting unsafe boot-time edits.

安装 jetson-memory-audit

下载技能文件并将其解压到 .claude/skills/ 目录中。

下载ZIP

克隆仓库并复制技能文件到您的项目中。

git clone https://github.com/NVIDIA/skills/tree/main/skills/jetson-memory-audit # Copy SKILL.md to your .claude/skills/ directory

复制 复制
快速设置: 将技能文件夹复制到 .claude/skills/ Claude 会自动检测并使用该技能
仓库 NVIDIA/skills

相关技能

gmgn-portfolio
更新时间 2026-07-01
device-integrity
更新时间 2026-06-29
zeroize-audit
更新时间 2026-07-01
flutter-use-http-package
更新时间 2026-06-30
OR