jetson-memory-audit
NVIDIA/skills
測量 Jetson 的 DRAM 及 NvMap 使用狀況,擷取基準值(實施前後),並透過即時稽核資料驗證記憶體回收成效。
...展開全部Jetson 記憶體稽核
針對 Jetson 的唯讀記憶體(ROM)快照,以及名為 `drop_caches`的驗證迴圈輔助程式,用以確認已釋放的記憶體確實顯示為可用狀態,而非仍被快取佔用。
目的
測量當前 Jetson 的記憶體消耗來源,擷取變更前後的基準值,並驗證使用者核准的變更是否確實釋放了 DRAM。採用即時裝置資料,而非根據容器大小、模型大小或通用程序記憶體所做的估算。
重要:停止 vLLM / sglang 後記憶體似乎卡住(JetPack 低於 7.2 / L4T 低於 r39.0)
這是 JetPack 7.2 之前或 L4T r39.0 之前 Jetson 版本中最常見的記憶體混淆問題。
當您停止 vLLM、sglang 或 Ollama 伺服器(或任何 CUDA 工作負載)後,即使進程已結束,透過`free -h`或`tegrastats`顯示為「可用」的記憶體可能仍無法釋放。此外,`nvidia-smi`可能會顯示誤導性的低 GPU 可用記憶體值。
根本原因:當 CUDA 執行上下文結束後,Thor RM(資源管理器)會將已釋放的 sysmem 頁面保留在其自身的池中。在 Jetson 等採用統一記憶體架構(UMA)的裝置上,cudaMemGetInfo會讀取 RM 池的狀態,並回報的可用記憶體遠低於新進程實際可用的記憶體量。
解決方法(適用於 JetPack 7.2 以下版本或 L4T r39.0 以下版本):
sudo sync && sudo sysctl -w vm.drop_caches=3
請在主機上執行此指令,而非在容器內。關鍵操作是`sudo sysctl -w vm.drop_caches=3`;請務必在該指令前立即執行`sudo sync`,以確保在釋放可回收的頁面/dentry/inode 快取之前,先將未寫入的資料沖洗掉。 執行後,free -h和tegrastats將顯示真實的可用記憶體。
針對受影響的發行版,當使用者反映以下情況時,建議執行此指令:
- 「停止 vLLM/sglang 後,記憶體並未釋放」
- 「為什麼我的容器退出後,tegrastats 仍顯示高使用率?」
- 「明明沒有任何程式在執行,卻出現 OOM 錯誤」
- 「昨天記憶體還好好的,現在卻滿了」
在 JetPack 7.2 以下版本或 L4T r39.0 以下版本中,若 CUDA 工作負載退出後記憶體似乎卡住,執行`drop_caches`指令是可靠的解決方法;在新版發行版本中,僅在觀察到相同症狀且經使用者同意的情況下才使用此方法。
何時使用
- 「這台 Jetson 目前佔用了多少記憶體?是什麼在佔用這些記憶體?」
- 「我已停用 GUI / 停止 vLLM / 退出容器 —— 記憶體真的釋放了嗎?」
- 「為何在停止工作負載後,
free -h仍顯示可用記憶體不足?」 - 作為套用
jetson-headless-mode或其他與記憶體相關變更前的基準值,並在變更後再次執行以計算實際的變化量。
先決條件
- 請在 Jetson 主機上執行,或在具備主機可見
/proc、/etc/nv_tegra_release、tegrastats及程序資料的沙箱/容器中執行。 - NvMap debugfs 讀取可能需要 root 權限。若無法取得,請回報 GPU 記憶體歸屬受限,而非進行推測。
drop_caches.sh需具備 root 權限或無密碼的sudo -n權限;僅在使用者明確授權清除快取後方可執行。
可用腳本
| 腳本 | 用途 | 參數 |
|---|---|---|
scripts/audit.sh |
從jetson-diagnostic/scripts/snapshot.sh產生 JSON 快照,供記憶體稽核工作流程使用。 |
無參數。 |
scripts/drop_caches.sh |
清除可回收的頁面/dentry/inode 快取,並輸出記憶體變動的「前」與「後」數值。 | --mode 1|2|3,--quiet。 |
若您的代理程式執行環境支援run_script,請使用該功能執行scripts/audit.sh或scripts/drop_caches.sh,並彙整回傳的輸出結果。否則,請從儲存庫根目錄使用bash執行這些腳本。
操作說明
若要查詢「目前佔用多少記憶體?」,請執行scripts/audit.sh,並僅回報 JSON 快照中的數值。
報告指引
請勿僅列印或提及輔助程式的路徑。應實際呼叫該輔助程式,並彙整其回傳的資料。
- 針對「目前佔用多少記憶體」的提示,請執行
scripts/audit.sh,並引用mem_total_gb、memory_kb.available,以及排在首位的procrank_top進程或nvmap.top_clients資源消耗者。 - 針對 GUI/桌面記憶體的提示,請執行
scripts/audit.sh,並回報default_systemd_target以及candidate_services中的任何顯示管理員(gdm3、gdm、lightdm、sddm或display-manager)。請勿停用任何項目;請交由jetson-headless-mode擬定方案。 - 若提示明確授權在工作負載停止後清除快取,請執行
scripts/drop_caches.sh(預設等同於sudo sync && sudo sysctl -w vm.drop_caches=3),並回報執行前後的可用空間、空閒空間及快取量的變化值。 若無法取得 root 權限,請說明必須在主機上使用 sudo 執行該指令。
若您的代理程式執行環境無法執行此技能目錄下的輔助腳本,請使用 AgentSkills{baseDir}佔位符來解析腳本路徑:
{baseDir}/scripts/audit.sh
{baseDir}/scripts/drop_caches.sh
請勿將 jetson-memory-audit 作為工具名稱,除非執行環境已明確將技能註冊為可呼叫的工具;代理技能通常由指令與檔案組成,而非直接的工具函式。
針對代理程式的沙箱注意事項:檢視此技能檔案並不保證能存取 Jetson 主機的記憶體資料。若在 NemoClaw/OpenClaw 沙箱內缺少/proc/device-tree/model、/etc/nv_tegra_release、tegrastats、/sys/kernel/debug/nvmap 或主機程序資料,請說明該沙箱缺乏 Jetson 主機可見性,並請使用者在 Jetson 主機上執行,或使用具備主機可見性的沙箱設定檔重新啟動。 請勿捏造總記憶體量、可用記憶體、PSS、NvMap 或回收增量。
針對「此變更釋放了多少記憶體?」的提問,請使用變更前後的差異值。請勿根據容器大小、映像檔大小、RSS 或單一變更後快照來估算釋放的記憶體。
- 變更前,請執行
scripts/audit.sh並儲存 JSON 基準資料。 - 執行經使用者核准的變更(停止容器、切換模式、套用調校建議等)。
- 若使用 JetPack 7.2 以下版本/L4T r39.0 以下版本,或在新版發行版上觀察到相同的記憶體卡住症狀時,請在主機上(而非容器內部)清除可回收的頁面快取,以便釋放的頁面顯示為「空閒」而非「快取」狀態:
sudo sync && sudo sysctl -w vm.drop_caches=3 - 重新執行
scripts/audit.sh,並比較執行前後的memory_kb.available 數值——該差異即為實際回收的記憶體量。
若使用者已進行變更且無基準值,請說明僅憑當前快照無法精確推算出釋放的記憶體量。請立即擷取新的基準值,以便日後能測量下次變更的效果。
請以即時稽核資料作為最終依據。 記憶體總量、可用記憶體、NvMap 總量、PSS 值、顯示管理員狀態以及節省量差異,必須來自實際裝置上的scripts/audit.sh、free -h 或tegrastats。若某個數值未出現在上述輸出中,請勿憑空推測。
audit.sh的輸出規範
{
"sku": "orin-nano",
"variant": "orin-nano-8gb",
"mem_total_gb": 8,
"l4t_version": "36.4.0",
"product_model": "nvidia jetson orin nano developer kit",
"memory_kb": { "total": 8123456, "available": 4123456, "free": 1023456, "cached": 1234567, "swap_total": 0, "swap_free": 0 },
"default_systemd_target": "graphical.target",
"candidate_services": { "gdm3": { "active": "active", "enabled": "enabled" } },
"tegrastats_sample": "RAM 4011/8138MB (lfb 8x4MB) ...",
"nvmap": { "readable": false, "total_kb": 0, "top_clients": [] },
"procrank_top": [ { "pid": 4321, "pss_kb": 4000000, "cmd": "vllm" } ]
}
限制事項
- 要精確計算釋放記憶體的增減值,需具備以下條件:事前快照、經使用者核准的變更、適時執行快取清除,以及事後快照。
- NvMap 的歸因取決於能否存取主機可見的 debugfs;若無法存取,則應回報有限的 GPU 記憶體歸因,而非進行推測。
- 沙箱/容器執行環境可能無法存取主機的
/proc、tegrastats、systemd 或 NvMap 資料,除非執行環境將其公開。
錯誤處理
- 若
scripts/audit.sh無法存取主機 Jetson 資料,請回報可見性缺失,並要求在 Jetson 主機上或可存取主機資料的沙箱中重新執行。 - 若
scripts/drop_caches.sh缺乏 root 權限或無密碼的sudo -n權限,請回報必須在獲得 sudo 授權的主機上執行快取清除作業。 - 若不存在先前快照,則應說明無法僅憑當前狀態恢復確切的回收量,並為下一次變更建立新的基準。
安全性
唯讀模式。drop_caches為非破壞性操作(核心僅會釋放那些在壓力下本可回收的頁面;會先執行sync以保留未寫入的資料)。
交由
jetson-headless-mode— 對於仍以graphical.target啟動的系統而言,這是使用者空間層面最大的單一優化。jetson-inference-mem-tune—— 當模型伺服器是 NvMap / PSS 的最大消耗者時。- 若執行時變更無法達到目標,應回報「進一步的資源回收超出此技能的範圍」,而非建議進行不安全的開機時編輯。
---
name: jetson-memory-audit
description: Measure Jetson DRAM and NvMap usage, capture before/after baselines, and verify memory reclamation with live audit data.
license: Apache-2.0
---
# Jetson Memory Audit
Read-only memory-focused snapshot for a Jetson, plus the `drop_caches` verify-loop helper that confirms freed memory actually shows up as free instead of cached.
## Purpose
Measure current Jetson memory consumers, capture before/after baselines, and verify whether user-approved changes actually reclaimed DRAM. Use live device data rather than estimates from container size, model size, or generic process memory.
## CRITICAL: Memory appears stuck after stopping vLLM / sglang (JetPack below 7.2 / L4T below r39.0)
**This is the most common memory confusion on Jetson releases before JetPack 7.2 or before L4T r39.0.**
After you stop a vLLM, sglang, or Ollama server (or any CUDA workload), the memory shown as free by `free -h` or `tegrastats` may not recover — even though the process is gone. `nvidia-smi` may also show misleadingly low free GPU memory.
**Root cause:** The Thor RM (resource manager) holds freed sysmem pages in its own pool after a CUDA context exits. On Unified Memory Architecture (UMA) devices like Jetson, `cudaMemGetInfo` reads RM pool state and reports far less free memory than is actually available to a new process.
**Workaround (for JetPack below 7.2 or L4T below r39.0):**
```bash
sudo sync && sudo sysctl -w vm.drop_caches=3
```
Run this on the **host**, not inside a container. The important operation is `sudo sysctl -w vm.drop_caches=3`; keep `sudo sync` immediately before it so dirty data is flushed before reclaimable page/dentry/inode caches are dropped. After running it, `free -h` and `tegrastats` will reflect the true available memory.
For affected releases, recommend this command when a user says:
- "Memory didn't free after I stopped vLLM/sglang"
- "Why does tegrastats still show high usage after my container exited?"
- "OOM even though nothing is running"
- "Memory was fine yesterday but now it's full"
On JetPack below 7.2 or L4T below r39.0, `drop_caches` is the reliable workaround when memory appears stuck after a CUDA workload exits; on newer releases, use it only if the same symptom is observed and the user approves.
## When to use
- "How much memory is in use on this Jetson? What's holding it?"
- "I disabled the GUI / stopped vLLM / quit my container — did memory actually free?"
- "Why does `free -h` still show low free memory after I stopped my workload?"
- As the **baseline** before applying `jetson-headless-mode` or other memory-related changes, and again **after** to compute the actual delta.
## Prerequisites
- Run on the Jetson host, or in a sandbox/container with host-visible `/proc`, `/etc/nv_tegra_release`, `tegrastats`, and process data.
- NvMap debugfs reads may require root. If unavailable, report that GPU memory attribution is limited rather than guessing.
- `drop_caches.sh` requires root or passwordless `sudo -n`; run it only after the user explicitly authorizes cache dropping.
## Available Scripts
| Script | Purpose | Arguments |
|--------|---------|-----------|
| `scripts/audit.sh` | Emits a JSON snapshot from `jetson-diagnostic/scripts/snapshot.sh` for memory audit workflows. | No arguments. |
| `scripts/drop_caches.sh` | Flushes reclaimable page/dentry/inode caches and prints before/after memory deltas. | `--mode 1\|2\|3`, `--quiet`. |
If your agent runtime supports `run_script`, use it to run `scripts/audit.sh` or `scripts/drop_caches.sh` and summarize the returned output. Otherwise run the scripts with `bash` from the repository root.
## Instructions
For "how much memory is in use right now?" questions, run `scripts/audit.sh` and report only values from the JSON snapshot.
## Reporting guidance
Do not only print or mention the path to a helper. Invoke the helper and then summarize the returned data.
- For "how much memory is in use" prompts, run `scripts/audit.sh` and quote `mem_total_gb`, `memory_kb.available`, and the leading `procrank_top` process or `nvmap.top_clients` consumer.
- For GUI/desktop memory prompts, run `scripts/audit.sh` and report `default_systemd_target` plus any display manager in `candidate_services` (`gdm3`, `gdm`, `lightdm`, `sddm`, or `display-manager`). Do not disable anything; hand off to `jetson-headless-mode` for a plan.
- For prompts that explicitly authorize cache dropping after a stopped workload, run `scripts/drop_caches.sh` (equivalent to `sudo sync && sudo sysctl -w vm.drop_caches=3` by default) and report its before/after free, available, and cached deltas. If root is unavailable, explain that it must be run on the host with sudo.
If your agent runtime does not execute helper scripts relative to this skill directory, resolve script paths with the AgentSkills `{baseDir}` placeholder:
```bash
{baseDir}/scripts/audit.sh
{baseDir}/scripts/drop_caches.sh
```
Do not call `jetson-memory-audit` as a tool name unless the runtime explicitly registers skills as callable tools; Agent Skills are normally instructions plus files, not direct tool functions.
Sandbox note for agents: seeing this skill file does not guarantee access to Jetson host memory data. If `/proc/device-tree/model`, `/etc/nv_tegra_release`, `tegrastats`, `/sys/kernel/debug/nvmap`, or host process data are missing inside a NemoClaw/OpenClaw sandbox, say the sandbox lacks Jetson host visibility and ask the user to run on the Jetson host or relaunch with a host-visible sandbox profile. Do not fabricate memory totals, available memory, PSS, NvMap, or reclamation deltas.
For "how much memory did this change free?" questions, use a before/after delta. Do not estimate freed memory from container size, image size, RSS, or a single post-change snapshot.
1. Before the change, run `scripts/audit.sh` and save the JSON baseline.
2. Make the user-approved change (stop the container, switch mode, apply a tuning recommendation, etc.).
3. On JetPack below 7.2 / L4T below r39.0, or when the same stuck-memory symptom is observed on a newer release, flush reclaimable page cache on the **host** (not inside a container) so freed pages show up as free instead of cached:
```bash
sudo sync && sudo sysctl -w vm.drop_caches=3
```
4. Re-run `scripts/audit.sh` and compare `memory_kb.available` before vs after — that delta is the real reclamation.
If the user already made the change and no baseline exists, say that the exact freed amount cannot be recovered from the current snapshot alone. Capture a new baseline now so the next change can be measured.
Use live audit data as the source of truth. Memory totals, available memory, NvMap totals, PSS values, display-manager state, and savings deltas must come from `scripts/audit.sh`, `free -h`, or `tegrastats` on the actual device. If a number is not present in those outputs, do not guess it.
## Output contract for `audit.sh`
```json
{
"sku": "orin-nano",
"variant": "orin-nano-8gb",
"mem_total_gb": 8,
"l4t_version": "36.4.0",
"product_model": "nvidia jetson orin nano developer kit",
"memory_kb": { "total": 8123456, "available": 4123456, "free": 1023456, "cached": 1234567, "swap_total": 0, "swap_free": 0 },
"default_systemd_target": "graphical.target",
"candidate_services": { "gdm3": { "active": "active", "enabled": "enabled" } },
"tegrastats_sample": "RAM 4011/8138MB (lfb 8x4MB) ...",
"nvmap": { "readable": false, "total_kb": 0, "top_clients": [] },
"procrank_top": [ { "pid": 4321, "pss_kb": 4000000, "cmd": "vllm" } ]
}
```
## Limitations
- Exact freed-memory deltas require a before snapshot, the user-approved change, cache flush when appropriate, and an after snapshot.
- NvMap attribution depends on host-visible debugfs access; if it is unavailable, report limited GPU memory attribution instead of guessing.
- Sandbox/container runs may not see host `/proc`, `tegrastats`, systemd, or NvMap data unless the runtime exposes them.
## Error handling
- If `scripts/audit.sh` cannot access host Jetson data, report the missing visibility and ask to rerun on the Jetson host or in a host-visible sandbox.
- If `scripts/drop_caches.sh` lacks root or passwordless `sudo -n`, report that cache dropping must be run on the host with sudo approval.
- If no before snapshot exists, say the exact reclaimed amount cannot be recovered from the current state alone and capture a new baseline for the next change.
## Safety
Read-only. `drop_caches` is non-destructive (kernel only releases pages it could reclaim under pressure anyway; `sync` runs first to preserve dirty data).
## Hand off to
- `jetson-headless-mode` — biggest single user-space win on systems still booting `graphical.target`.
- `jetson-inference-mem-tune` — when a model server is the top NvMap / PSS consumer.
- If runtime changes cannot hit the target, report that further reclamation is outside this skill's scope rather than suggesting unsafe boot-time edits.
所有檔案
8 個檔案安裝 jetson-memory-audit
請下載並將技能檔案解壓縮至您的 .claude/skills/ 目錄中。
下載 ZIP複製儲存庫並將技能檔案複製到您的專案中。
git clone https://github.com/NVIDIA/skills/tree/main/skills/jetson-memory-audit # Copy SKILL.md to your .claude/skills/ directory
複製





首頁
