jetson-memory-audit
NVIDIA/skills
JetsonのDRAMおよびNvMapの使用状況を測定し、変更前後のベースラインを収集した上で、実際の監査データを用いてメモリの回収状況を検証します。
...すべて拡張しますJetson メモリ監査
Jetson向けの読み取り専用メモリに焦点を当てたスナップショットに加え、解放されたメモリがキャッシュされずに実際に空きメモリとして認識されることを確認する「drop_caches」検証ループヘルパー。
目的
現在の Jetson メモリ消費元を測定し、変更前後のベースラインをキャプチャして、ユーザーが承認した変更によって実際に DRAM が解放されたかどうかを検証します。コンテナサイズ、モデルサイズ、または一般的なプロセスメモリからの推定値ではなく、ライブデバイスデータを使用します。
重要:vLLM / sglang の停止後にメモリが固定された状態になる(JetPack 7.2 未満 / L4T r39.0 未満)
これは、JetPack 7.2 以前または L4T r39.0 以前の Jetson リリースにおいて、最もよく見られるメモリに関する誤解です。
vLLM、sglang、またはOllamaサーバー(あるいは任意のCUDAワークロード)を停止した後、プロセスが終了しているにもかかわらず、`free -h`や`tegrastats`で空きメモリとして表示されるメモリが解放されない場合があります。また、`nvidia-smi`でも、GPUの空きメモリが実際よりも低い値として誤って表示されることがあります。
根本原因:Thor RM(リソースマネージャー)は、CUDAコンテキストが終了した後も、解放されたsysmemページを独自のプールに保持し続けます。JetsonのようなUnified Memory Architecture(UMA)デバイスでは、cudaMemGetInfoがRMプールの状態を読み取り、新しいプロセスが実際に利用可能な空きメモリよりもはるかに少ない値を報告してしまいます。
回避策(JetPack 7.2未満またはL4T r39.0未満の場合):
sudo sync && sudo sysctl -w vm.drop_caches=3
このコマンドはコンテナ内ではなく、ホスト上で実行してください。重要な操作は`sudo sysctl -w vm.drop_caches=3` です。その直前に`sudo sync` を実行し、回収可能なページ/dentry/inode キャッシュが破棄される前に、未コミットデータがフラッシュされるようにしてください。 実行後、`free -h` および`tegrastats` には、実際の利用可能メモリが反映されます。
影響を受けるリリースについては、ユーザーから次のような報告があった場合に、このコマンドの実行を推奨します:
- 「vLLM/sglangを停止してもメモリが解放されませんでした」
- 「コンテナが終了したのに、なぜ tegrastats の使用率が高いままなのか?」
- 「何も実行していないのにOOMが発生する」
- 「昨日はメモリに問題なかったのに、今は満杯だ」
JetPack 7.2 未満または L4T r39.0 未満では、CUDA ワークロードの終了後にメモリが固定されたように見える場合、drop_caches が確実な回避策となります。新しいリリースでは、同様の症状が確認され、かつユーザーが了承した場合にのみこれを使用してください。
使用すべき場合
- 「このJetsonではどのくらいのメモリが使用中か?何がメモリを占有しているのか?」
- 「GUIを無効化したり、vLLMを停止したり、コンテナを終了したりしましたが、メモリは実際に解放されましたか?」
- 「ワークロードを停止したのに、なぜ `
free -h` では空きメモリがまだ少ないと表示されるのか?」 jetson-headless-modeやその他のメモリ関連の変更を適用する前のベースラインとして、また適用後に実際の変化量を算出するために再度実行します。
前提条件
- Jetsonホスト上で、またはホストから
/proc、/etc/nv_tegra_release、tegrastats、およびプロセスデータにアクセス可能なサンドボックス/コンテナ内で実行してください。 - NvMapによるdebugfsの読み取りにはroot権限が必要な場合があります。利用できない場合は、推測するのではなく、GPUメモリの割り当てが制限されていることを報告してください。
drop_caches.shを実行するには、root 権限またはパスワード不要のsudo -nが必要です。ユーザーがキャッシュの削除を明示的に承認した後にのみ実行してください。
利用可能なスクリプト
| スクリプト | 目的 | 引数 |
|---|---|---|
scripts/audit.sh |
メモリアuditワークフロー向けに、jetson-diagnostic/scripts/snapshot.shから JSON スナップショットを出力します。 |
引数なし。 |
scripts/drop_caches.sh |
回収可能なページ/dentry/inode キャッシュをフラッシュし、処理前後のメモリの差分を出力します。 | --mode 1|2|3、--quiet。 |
エージェントランタイムがrun_script をサポートしている場合は、これを使用してscripts/audit.shまたはscripts/drop_caches.shを実行し、返された出力を要約してください。そうでない場合は、リポジトリのルートからbashを使用してスクリプトを実行してください。
手順
「現在、どのくらいのメモリが使用されているか」という質問については、scripts/audit.shを実行し、JSON スナップショットの値のみを報告してください。
レポート作成の指針
ヘルパーへのパスを単に表示したり言及したりするだけではいけません。ヘルパーを呼び出し、返されたデータを要約してください。
- 「現在、メモリはどれくらい使用されているか」というプロンプトに対しては、
scripts/audit.shを実行し、mem_total_gb、memory_kb.available、およびprocrank_topプロセスまたはnvmap.top_clientsコンシューマーのうち、最もメモリを消費しているものを引用してください。 - GUI/デスクトップのメモリに関するプロンプトについては、
scripts/audit.shを実行し、default_systemd_targetおよびcandidate_services内のディスプレイマネージャー(gdm3、gdm、lightdm、sddm、またはdisplay-manager)を報告してください。何も無効化しないでください。対応策についてはjetson-headless-modeに引き継いでください。 - ワークロードの停止後にキャッシュの削除を明示的に許可するプロンプトについては、
scripts/drop_caches.shを実行し(デフォルトではsudo sync && sudo sysctl -w vm.drop_caches=3と同等)、実行前後の free、available、および cached の差分を報告してください。 root 権限が利用できない場合は、ホスト上で sudo を使用して実行する必要があることを説明してください。
エージェントランタイムがこのスキルディレクトリを基準としてヘルパースクリプトを実行しない場合は、AgentSkills の{baseDir}プレースホルダーを使用してスクリプトのパスを解決してください:
{baseDir}/scripts/audit.sh
{baseDir}/scripts/drop_caches.sh
ランタイムがスキルを明示的に呼び出し可能なツールとして登録していない限り、 jetson-memory-audit をツール名として使用しないでください。ランタイムがスキルを明示的に呼び出し可能なツールとして登録していない限り、Agent Skillsは通常、指示とファイルの組み合わせであり、直接的なツール関数ではありません。
エージェント向けのサンドボックスに関する注意:このスキルファイルが表示されても、Jetsonホストのメモリデータへのアクセスが保証されるわけではありません。NemoClaw/OpenClawサンドボックス内で、/proc/device-tree/model、/etc/nv_tegra_release、tegrastats、/sys/kernel/debug/nvmap、またはホストプロセスデータが NemoClaw/OpenClaw サンドボックス内に存在しない場合は、そのサンドボックスには Jetson ホストへの可視性がないものとみなして、ユーザーに Jetson ホスト上で実行するか、ホスト可視のサンドボックスプロファイルを使用して再起動するよう促してください。 メモリの合計量、利用可能メモリ、PSS、NvMap、または再利用の差分をでっち上げてはいけません。
「この変更によってどのくらいのメモリが解放されたか」という質問に対しては、変更前後の差分を使用してください。コンテナサイズ、イメージサイズ、RSS、または変更後の単一のスナップショットから解放されたメモリ量を推定してはなりません。
- 変更前に、
scripts/audit.shを実行し、JSON ベースラインを保存してください。 - ユーザーが承認した変更(コンテナの停止、モードの切り替え、チューニング推奨事項の適用など)を行います。
- JetPack 7.2未満/L4T r39.0未満の場合、または新しいリリースで同様のメモリ停滞現象が確認された場合は、ホスト上(コンテナ内ではなく)で回収可能なページキャッシュをフラッシュし、解放されたページがキャッシュ済みではなく空きとして表示されるようにします:
sudo sync && sudo sysctl -w vm.drop_caches=3 scripts/audit.sh を再実行し、変更前後のmemory_kb.available を比較してください。その差分が実際の回収量となります。
ユーザーがすでに変更を行っており、ベースラインが存在しない場合は、現在のスナップショットだけでは正確に解放されたメモリ量を特定できないことを伝えます。次回の変更を測定できるよう、今すぐ新しいベースラインをキャプチャしてください。
ライブ監査データを「真実の源」として使用してください。 メモリ合計、使用可能メモリ、NvMap合計、PSS値、ディスプレイマネージャーの状態、および節約量の差分は、実際のデバイス上のscripts/audit.sh、free -h、またはtegrastatsから取得する必要があります。これらの出力に数値が含まれていない場合は、推測しないでください。
audit.shの出力例
{
"sku": "orin-nano",
"variant": "orin-nano-8gb",
"mem_total_gb": 8,
"l4t_version": "36.4.0",
"product_model": "nvidia jetson orin nano developer kit",
"memory_kb": { "total": 8123456, "available": 4123456, "free": 1023456, "cached": 1234567, "swap_total": 0, "swap_free": 0 },
"default_systemd_target": "graphical.target",
"candidate_services": { "gdm3": { "active": "active", "enabled": "enabled" } },
"tegrastats_sample": "RAM 4011/8138MB (lfb 8x4MB) ...",
"nvmap": { "readable": false, "total_kb": 0, "top_clients": [] },
"procrank_top": [ { "pid": 4321, "pss_kb": 4000000, "cmd": "vllm" } ]
}
制限事項
- 正確な解放メモリの差分を算出するには、変更前のスナップショット、ユーザー承認済みの変更、適切なタイミングでのキャッシュのフラッシュ、および変更後のスナップショットが必要です。
- NvMapによる割り当て情報の特定は、ホストから見えるdebugfsへのアクセスに依存します。利用できない場合は、推測する代わりに、限定的なGPUメモリ割り当て情報を報告します。
- サンドボックスやコンテナでの実行では、ランタイムによって公開されていない限り、ホストの
/proc、tegrastats、systemd、または NvMap データを確認できない場合があります。
エラー処理
scripts/audit.shがホストの Jetson データにアクセスできない場合は、可視性が欠落していることを報告し、Jetson ホスト上、またはホストから可視なサンドボックスで再実行するよう依頼してください。scripts/drop_caches.sh でroot 権限またはパスワード不要のsudo -nが利用できない場合は、sudo の承認を得てホスト上でキャッシュの削除を実行する必要があることを報告してください。- 以前のスナップショットが存在しない場合は、現在の状態だけでは正確な回収量を復元できないことを伝え、次の変更に向けて新しいベースラインをキャプチャしてください。
安全性
読み取り専用です。drop_caches は非破壊的です(カーネルは、負荷がかかった場合に回収可能なページのみを解放します。また、ダーティデータを保持するために、まずsync が実行されます)。
以下に引き継ぐ
jetson-headless-mode— まだgraphical.target を起動しているシステムにおいて、ユーザー空間で得られる最大のメリット。jetson-inference-mem-tune— モデルサーバーが NvMap / PSS の最大の消費元である場合。- 実行時の変更で目標を達成できない場合は、安全でない起動時の編集を提案するのではなく、それ以上のメモリ回収はこのスキルの範囲外であることを報告してください。
---
name: jetson-memory-audit
description: Measure Jetson DRAM and NvMap usage, capture before/after baselines, and verify memory reclamation with live audit data.
license: Apache-2.0
---
# Jetson Memory Audit
Read-only memory-focused snapshot for a Jetson, plus the `drop_caches` verify-loop helper that confirms freed memory actually shows up as free instead of cached.
## Purpose
Measure current Jetson memory consumers, capture before/after baselines, and verify whether user-approved changes actually reclaimed DRAM. Use live device data rather than estimates from container size, model size, or generic process memory.
## CRITICAL: Memory appears stuck after stopping vLLM / sglang (JetPack below 7.2 / L4T below r39.0)
**This is the most common memory confusion on Jetson releases before JetPack 7.2 or before L4T r39.0.**
After you stop a vLLM, sglang, or Ollama server (or any CUDA workload), the memory shown as free by `free -h` or `tegrastats` may not recover — even though the process is gone. `nvidia-smi` may also show misleadingly low free GPU memory.
**Root cause:** The Thor RM (resource manager) holds freed sysmem pages in its own pool after a CUDA context exits. On Unified Memory Architecture (UMA) devices like Jetson, `cudaMemGetInfo` reads RM pool state and reports far less free memory than is actually available to a new process.
**Workaround (for JetPack below 7.2 or L4T below r39.0):**
```bash
sudo sync && sudo sysctl -w vm.drop_caches=3
```
Run this on the **host**, not inside a container. The important operation is `sudo sysctl -w vm.drop_caches=3`; keep `sudo sync` immediately before it so dirty data is flushed before reclaimable page/dentry/inode caches are dropped. After running it, `free -h` and `tegrastats` will reflect the true available memory.
For affected releases, recommend this command when a user says:
- "Memory didn't free after I stopped vLLM/sglang"
- "Why does tegrastats still show high usage after my container exited?"
- "OOM even though nothing is running"
- "Memory was fine yesterday but now it's full"
On JetPack below 7.2 or L4T below r39.0, `drop_caches` is the reliable workaround when memory appears stuck after a CUDA workload exits; on newer releases, use it only if the same symptom is observed and the user approves.
## When to use
- "How much memory is in use on this Jetson? What's holding it?"
- "I disabled the GUI / stopped vLLM / quit my container — did memory actually free?"
- "Why does `free -h` still show low free memory after I stopped my workload?"
- As the **baseline** before applying `jetson-headless-mode` or other memory-related changes, and again **after** to compute the actual delta.
## Prerequisites
- Run on the Jetson host, or in a sandbox/container with host-visible `/proc`, `/etc/nv_tegra_release`, `tegrastats`, and process data.
- NvMap debugfs reads may require root. If unavailable, report that GPU memory attribution is limited rather than guessing.
- `drop_caches.sh` requires root or passwordless `sudo -n`; run it only after the user explicitly authorizes cache dropping.
## Available Scripts
| Script | Purpose | Arguments |
|--------|---------|-----------|
| `scripts/audit.sh` | Emits a JSON snapshot from `jetson-diagnostic/scripts/snapshot.sh` for memory audit workflows. | No arguments. |
| `scripts/drop_caches.sh` | Flushes reclaimable page/dentry/inode caches and prints before/after memory deltas. | `--mode 1\|2\|3`, `--quiet`. |
If your agent runtime supports `run_script`, use it to run `scripts/audit.sh` or `scripts/drop_caches.sh` and summarize the returned output. Otherwise run the scripts with `bash` from the repository root.
## Instructions
For "how much memory is in use right now?" questions, run `scripts/audit.sh` and report only values from the JSON snapshot.
## Reporting guidance
Do not only print or mention the path to a helper. Invoke the helper and then summarize the returned data.
- For "how much memory is in use" prompts, run `scripts/audit.sh` and quote `mem_total_gb`, `memory_kb.available`, and the leading `procrank_top` process or `nvmap.top_clients` consumer.
- For GUI/desktop memory prompts, run `scripts/audit.sh` and report `default_systemd_target` plus any display manager in `candidate_services` (`gdm3`, `gdm`, `lightdm`, `sddm`, or `display-manager`). Do not disable anything; hand off to `jetson-headless-mode` for a plan.
- For prompts that explicitly authorize cache dropping after a stopped workload, run `scripts/drop_caches.sh` (equivalent to `sudo sync && sudo sysctl -w vm.drop_caches=3` by default) and report its before/after free, available, and cached deltas. If root is unavailable, explain that it must be run on the host with sudo.
If your agent runtime does not execute helper scripts relative to this skill directory, resolve script paths with the AgentSkills `{baseDir}` placeholder:
```bash
{baseDir}/scripts/audit.sh
{baseDir}/scripts/drop_caches.sh
```
Do not call `jetson-memory-audit` as a tool name unless the runtime explicitly registers skills as callable tools; Agent Skills are normally instructions plus files, not direct tool functions.
Sandbox note for agents: seeing this skill file does not guarantee access to Jetson host memory data. If `/proc/device-tree/model`, `/etc/nv_tegra_release`, `tegrastats`, `/sys/kernel/debug/nvmap`, or host process data are missing inside a NemoClaw/OpenClaw sandbox, say the sandbox lacks Jetson host visibility and ask the user to run on the Jetson host or relaunch with a host-visible sandbox profile. Do not fabricate memory totals, available memory, PSS, NvMap, or reclamation deltas.
For "how much memory did this change free?" questions, use a before/after delta. Do not estimate freed memory from container size, image size, RSS, or a single post-change snapshot.
1. Before the change, run `scripts/audit.sh` and save the JSON baseline.
2. Make the user-approved change (stop the container, switch mode, apply a tuning recommendation, etc.).
3. On JetPack below 7.2 / L4T below r39.0, or when the same stuck-memory symptom is observed on a newer release, flush reclaimable page cache on the **host** (not inside a container) so freed pages show up as free instead of cached:
```bash
sudo sync && sudo sysctl -w vm.drop_caches=3
```
4. Re-run `scripts/audit.sh` and compare `memory_kb.available` before vs after — that delta is the real reclamation.
If the user already made the change and no baseline exists, say that the exact freed amount cannot be recovered from the current snapshot alone. Capture a new baseline now so the next change can be measured.
Use live audit data as the source of truth. Memory totals, available memory, NvMap totals, PSS values, display-manager state, and savings deltas must come from `scripts/audit.sh`, `free -h`, or `tegrastats` on the actual device. If a number is not present in those outputs, do not guess it.
## Output contract for `audit.sh`
```json
{
"sku": "orin-nano",
"variant": "orin-nano-8gb",
"mem_total_gb": 8,
"l4t_version": "36.4.0",
"product_model": "nvidia jetson orin nano developer kit",
"memory_kb": { "total": 8123456, "available": 4123456, "free": 1023456, "cached": 1234567, "swap_total": 0, "swap_free": 0 },
"default_systemd_target": "graphical.target",
"candidate_services": { "gdm3": { "active": "active", "enabled": "enabled" } },
"tegrastats_sample": "RAM 4011/8138MB (lfb 8x4MB) ...",
"nvmap": { "readable": false, "total_kb": 0, "top_clients": [] },
"procrank_top": [ { "pid": 4321, "pss_kb": 4000000, "cmd": "vllm" } ]
}
```
## Limitations
- Exact freed-memory deltas require a before snapshot, the user-approved change, cache flush when appropriate, and an after snapshot.
- NvMap attribution depends on host-visible debugfs access; if it is unavailable, report limited GPU memory attribution instead of guessing.
- Sandbox/container runs may not see host `/proc`, `tegrastats`, systemd, or NvMap data unless the runtime exposes them.
## Error handling
- If `scripts/audit.sh` cannot access host Jetson data, report the missing visibility and ask to rerun on the Jetson host or in a host-visible sandbox.
- If `scripts/drop_caches.sh` lacks root or passwordless `sudo -n`, report that cache dropping must be run on the host with sudo approval.
- If no before snapshot exists, say the exact reclaimed amount cannot be recovered from the current state alone and capture a new baseline for the next change.
## Safety
Read-only. `drop_caches` is non-destructive (kernel only releases pages it could reclaim under pressure anyway; `sync` runs first to preserve dirty data).
## Hand off to
- `jetson-headless-mode` — biggest single user-space win on systems still booting `graphical.target`.
- `jetson-inference-mem-tune` — when a model server is the top NvMap / PSS consumer.
- If runtime changes cannot hit the target, report that further reclamation is outside this skill's scope rather than suggesting unsafe boot-time edits.
すべてのファイル
8件のファイルjetson-memory-auditをインストール
スキルファイルをダウンロードし、.claude/skills/ ディレクトリに解凍してください。
ZIPをダウンロードリポジトリをクローンし、スキルファイルをプロジェクトにコピーしてください。
git clone https://github.com/NVIDIA/skills/tree/main/skills/jetson-memory-audit # Copy SKILL.md to your .claude/skills/ directory
コピー





家
