옵션
집집 Skill 보안 jetson-memory-audit

jetson-memory-audit

NVIDIA/skills NVIDIA/skills

Jetson DRAM 및 NvMap 사용량을 측정하고, 기준치(전후)를 캡처한 뒤, 실시간 감사 데이터를 통해 메모리 회수 여부를 확인합니다.

...모든 것을 확장하십시오
15
업데이트 된 시간 2026년 9월 24일

Jetson 메모리 감사

Jetson용 읽기 전용 메모리(ROM)에 초점을 맞춘 스냅샷과, 해제된 메모리가 캐시된 상태가 아닌 실제로 사용 가능한 메모리로 표시되는지 확인하는 drop_caches 검증 루프 헬퍼.

목적

현재 Jetson 메모리 소비량을 측정하고, 변경 전후의 기준치를 확보하며, 사용자가 승인한 변경 사항이 실제로 DRAM을 회수했는지 확인합니다. 컨테이너 크기, 모델 크기 또는 일반적인 프로세스 메모리 추정치 대신 실시간 디바이스 데이터를 사용합니다.

중요: vLLM / sglang 중지 후 메모리가 고착된 것처럼 보임 (JetPack 7.2 미만 / L4T r39.0 미만)

이는 JetPack 7.2 이전 또는 L4T r39.0 이전 버전의 Jetson 릴리스에서 가장 흔히 발생하는 메모리 관련 혼란입니다.

vLLM, sglang 또는 Ollama 서버(혹은 기타 CUDA 워크로드)를 중지한 후에도, 프로세스가 종료되었음에도 불구하고 free -h 또는 tegrastats에서 사용 가능 메모리로 표시되는 메모리가 회수되지 않을 수 있습니다. 또한 nvidia-smi에서도 GPU 사용 가능 메모리가 오해의 소지가 있을 정도로 낮게 표시될 수 있습니다.

근본 원인: CUDA 컨텍스트가 종료된 후, Thor RM(리소스 관리자)은 해제된 sysmem 페이지를 자체 풀에 보관합니다. Jetson과 같은 UMA(Unified Memory Architecture) 장치에서 cudaMemGetInfo는 RM 풀 상태를 읽으며, 새로운 프로세스가 실제로 사용할 수 있는 양보다 훨씬 적은 양의 여유 메모리를 보고합니다.

해결 방법(7.2 미만 JetPack 또는 r39.0 미만 L4T의 경우):

sudo sync && sudo sysctl -w vm.drop_caches=3

이 명령은 컨테이너 내부가 아닌 호스트에서 실행해야 합니다. 중요한 작업은 sudo sysctl -w vm.drop_caches=3이며, 재사용 가능한 페이지/dentry/inode 캐시가 삭제되기 전에 더티 데이터가 플러시되도록 이 명령 바로 앞에 sudo sync를 반드시 포함해야 합니다. 이 명령을 실행한 후, free -h 및 tegrastats 명령을 실행하면 실제 사용 가능한 메모리 용량이 반영됩니다.

이 문제가 발생하는 릴리스의 경우, 사용자가 다음과 같이 말할 때 이 명령을 권장합니다:

  • "vLLM/sglang을 중지했는데도 메모리가 해제되지 않았습니다"
  • "컨테이너가 종료된 후에도 tegrastats에 여전히 높은 사용량이 표시되는 이유는 무엇인가요?"
  • "아무것도 실행 중이 아닌데도 OOM 오류가 발생합니다"
  • "어제는 메모리 문제가 없었는데 지금은 꽉 찼어요"

7.2 미만의 JetPack 또는 r39.0 미만의 L4T에서, CUDA 워크로드 종료 후 메모리가 고정된 것처럼 보일 때 drop_caches를 사용하는 것이 신뢰할 수 있는 해결 방법입니다. 최신 릴리스에서는 동일한 증상이 관찰되고 사용자가 동의하는 경우에만 이 방법을 사용하십시오.

사용 시점

  • "이 Jetson에서 사용 중인 메모리는 얼마나 되나요? 무엇이 메모리를 차지하고 있나요?"
  • "GUI를 비활성화했거나, vLLM을 중지했거나, 컨테이너를 종료했는데 메모리가 실제로 해제되었나요?"
  • "워크로드를 중지했는데도 free -h 명령어에 여전히 사용 가능한 메모리가 적게 표시되는 이유는 무엇인가요?"
  • jetson-headless-mode 또는 기타 메모리 관련 변경 사항을 적용하기 전의 기준값으로 사용하며, 적용 후에도 실제 차이를 계산하기 위해 다시 확인합니다.

필수 조건

  • Jetson 호스트에서, 또는 호스트에서 /proc, /etc/nv_tegra_release, tegrastats 및 프로세스 데이터를 확인할 수 있는 샌드박스/컨테이너에서 실행하십시오.
  • NvMap debugfs 읽기 작업에는 루트 권한이 필요할 수 있습니다. 해당 권한이 없는 경우, 추측하기보다는 GPU 메모리 할당량이 제한적이라고 보고하십시오.
  • drop_caches.sh를 실행하려면 root 권한 또는 비밀번호가 필요 없는 sudo -n이 필요합니다. 사용자가 캐시 삭제를 명시적으로 승인한 후에만 실행하십시오.

사용 가능한 스크립트

스크립트 목적 인수
scripts/audit.sh 메모리 감사 워크플로우를 위해 jetson-diagnostic/scripts/snapshot.sh 에서 JSON 스냅샷을 출력합니다. 인수가 없습니다.
scripts/drop_caches.sh 회수 가능한 페이지/dentry/inode 캐시를 플러시하고, 메모리 변화량(전후)을 출력합니다. --mode 1|2|3, --quiet.

에이전트 런타임이 run_script를 지원하는 경우, 이를 사용하여 scripts/audit.sh 또는 scripts/drop_caches.sh를 실행하고 반환된 출력을 요약합니다. 그렇지 않은 경우, 리포지토리 루트에서 bash를 사용하여 스크립트를 실행합니다.

사용 방법

"현재 사용 중인 메모리 용량은 얼마인가?"와 같은 질문에 대해서는 scripts/audit.sh를 실행하고 JSON 스냅샷의 값만 보고하십시오.

보고 지침

헬퍼의 경로만 출력하거나 언급하지 마십시오. 헬퍼를 호출한 후 반환된 데이터를 요약하여 보고하십시오.

  • "현재 사용 중인 메모리 용량은 얼마인가?"라는 프롬프트의 경우, scripts/audit.sh를 실행하고 mem_total_gb, memory_kb.available, 그리고 procrank_top 프로세스 또는 nvmap.top_clients 소비자 중 가장 앞선 항목의 값을 인용하십시오.
  • GUI/데스크톱 메모리 관련 프롬프트의 경우, scripts/audit.sh를 실행하고 default_systemd_target과 candidate_services에 포함된 디스플레이 매니저(gdm3, gdm, lightdm, sddm 또는 display-manager) 를 보고하십시오. 어떤 것도 비활성화하지 말고, 해결 방안은 jetson-headless-mode에 위임하십시오.
  • 중지된 워크로드 후 캐시 삭제를 명시적으로 허용하는 경우, scripts/drop_caches.sh를 실행하고(기본적으로 sudo sync && sudo sysctl -w vm.drop_caches=3과 동일), 실행 전후의 free, available 및 cached 값의 변화를 보고하십시오. root 권한을 사용할 수 없는 경우, 호스트에서 sudo를 사용하여 실행해야 함을 설명하십시오.

에이전트 런타임이 이 스킬 디렉터리를 기준으로 헬퍼 스크립트를 실행하지 않는 경우, AgentSkills {baseDir} 자리 표시자를 사용하여 스크립트 경로를 해결하십시오:

{baseDir}/scripts/audit.sh
{baseDir}/scripts/drop_caches.sh

런타임이 스킬을 호출 가능한 도구로 명시적으로 등록하지 않는 한, jetson-memory-audit 를 도구 이름으로 호출하지 마십시오. 런타임이 스킬을 호출 가능한 도구로 명시적으로 등록한 경우가 아니라면, 에이전트 스킬은 일반적으로 명령어와 파일의 조합이며 직접적인 도구 함수가 아닙니다.

에이전트용 샌드박스 참고 사항: 이 스킬 파일을 볼 수 있다고 해서 Jetson 호스트 메모리 데이터에 대한 접근이 보장되는 것은 아닙니다. NemoClaw/OpenClaw 샌드박스 내에서 /proc/device-tree/model, /etc/nv_tegra_release, tegrastats, /sys/kernel/debug/nvmap 또는 호스트 프로세스 데이터가 누락된 경우, 샌드박스가 Jetson 호스트를 볼 수 없다고 판단하고 사용자에게 Jetson 호스트에서 실행하거나 호스트가 보이는 샌드박스 프로필로 다시 실행하도록 요청하십시오. 메모리 총량, 사용 가능 메모리, PSS, NvMap 또는 회수량 차이를 임의로 조작하지 마십시오.

"이 변경으로 인해 메모리가 얼마나 확보되었나요?"라는 질문에는 변경 전후의 차이를 사용하십시오. 컨테이너 크기, 이미지 크기, RSS 또는 변경 후 단일 스냅샷을 바탕으로 확보된 메모리를 추산해서는 안 됩니다.

  1. 변경 전, scripts/audit.sh를 실행하고 JSON 기준값을 저장하십시오.
  2. 사용자가 승인한 변경 사항(컨테이너 중지, 모드 전환, 튜닝 권장 사항 적용 등)을 적용하십시오.
  3. JetPack 7.2 미만 / L4T r39.0 미만 환경에서, 또는 최신 릴리스에서 동일한 메모리 고착 현상이 관찰될 경우, 호스트 (컨테이너 내부가 아님)에서 회수 가능한 페이지 캐시를 플러시하여 해제된 페이지가 캐시된 상태가 아닌 사용 가능한 상태로 표시되도록 하십시오:
    sudo sync && sudo sysctl -w vm.drop_caches=3
    
    
  4. scripts/audit.sh를 다시 실행하고 변경 전후의 memory_kb.available 값을 비교하십시오. 그 차이가 실제 회수된 메모리 양입니다.

사용자가 이미 변경을 적용했고 기준값이 없는 경우, 현재 스냅샷만으로는 정확히 확보된 용량을 파악할 수 없다고 설명하십시오. 다음 변경 사항을 측정할 수 있도록 지금 새로운 기준값을 캡처하십시오.

실시간 감사 데이터를 신뢰할 수 있는 기준으로 사용하십시오. 총 메모리, 사용 가능 메모리, NvMap 총량, PSS 값, 디스플레이 매니저 상태 및 절약된 용량 차이는 실제 기기에서 scripts/audit.sh, free -h 또는 tegrastats를 통해 얻어야 합니다. 해당 출력 결과에 수치가 나타나지 않는 경우, 임의로 추측하지 마십시오.

audit.sh의 출력 형식

{
  "sku": "orin-nano",
  "variant": "orin-nano-8gb",
  "mem_total_gb": 8,
  "l4t_version": "36.4.0",
  "product_model": "nvidia jetson orin nano developer kit",
  "memory_kb": { "total": 8123456, "available": 4123456, "free": 1023456, "cached": 1234567, "swap_total": 0, "swap_free": 0 },
  "default_systemd_target": "graphical.target",
  "candidate_services": { "gdm3": { "active": "active", "enabled": "enabled" } },
  "tegrastats_sample": "RAM 4011/8138MB (lfb 8x4MB) ...",
  "nvmap": { "readable": false, "total_kb": 0, "top_clients": [] },
  "procrank_top": [ { "pid": 4321, "pss_kb": 4000000, "cmd": "vllm" } ]
}

제한 사항

  • 정확한 해제 메모리 차이를 파악하려면 사전 스냅샷, 사용자가 승인한 변경 사항, 적절한 시점의 캐시 플러시, 그리고 사후 스냅샷이 필요합니다.
  • NvMap 할당 분석은 호스트에서 볼 수 있는 debugfs 액세스에 의존합니다. 이 액세스가 불가능한 경우, 추측 대신 제한된 GPU 메모리 할당 정보를 보고합니다.
  • 런타임이 노출하지 않는 한, 샌드박스/컨테이너 실행 환경에서는 호스트의 /proc, tegrastats, systemd 또는 NvMap 데이터를 볼 수 없을 수 있습니다.

오류 처리

  • scripts/audit.sh가 호스트 Jetson 데이터에 액세스할 수 없는 경우, 가시성이 부족함을 보고하고 Jetson 호스트 또는 호스트가 보이는 샌드박스에서 다시 실행할 것을 요청합니다.
  • scripts/drop_caches.sh에 루트 권한이나 비밀번호 없이 실행 가능한 sudo -n 권한이 없는 경우, sudo 승인을 받은 호스트에서 캐시 삭제를 실행해야 함을 보고하십시오.
  • 이전 스냅샷이 없는 경우, 현재 상태만으로는 정확한 회수량을 파악할 수 없음을 알리고 다음 변경을 위해 새로운 기준선을 캡처하십시오.

안전성

읽기 전용입니다. drop_caches는 비파괴적입니다(커널은 어차피 부하 상황에서 회수할 수 있는 페이지만 해제하며, 더티 데이터를 보존하기 위해 먼저 sync가 실행됩니다).

다음으로 인계

  • jetson-headless-mode — 여전히 graphical.target로 부팅되는 시스템에서 사용자 공간 측면에서 얻을 수 있는 가장 큰 이점입니다.
  • jetson-inference-mem-tune — 모델 서버가 NvMap / PSS의 최대 소비자일 때.
  • 런타임 변경 사항이 목표치를 달성할 수 없는 경우, 안전하지 않은 부팅 시 편집을 제안하기보다는 추가적인 메모리 회수가 이 스킬의 범위를 벗어난다는 사실을 보고하십시오.
GitHub에서 보기
---
name: jetson-memory-audit
description: Measure Jetson DRAM and NvMap usage, capture before/after baselines, and verify memory reclamation with live audit data.
license: Apache-2.0
---

# Jetson Memory Audit

Read-only memory-focused snapshot for a Jetson, plus the `drop_caches` verify-loop helper that confirms freed memory actually shows up as free instead of cached.

## Purpose

Measure current Jetson memory consumers, capture before/after baselines, and verify whether user-approved changes actually reclaimed DRAM. Use live device data rather than estimates from container size, model size, or generic process memory.

## CRITICAL: Memory appears stuck after stopping vLLM / sglang (JetPack below 7.2 / L4T below r39.0)

**This is the most common memory confusion on Jetson releases before JetPack 7.2 or before L4T r39.0.**

After you stop a vLLM, sglang, or Ollama server (or any CUDA workload), the memory shown as free by `free -h` or `tegrastats` may not recover — even though the process is gone. `nvidia-smi` may also show misleadingly low free GPU memory.

**Root cause:** The Thor RM (resource manager) holds freed sysmem pages in its own pool after a CUDA context exits. On Unified Memory Architecture (UMA) devices like Jetson, `cudaMemGetInfo` reads RM pool state and reports far less free memory than is actually available to a new process.

**Workaround (for JetPack below 7.2 or L4T below r39.0):**
```bash
sudo sync && sudo sysctl -w vm.drop_caches=3
```
Run this on the **host**, not inside a container. The important operation is `sudo sysctl -w vm.drop_caches=3`; keep `sudo sync` immediately before it so dirty data is flushed before reclaimable page/dentry/inode caches are dropped. After running it, `free -h` and `tegrastats` will reflect the true available memory.

For affected releases, recommend this command when a user says:
- "Memory didn't free after I stopped vLLM/sglang"
- "Why does tegrastats still show high usage after my container exited?"
- "OOM even though nothing is running"
- "Memory was fine yesterday but now it's full"

On JetPack below 7.2 or L4T below r39.0, `drop_caches` is the reliable workaround when memory appears stuck after a CUDA workload exits; on newer releases, use it only if the same symptom is observed and the user approves.

## When to use

- "How much memory is in use on this Jetson? What's holding it?"
- "I disabled the GUI / stopped vLLM / quit my container — did memory actually free?"
- "Why does `free -h` still show low free memory after I stopped my workload?"
- As the **baseline** before applying `jetson-headless-mode` or other memory-related changes, and again **after** to compute the actual delta.

## Prerequisites

- Run on the Jetson host, or in a sandbox/container with host-visible `/proc`, `/etc/nv_tegra_release`, `tegrastats`, and process data.
- NvMap debugfs reads may require root. If unavailable, report that GPU memory attribution is limited rather than guessing.
- `drop_caches.sh` requires root or passwordless `sudo -n`; run it only after the user explicitly authorizes cache dropping.

## Available Scripts

| Script | Purpose | Arguments |
|--------|---------|-----------|
| `scripts/audit.sh` | Emits a JSON snapshot from `jetson-diagnostic/scripts/snapshot.sh` for memory audit workflows. | No arguments. |
| `scripts/drop_caches.sh` | Flushes reclaimable page/dentry/inode caches and prints before/after memory deltas. | `--mode 1\|2\|3`, `--quiet`. |

If your agent runtime supports `run_script`, use it to run `scripts/audit.sh` or `scripts/drop_caches.sh` and summarize the returned output. Otherwise run the scripts with `bash` from the repository root.

## Instructions

For "how much memory is in use right now?" questions, run `scripts/audit.sh` and report only values from the JSON snapshot.

## Reporting guidance

Do not only print or mention the path to a helper. Invoke the helper and then summarize the returned data.

- For "how much memory is in use" prompts, run `scripts/audit.sh` and quote `mem_total_gb`, `memory_kb.available`, and the leading `procrank_top` process or `nvmap.top_clients` consumer.
- For GUI/desktop memory prompts, run `scripts/audit.sh` and report `default_systemd_target` plus any display manager in `candidate_services` (`gdm3`, `gdm`, `lightdm`, `sddm`, or `display-manager`). Do not disable anything; hand off to `jetson-headless-mode` for a plan.
- For prompts that explicitly authorize cache dropping after a stopped workload, run `scripts/drop_caches.sh` (equivalent to `sudo sync && sudo sysctl -w vm.drop_caches=3` by default) and report its before/after free, available, and cached deltas. If root is unavailable, explain that it must be run on the host with sudo.

If your agent runtime does not execute helper scripts relative to this skill directory, resolve script paths with the AgentSkills `{baseDir}` placeholder:

```bash
{baseDir}/scripts/audit.sh
{baseDir}/scripts/drop_caches.sh
```

Do not call `jetson-memory-audit` as a tool name unless the runtime explicitly registers skills as callable tools; Agent Skills are normally instructions plus files, not direct tool functions.

Sandbox note for agents: seeing this skill file does not guarantee access to Jetson host memory data. If `/proc/device-tree/model`, `/etc/nv_tegra_release`, `tegrastats`, `/sys/kernel/debug/nvmap`, or host process data are missing inside a NemoClaw/OpenClaw sandbox, say the sandbox lacks Jetson host visibility and ask the user to run on the Jetson host or relaunch with a host-visible sandbox profile. Do not fabricate memory totals, available memory, PSS, NvMap, or reclamation deltas.

For "how much memory did this change free?" questions, use a before/after delta. Do not estimate freed memory from container size, image size, RSS, or a single post-change snapshot.

1. Before the change, run `scripts/audit.sh` and save the JSON baseline.
2. Make the user-approved change (stop the container, switch mode, apply a tuning recommendation, etc.).
3. On JetPack below 7.2 / L4T below r39.0, or when the same stuck-memory symptom is observed on a newer release, flush reclaimable page cache on the **host** (not inside a container) so freed pages show up as free instead of cached:
   ```bash
   sudo sync && sudo sysctl -w vm.drop_caches=3
   ```
4. Re-run `scripts/audit.sh` and compare `memory_kb.available` before vs after — that delta is the real reclamation.

If the user already made the change and no baseline exists, say that the exact freed amount cannot be recovered from the current snapshot alone. Capture a new baseline now so the next change can be measured.

Use live audit data as the source of truth. Memory totals, available memory, NvMap totals, PSS values, display-manager state, and savings deltas must come from `scripts/audit.sh`, `free -h`, or `tegrastats` on the actual device. If a number is not present in those outputs, do not guess it.

## Output contract for `audit.sh`

```json
{
  "sku": "orin-nano",
  "variant": "orin-nano-8gb",
  "mem_total_gb": 8,
  "l4t_version": "36.4.0",
  "product_model": "nvidia jetson orin nano developer kit",
  "memory_kb": { "total": 8123456, "available": 4123456, "free": 1023456, "cached": 1234567, "swap_total": 0, "swap_free": 0 },
  "default_systemd_target": "graphical.target",
  "candidate_services": { "gdm3": { "active": "active", "enabled": "enabled" } },
  "tegrastats_sample": "RAM 4011/8138MB (lfb 8x4MB) ...",
  "nvmap": { "readable": false, "total_kb": 0, "top_clients": [] },
  "procrank_top": [ { "pid": 4321, "pss_kb": 4000000, "cmd": "vllm" } ]
}
```

## Limitations

- Exact freed-memory deltas require a before snapshot, the user-approved change, cache flush when appropriate, and an after snapshot.
- NvMap attribution depends on host-visible debugfs access; if it is unavailable, report limited GPU memory attribution instead of guessing.
- Sandbox/container runs may not see host `/proc`, `tegrastats`, systemd, or NvMap data unless the runtime exposes them.

## Error handling

- If `scripts/audit.sh` cannot access host Jetson data, report the missing visibility and ask to rerun on the Jetson host or in a host-visible sandbox.
- If `scripts/drop_caches.sh` lacks root or passwordless `sudo -n`, report that cache dropping must be run on the host with sudo approval.
- If no before snapshot exists, say the exact reclaimed amount cannot be recovered from the current state alone and capture a new baseline for the next change.

## Safety

Read-only. `drop_caches` is non-destructive (kernel only releases pages it could reclaim under pressure anyway; `sync` runs first to preserve dirty data).

## Hand off to

- `jetson-headless-mode` — biggest single user-space win on systems still booting `graphical.target`.
- `jetson-inference-mem-tune` — when a model server is the top NvMap / PSS consumer.
- If runtime changes cannot hit the target, report that further reclamation is outside this skill's scope rather than suggesting unsafe boot-time edits.

jetson-memory-audit 설치

스킬 파일을 다운로드하여 .claude/skills/ 디렉터리에 압축을 풀어주세요.

ZIP 다운로드

저장소를 클론하고 스킬 파일을 프로젝트에 복사하세요.

git clone https://github.com/NVIDIA/skills/tree/main/skills/jetson-memory-audit # Copy SKILL.md to your .claude/skills/ directory

복사 복사
빠른 설정: 스킬 폴더를 .claude/skills/로 복사하세요. Claude가 해당 스킬을 자동으로 감지하여 사용할 것입니다.
저장소 NVIDIA/skills

관련 스킬

gmgn-portfolio
업데이트 된 시간 2026년 7월 1일
device-integrity
업데이트 된 시간 2026년 6월 29일
zeroize-audit
업데이트 된 시간 2026년 7월 1일
flutter-use-http-package
업데이트 된 시간 2026년 6월 30일
OR