August 18, 2026
Android performance via ADB: how to read CPU, GPU, and NPU utilization
Performance work on Android lives or dies on one question: which chip is actually doing the work? A model you believe runs on the NPU may be silently falling back to CPU. A janky screen may be GPU-starved or CPU-bound — same symptom, opposite fixes. The answer is rarely where the docs say it should be, because Android’s hardware landscape is fragmented across SoC vendors, and each vendor exposes its performance counters differently — or not at all.
This guide is the map of what you can read over plain ADB, with no root and no vendor SDK where possible. CPU is the mature, well-standardized part: dumpsys cpuinfo, top, and /proc/stat work on essentially every device. GPU is where fragmentation begins — Qualcomm Adreno, ARM Mali, and the generic devfreq interface each have their own sysfs nodes — and we cover all three. NPU is the honest part of this guide: there is no standard interface, most vendors expose nothing, and the practical approach is indirect signals plus purpose-built profilers. Every command below runs through adb shell, so everything works over USB or Wi-Fi, from your desk, against production hardware.
Setup: USB debugging and device discovery#
On the phone: Settings → About phone → tap Build number seven times to unlock developer options, then Settings → System → Developer options → enable USB debugging. Plug in over USB, and the device will show an authorization prompt — accept it with your machine’s RSA fingerprint, or every subsequent command will fail against an “unauthorized” device.
Confirm the connection:
adb devices
A healthy line looks like ZX1G22B7QK device. If it says unauthorized, look at the phone and accept the prompt; if the list is empty, the two usual causes are a charge-only USB cable (data lines missing — maddeningly common) and a driver problem on the PC side.
For wireless work (Android 11+): Developer options → Wireless debugging → Pair device with pairing code shows an IP, port, and six-digit code:
adb pair 192.168.1.42:37851 # then type the pairing code when asked
adb connect 192.168.1.42:35577 # the port shown on the main wireless debugging screen
Pairing port and connect port are different numbers — the pair dialog shows one, the main screen the other. Once connected, every command below works identically over Wi-Fi.
CPU utilization: the mature part#
Three tools, three altitudes: top for live process hunting, dumpsys cpuinfo for the aggregated trustworthy number, and /proc/stat for computing overall utilization yourself.
top: the live process view#
adb shell top -m 10
-m 10 limits output to the top 10 processes by CPU; add -n 1 for a single snapshot (otherwise top refreshes forever, which is what you want interactively and wrong for scripts) and -d 5 to set a 5-second refresh. The columns worth your attention: %CPU (per-process, can exceed 100% — it is summed across cores, so 400% on an 8-core device means half the machine), VIRT/RES (memory), and the thread count. This is the tool for “what is eating my battery right now” questions.
Two limitations to know. Android’s top differs from desktop Linux flags across versions (older builds want -m, newer accept -n/-d with different semantics — if a flag errors, run top --help once and read your device’s dialect). And per-process numbers attribute kernel time somewhat loosely on some kernels; treat them as a hunting hint, then confirm with dumpsys.
dumpsys cpuinfo: the number you quote#
adb shell dumpsys cpuinfo
This aggregates CPU time from the kernel’s accounting over a rolling window and prints, per process, the user/system split — plus a TOTAL line with overall utilization against core count. It is more reliable than a single top snapshot because it averages over the window rather than sampling one instant, and its per-process attribution comes from the same counters the platform’s own tooling uses. When you need one reproducible number for a bug report — “our app idles at 3% CPU on device X” — this is the command.
/proc/stat: computing overall utilization yourself#
For the ground truth of whole-SoC utilization — including all the work top attributes to no process — read the kernel’s cumulative counters:
adb shell cat /proc/stat
The first line, cpu, aggregates all cores; the following cpu0…cpuN lines are per-core. The first four numbers of each line are, in order: user, nice, system, idle (kernel versions add more fields after — irq, softirq, steal — which you can include or ignore as long as you are consistent).
These are cumulative jiffies since boot, so a single read tells you nothing about the current rate. Utilization is computed from two reads and the delta between them:
busy = (user2 - user1) + (nice2 - nice1) + (system2 - system1)
total = busy + (idle2 - idle1)
utilization% = 100 * busy / total
Read once, wait a fixed interval (say 5 seconds), read again, subtract paired fields, and the ratio is the average utilization over exactly that window. Two practical notes: include the extra fields (irq, softirq, steal) in the total if your kernel reports them — they are time the CPU was not idle — and remember the result is an average; a 100% spike for 200 ms inside a 5-second window reads as 4%. Per-core utilization is the same arithmetic on the cpu0…cpuN lines, and it is worth doing: Android schedules threads per-core and big-core-saturation with idle little cores has a completely different diagnosis than even spread.
Single-process focus: top -p and finding the PID#
adb shell pidof com.example.app
adb shell ps -A | grep example
ps -A lists all processes; grep narrows it. On older Android builds, plain ps already lists everything — the -A is the POSIX form that works everywhere modern. With a PID in hand:
adb shell top -p 12345 -d 2 -n 3
Three samples, two seconds apart, one process. Combine with the loop technique at the end of this guide to log a single process for minutes.
Per-core frequencies#
adb shell cat /sys/devices/system/cpu/cpu0/cpufreq/scaling_cur_freq
Frequencies are in kHz (1900800 = 1.9 GHz). The wildcard form reads every core at once:
adb shell 'for c in /sys/devices/system/cpu/cpu*/cpufreq/scaling_cur_freq; do echo "$c: $(cat $c)"; done'
Frequency next to utilization is the diagnostic pair that separates “busy” from “busy and boosting”: 30% utilization at max clock is a thermal or scheduler story; 30% at idle clocks is a quiet machine. scaling_max_freq and cpuinfo_max_freq sit in the same directories and tell you the ceiling each core is currently allowed versus its hardware maximum — the two diverge during thermal throttling, which is exactly when you want to be reading them.
GPU utilization: the fragmented part#
There is no dumpsys gpuinfo. GPU counters live in sysfs nodes that differ per vendor, because each vendor’s driver exposes its own instrumentation. Three families cover nearly every device on the market.
Qualcomm Adreno: the kgsl nodes#
Qualcomm’s Adreno driver exposes itself through the KGSL interface, and on the overwhelming majority of Snapdragon devices:
adb shell cat /sys/class/kgsl/kgsl-3d0/gpubusy
Output is two numbers, e.g.:
43 128
The first is busy ticks, the second total ticks over the window — utilization is their ratio: 43/128 ≈ 34%. This is the canonical Adreno GPU usage number, and the one system monitors on Snapdragon devices read.
Frequency and model sit next door:
adb shell cat /sys/class/kgsl/kgsl-3d0/gpuclk
adb shell cat /sys/class/kgsl/kgsl-3d0/max_gpuclk
adb shell cat /sys/class/kgsl/kgsl-3d0/gpu_model
Some firmware builds rename or relocate the frequency file (devfreq/gpuclk, clock_mhz); if one path errors, ls /sys/class/kgsl/kgsl-3d0/ shows what your device actually offers — a habit that pays off across this entire section, because sysfs layout varies by kernel even within one vendor.
ARM Mali: device-specific nodes#
Mali (MediaTek, Exynos, Kirin, many Snapdragons’ siblings) exposes counters through its device node, but the path moved between driver generations. Two places to look:
adb shell cat /sys/class/misc/mali0/device/utilisation
adb shell ls /sys/class/misc/mali0/device/devfreq/
On kernels that provide it, utilisation prints a ready-made percentage. More universally, the Mali driver registers with the kernel’s devfreq framework — the standard interface for frequency-scalable devices — and its directory contains the counters (below).
The generic devfreq interface: works on most SoCs#
Devfreq is the kernel’s generic frequency-scaling framework, and nearly every modern SoC’s GPU registers with it — which makes this the most portable approach across vendors:
adb shell ls /sys/class/devfreq/
You will see device nodes named after their hardware, e.g. 1c00000.qcom,kgsl-3d0 (Adreno), ff9a0000.gpu (Rockchip Mali), 13000000.mali (MediaTek). Identify the GPU by the name — mali, kgsl, gpu — then read its directory:
adb shell cat /sys/class/devfreq/ff9a0000.gpu/cur_freq
adb shell cat /sys/class/devfreq/ff9a0000.gpu/available_frequencies
adb shell cat /sys/class/devfreq/ff9a0000.gpu/load
cur_freq is the current frequency; available_frequencies lists the ladder the GPU can climb; load (or on some kernels a busy_time/total_time pair in a governor subdirectory) gives the utilization figure. When you find busy_time and total_time, the ratio between consecutive reads — reset or accumulate depending on kernel — is your utilization; the same two-read delta logic from /proc/stat applies, since these too are cumulative counters on most implementations.
dumpsys gfxinfo: frame timing, not utilization#
adb shell dumpsys gfxinfo com.example.app framestats
Android 6.0+ provides per-app frame statistics: render time distribution, jank counts, and (with framestats) per-frame timestamps precise enough to profile the render pipeline. It is not GPU utilization — frames can be slow because of CPU-side layout work, shader compilation, or GPU saturation — but it is the consequence metric: users feel frame time, not utilization. The strongest diagnostic move available over plain ADB is reading gfxinfo and the GPU busy node together: slow frames plus idle GPU points at CPU/layout; slow frames plus saturated GPU points at shaders or resolution.
NPU utilization: the honest part#
Here is where the story gets uncomfortable, and pretending otherwise would make this guide useless. There is no standard Android interface for NPU utilization. The NPU (or DSP, or “AI accelerator” under any of its marketing names) is vendor silicon behind vendor drivers, and the vendors do not expose utilization counters through sysfs or dumpsys:
- Qualcomm Hexagon DSP (the cDSP that runs most QNN/Verso inference): no documented utilization node. Some kernels expose clock-related entries under
cdsp/dspclkpaths, but the overwhelming majority of production devices expose nothing readable. - MediaTek APU: no public interface. The counter exists — MediaTek’s own tools read it — but it is not exposed for third-party ADB access.
- Samsung’s NPU (the Eniso/Exynos NPU stack): similarly closed.
- Huawei’s Da Vinci NPU: no public counter interface.
The consequence: over plain ADB, you cannot directly read NPU utilization on essentially any production device. What you can do is reason from indirect signals, and they are genuinely informative:
- CPU usage during inference (
dumpsys cpuinfowhile running your workload): if CPU stays high while the model runs, your inference is executing on CPU — the fallback path — not the accelerator you targeted. Falling back to CPU is the single most common “why is my model slow” cause, and CPU utilization is the cleanest fingerprint of it. - GPU busy during inference (the nodes above): high Adreno/Mali utilization while a model runs means the model is running on the GPU delegate, not the NPU.
- Native memory growth (
adb shell dumpsys meminfo <package>): accelerator runtimes allocate large DMA/ion buffers when they genuinely engage the hardware; memory profile rising into hundreds of MB is soft evidence the NPU path is active, though it is inference-context allocation, not proof of sustained utilization.
When you need the real counters, you need the vendor profilers — named here so you know what to search for. Android GPU Inspector (AGI), Google’s official profiler, reads GPU internals and frame timing on a supported subset of Adreno and Mali devices. Snapdragon Profiler is Qualcomm’s tool and can see Hexagon/cDSP activity on Snapdragon. ARM Streamline covers Mali-based SoCs. The pattern to internalize: these tools install privileged components or use vendor-private interfaces to reach counters that plain ADB cannot — the fragmentation isn’t an oversight to work around; it is the current shape of the ecosystem.
Combining it: sampling loops and logging#
One-shot reads answer “what is it doing right now”. Performance questions need trends, and ADB shell gives you standard shell tools to build them.
The one-line monitor — read a node every second, forever:
adb shell 'while true; do cat /sys/class/kgsl/kgsl-3d0/gpubusy; sleep 1; done'
Everything inside the single quotes runs on the device; sleep 1 sets the cadence, and swapping the cat target turns it into a CPU-frequency monitor or anything else. On devices whose toybox includes watch (many do):
adb shell 'watch -n 1 cat /sys/class/kgsl/kgsl-3d0/gpubusy'
Timestamped output redirected to a file on the PC gives you a log for graphing:
adb shell 'while true; do echo "$(date +%H:%M:%S) $(cat /sys/class/kgsl/kgsl-3d0/gpubusy)"; sleep 2; done' > gpu-log.txt
Note the quoting: the while loop runs on the device, the redirect happens on the PC, so the file lands on your machine without any pull step. A practical workflow for inference benchmarking: start one logged loop for GPU busy, another for /proc/stat, run the model for 60 seconds, stop both, and the two CSVs side by side tell you which chip ran it and how hard. If you want timestamps the device computes (immune to USB latency jitter), put the date inside the loop as above rather than stamping on the PC side.
FAQ#
My device has no /sys/class/kgsl directory. Why?#
Because it is not a Qualcomm device. KGSL is the Adreno driver’s interface, so it exists on Snapdragon SoCs and nothing else. On a MediaTek or Exynos device, go straight to the devfreq listing (ls /sys/class/devfreq/) and look for a node with mali or gpu in the name, or check the Mali paths in the GPU section. Every GPU made in the last decade registers with devfreq, which is why that interface is the portable first move on unfamiliar hardware.
Do I need root for any of this?#
Mostly no, with a per-node caveat. dumpsys cpuinfo, top, /proc/stat, dumpsys gfxinfo, and dumpsys meminfo all work unrooted on every device — they are the standard debugging surface. Sysfs nodes are individually permissioned by the kernel builder: the kgsl gpubusy node is world-readable on most production Snapdragons, some devfreq entries are restricted, and vendor-specific counters occasionally root-only. You discover your device’s policy empirically: cat the node, and an empty result or permission error tells you. Rooting just to read a counter is almost never worth it — the dumpsys surface plus the world-readable sysfs nodes covers the vast majority of real questions.
Utilization reads zero while my model is running. What does that mean?#
Almost always: your model is not executing where you think it is. Zero GPU busy during inference means the GPU delegate is not engaged — the runtime fell back to CPU (check dumpsys cpuinfo: high CPU during inference confirms it) — and the absence of any readable NPU counter means you cannot distinguish “on NPU” from “on some path with no counter” directly. This is precisely the situation the indirect signals section addresses: CPU activity, GPU activity, and native memory growth triangulate the answer when a direct read is impossible. Delegate configuration (the NNAPI/GPU/CPU choice in your inference runtime) is the first thing to verify.
Frame rate versus GPU utilization — what is the difference?#
Frame rate is the rate of output: how many frames per second the display pipeline delivers. GPU utilization is how busy the GPU’s processing resources were, whether or not output resulted. They are related but not interchangeable: a device can render 60 fps with a nearly idle GPU (simple scene) or miss 60 fps with the GPU 90% saturated (heavy scene), and GPU utilization can sit high while frame rate drops — the exact signature of GPU-bound jank. Frame rate tells you what the user experienced; utilization tells you which component to blame and whether you have headroom. dumpsys gfxinfo measures the first, the sysfs nodes measure the second, and diagnosis means reading both: the failure modes “GPU too slow” and “GPU starved by CPU” produce the same frame rate and opposite utilization signatures.
Closing thoughts#
The three hardware families could not be more different in how knowable they are, and matching your expectations to that reality is most of the skill. CPU: standardized, well-attributed, dumpsys cpuinfo plus /proc/stat delta math gets you publication-grade numbers on any device. GPU: fragmented but readable — learn the three interfaces (kgsl for Adreno, the Mali nodes, devfreq for everything else) and the ls-first habit when a path 404s. NPU: no standard interface, and the professional move is not to wish for one but to triangulate — CPU and GPU utilization as fallback detectors, memory growth as engagement evidence, and vendor profilers when the question truly requires accelerator counters. Wire it together with a logged sampling loop, and a USB cable becomes a respectable performance lab.