Build resource monitoring#
The cross container builds are long and QEMU-heavy, and the failure that hurts
most is silent resource exhaustion — an OOM-killed cc1plus at hour six, or a
full disk mid-export. linux/scripts/01-core/resource-monitor.sh records what
the system was doing throughout a build so you can answer, after the fact:
When, and during which build activity, did RAM / disk / CPU headroom run out?
What it captures#
A CSV time-series (resources-<run-id>.csv), one row per tick (default every
15 s), plus a computed resources-<run-id>.summary.txt on exit.
Column |
Meaning |
|---|---|
|
sample time (absolute + seconds since start) |
|
load average and CPU-busy % (from |
|
|
|
swap in use — any value >0 means real memory pressure |
|
free space / % used on the build filesystem |
|
live |
|
coarse build stage + the exact build log line active then |
The context column is the key one: at every peak-pressure moment the
summary tells you which file/target was compiling, so you can trace an OOM or
disk spike straight to the offending stage (e.g. torch’s aten TUs).
Overhead#
Negligible. Each tick is a handful of /proc reads plus ~3 short-lived forks
(df, pgrep, tail|grep). At 15 s that is far below the noise floor of a
32-core build and never competes with the compilers.
Automatic use#
build-cross-chain.sh starts it automatically for the whole run (gated by
RESOURCE_MONITOR, default on) and it self-terminates when the orchestrator
exits (via --watch-pid), writing the summary. Output lands in --log-dir.
# default: monitoring on, 15s interval, CSV+summary in the run's log dir
bash linux/scripts/build-cross-chain.sh --from-stage media --log-dir ~/build-logs
# disable it
RESOURCE_MONITOR=0 bash linux/scripts/build-cross-chain.sh ...
Manual use (any build, or a surgical build-cross-stage.sh)#
# start in the background, following whichever log is newest, self-ending with
# the build process:
setsid bash linux/scripts/01-core/resource-monitor.sh \
--out-dir ~/build-logs --run-id my-run --interval 15 \
--stage-log-dir ~/build-logs --disk-path / --watch-pid "$BUILD_PID" &
# ...or re-analyze a finished run's CSV at any time:
bash linux/scripts/01-core/resource-monitor.sh summarize ~/build-logs/resources-my-run.csv
Thresholds for the summary’s “near-OOM” / “low-disk” sample counts are tunable
with --near-oom-mb (default 4096) and --low-disk-gb (default 20).
Reading the results#
swap_used_mbclimbing ormem_avail_mbnear zero → you are at the OOM edge. Cross-reference thecontextat that row and lower the relevant*_MB_PER_JOBper build-parallelism-memory-tuning.md, or add RAM.disk_avail_gbdropping toward the low-disk threshold → prune the local buildkit cache (~/.cache/kata-buildcache) before the next run.compilerswell below core count whilemem_availis low → the stage is RAM-bound (expected for torch); more RAM, not more-j, is the lever.The CSV plots directly in any spreadsheet /
gnuplotif you want a timeline.
The CSV is rewritten on each start, so give each build a distinct
--run-id(the orchestrator passesCROSS_RUN_ID) to preserve history.