DGX Spark uses unified memory shared between the CPU and the GPU. If a workload takes all of it, the operating system is left without memory and the whole machine can become unresponsive. This article covers how to keep headroom for the OS, how to cap a user's and a container's memory so that Linux terminates the runaway workload instead, and how to configure the hardware watchdog as a last-resort recovery from a complete freeze.
User names below are examples. Substitute your own.
1. Keep memory free for the operating system
Do not plan to use all available memory. Keep approximately 20 GB free for the operating system.
Before starting a heavy workload, check how much memory is available:
free -h
For vLLM, cap the share of memory it may claim:
vllm serve MODEL --gpu-memory-utilization 0.7
Avoid running several memory-intensive workloads at the same time if their combined usage approaches the system limit.
2. Limit a user's memory
We recommend limiting the working user account to approximately 108 GB, which leaves about 20 GB for the operating system. For a user named sayuser:
UIDN=$(id -u sayuser)
sudo systemctl set-property user-${UIDN}.slice \
MemoryHigh=100G \
MemoryMax=108G
Above MemoryHigh the user's processes are throttled and pushed to reclaim memory; at MemoryMax the kernel terminates a process inside that user's slice. The rest of the system is not affected.
Check the applied limits:
systemctl show user-${UIDN}.slice \
-p MemoryHigh \
-p MemoryMax
3. Limit Docker containers
Containers started through Docker run under the Docker service, not under the user's slice, so the limit above does not apply to them. Give every container its own limit:
docker run \
--memory=108g \
--memory-swap=108g \
...
Do not use --oom-kill-disable. The goal is to let Linux terminate an excessive workload instead of allowing the entire DGX Spark to become unresponsive.
4. Automatic reboot after a complete freeze
First check whether a hardware watchdog device is available:
ls -l /dev/watchdog*
sudo wdctl
If /dev/watchdog0 is present, configure systemd to use it:
sudo mkdir -p /etc/systemd/system.conf.d
printf '[Manager]\nRuntimeWatchdogSec=60s\nRebootWatchdogSec=5min\n' | \
sudo tee /etc/systemd/system.conf.d/10-watchdog.conf
Apply the configuration:
sudo systemctl daemon-reexec
Verify:
systemctl show --property=RuntimeWatchdogUSec
sudo wdctl
If the system freezes completely and stops servicing the hardware watchdog, it reboots automatically.
If /dev/watchdog0 is not present, do not change the watchdog configuration. Open a ticket from your client area instead.
5. Important: do not reboot on OOM
Do not configure an automatic reboot on every out-of-memory event (for example via vm.panic_on_oom). A workload that runs out of memory should be terminated while the DGX Spark stays online. Automatic reboot is only the last-resort recovery for a complete system freeze.