Preventing DGX Spark freezes under heavy memory load Print

  • dgx-spark, memory, oom, cgroups, watchdog, vllm
  • 0

DGX Spark uses unified memory shared between the CPU and the GPU. If a workload takes all of it, the operating system is left without memory and the whole machine can become unresponsive. This article covers how to keep headroom for the OS, how to cap a user's and a container's memory so that Linux terminates the runaway workload instead, and how to configure the hardware watchdog as a last-resort recovery from a complete freeze.

User names below are examples. Substitute your own.

1. Keep memory free for the operating system

Do not plan to use all available memory. Keep approximately 20 GB free for the operating system.

Before starting a heavy workload, check how much memory is available:

free -h

For vLLM, cap the share of memory it may claim:

vllm serve MODEL --gpu-memory-utilization 0.7

Avoid running several memory-intensive workloads at the same time if their combined usage approaches the system limit.

2. Limit a user's memory

We recommend limiting the working user account to approximately 108 GB, which leaves about 20 GB for the operating system. For a user named sayuser:

UIDN=$(id -u sayuser)

sudo systemctl set-property user-${UIDN}.slice \
    MemoryHigh=100G \
    MemoryMax=108G

Above MemoryHigh the user's processes are throttled and pushed to reclaim memory; at MemoryMax the kernel terminates a process inside that user's slice. The rest of the system is not affected.

Check the applied limits:

systemctl show user-${UIDN}.slice \
    -p MemoryHigh \
    -p MemoryMax

3. Limit Docker containers

Containers started through Docker run under the Docker service, not under the user's slice, so the limit above does not apply to them. Give every container its own limit:

docker run \
    --memory=108g \
    --memory-swap=108g \
    ...

Do not use --oom-kill-disable. The goal is to let Linux terminate an excessive workload instead of allowing the entire DGX Spark to become unresponsive.

4. Automatic reboot after a complete freeze

First check whether a hardware watchdog device is available:

ls -l /dev/watchdog*
sudo wdctl

If /dev/watchdog0 is present, configure systemd to use it:

sudo mkdir -p /etc/systemd/system.conf.d

printf '[Manager]\nRuntimeWatchdogSec=60s\nRebootWatchdogSec=5min\n' | \
sudo tee /etc/systemd/system.conf.d/10-watchdog.conf

Apply the configuration:

sudo systemctl daemon-reexec

Verify:

systemctl show --property=RuntimeWatchdogUSec
sudo wdctl

If the system freezes completely and stops servicing the hardware watchdog, it reboots automatically.

If /dev/watchdog0 is not present, do not change the watchdog configuration. Open a ticket from your client area instead.

5. Important: do not reboot on OOM

Do not configure an automatic reboot on every out-of-memory event (for example via vm.panic_on_oom). A workload that runs out of memory should be terminated while the DGX Spark stays online. Automatic reboot is only the last-resort recovery for a complete system freeze.


Was this answer helpful?

« Back