Tools

axon: use a GPU that is in another computer

Run your CUDA program. The same binary, not rebuilt, not patched, not put in a container. It executes on the GPU in the workstation across the room, or in the machine under someone else's desk, and the program never learns the difference.

What this looks like

Blender, Julia, llama.cpp or anything else that loads libcuda.so.1 runs on your machine and computes on a different one. The machine holding the GPU does not have to be running Nerv, and it does not have to be running Linux — a Windows box with a spare GPU is a valid host.

zsh
# the GPU in this machine
~ ❯ ./train.py

# the GPU in that machine. same binary.
~ ❯ AXON_HOST=zgx ./train.py

# and afterwards, where did the work actually run?
~ ❯ axon top

A kernel dispatched this way has been executed on a remote A800 through the full stack and returned results bit-identical to running it natively.

How it works, and what it costs

On a Nerv machine no program talks to the GPU driver. A single daemon, axond, owns the device, and everything else is a client of it over a socket. That includes programs on the same machine: there is no separate local path to fall back to.

This is the part that makes remote execution work rather than a retrofit. A local client and a remote client take the same code path, so which machine holds the GPU becomes a parameter instead of an assumption. The daemon handles discovery and pairing, then hands the program a connected socket and steps out of the data path, so it can restart without killing a running job.

The cost is one boundary crossing. Locally that round trip is 1.9 µs over a shared-memory ring; over TCP it is 29 µs. Neither figure matters much on its own, because the crossings get batched: collapsing a repeated sequence of them into a single replayed graph measured 4.6–6.3× faster than issuing them one at a time.

What else falls out of it

Allocation stops costing a driver call. cuMemAlloc is 63 µs for a single 4 KiB page, per call rather than per byte, and the driver does no pooling. Suballocating from one reservation measured 3.8× cheaper to acquire and 44.8× cheaper to release for the same 4 GiB. Minted addresses make the substitution invisible to the application.

One GPU job stops blocking the next one. On unified memory a long-running GPU process fragments the pool badly enough that the next process cannot start, and the usual fix is to quit the first one. Measured on a GB10: the count of free 2 MiB blocks went from zero, to 248, back to zero as a single server was stopped and restarted. Under axon no application process ever creates a driver context, so there is nothing to fragment the pool in the first place.

You can finally see who is using the GPU. One process observes every allocation and knows which client asked for it, so per-client accounting and cgroup v2 limits come for free. This is a problem the industry has open: the kernel's DRM cgroup work is unfinished, Windows reports per-process GPU use without letting you bound it, and MPS and MIG partition the device instead of scheduling it.

A running GPU job can be suspended and moved. Clients hold no device state, so checkpointing one costs almost nothing. Stop a job, snapshot it, resume it later or on a different machine.

Why this has to be the distribution

Policy is real only when bypass is impossible. A library shim is advisory: any process can dlopen the vendor library and walk around the accounting, the limits and the arena. Making it mandatory is three packaging decisions, and only whoever builds the package set can make them.

  • The only libcuda.so.1 in the package set is axon's.
  • /dev/nvidia* is permissioned so that only the daemon may open it.
  • The arena is reserved on the kernel command line, before anything has fragmented memory.

A container cannot do any of it. A normal distribution cannot do the first two.