Multi-GPU Requirements
Distributed training uses one worker JVM per GPU and NCCL for inter-process communication. The full validated stack for multi-GPU work is:- Two or more NVIDIA GPUs visible to the same Linux host
- A working NVIDIA driver with CUDA 13 support
- One distinct CUDA device per rank
- Enough host RAM and CUDA VRAM for one model replica per rank
- A writable checkpoint directory and an available local TCP port
nREPL Dev Server
The repository’sdeps.edn includes a :dev alias that starts an nREPL server. Clone the repository and run:
127.0.0.1:7891 by default. Override the port or bind address with environment variables: