Requirements
1
Install CUDA runtime packages
On an Ubuntu host configured with NVIDIA’s CUDA package repository:
2
Verify GPU visibility
3
Set environment variables
These variables must be present before the Clojure process starts.
JAVA_TOOL_OPTIONS is inherited by every worker JVM spawned by the launcher.4
Confirm CUDA from the REPL
Running the Training Example
The shipped example inexamples/distributed_training.clj demonstrates a complete DDP training loop with AMP and checkpointing.
:float16 to enable dynamic loss scaling. Use :bfloat16 for autocast without a scaler. :accumulation 2 performs two micro-batches per optimizer step and suppresses DDP synchronization until the final micro-batch.
Launching Workers with dist/launch!
For your own training namespace, call dist/launch! directly with a config map:
Job Management
Worker Entrypoint Map
Clorch invokes your:main function with a single map argument:
Environment Variables Set by the Launcher
The launcher sets the following for each child JVM:Process Groups and Collectives
Workers launched by Clorch receive an initialized process group. For a custom launcher, initialize from environment variables:dist/with-process-group initializes the group, runs the body, and always destroys the group, even on error.
Available Collective Operations
Every rank must call collectives in the same order with compatible shapes, dtypes, and split sizes. Collectives operate on CUDA tensors in place.- all-reduce!
- broadcast!
- reduce!
- gather / scatter
- send / receive!
- barrier!
Reduces tensors across every rank in place.Supported
op values: :sum, :avg, :min, :max, :band, :bor, :bxor.DistributedDataParallel (DDP)
DDP replicates a model across all ranks, synchronizes gradients via all-reduce after each backward pass, and averages them before the optimizer step.ddp/optimizer-step! waits for pending reductions, commits averaged gradients to the model, and steps the optimizer.
Gradient Accumulation
Useddp/no-sync around every micro-batch except the last to suppress expensive all-reduce on intermediate steps:
Distributed Sampling
Each rank needs a disjoint, deterministic subset of the dataset.data/distributed-sampler partitions indices across replicas using a seeded shuffle.
data/distributed-sampler defaults :num-replicas and :rank from the WORLD_SIZE and RANK environment variables when omitted, so worker entrypoints do not need to thread these values explicitly.
Mixed Precision
For AMP within distributed training, see the AMP page for full details. A brief integration example:Checkpoints
Only rank zero writes the checkpoint. All ranks participate in save and restore barriers to keep execution synchronized.Saving
Restoring
What Gets Saved
The tensor archive and EDN metadata file are written to temporary files first and then moved atomically, so a crash during writing never leaves a corrupt checkpoint.
Current Verification Scope
The release suite covers CPU behavior and CUDA execution paths including NCCL, DDP backward, AMP overflow handling, fused scaled-dot-product attention, checkpoints, worker failures, and process cleanup. Two-rank validation on two RTX A5000 GPUs covers NCCL gradient reduction, parameter synchronization, bfloat16 AMP, gradient accumulation, and rank-zero checkpoint creation.Run the GPU release check after configuring the host to confirm your environment: