Skip to main content
Clorch runs one JVM per CUDA device and coordinates ranks with NCCL. The distributed API covers process-group lifecycle, collective operations, local worker launch, distributed sampling, synchronous data-parallel training with DDP, automatic mixed precision, and rank-zero checkpoints. This page covers every aspect of that system in depth.

Requirements

Clorch currently supports only the :nccl backend. Gloo, RPC, FSDP, tensor parallelism, and elastic membership are not yet implemented.
1

Install CUDA runtime packages

On an Ubuntu host configured with NVIDIA’s CUDA package repository:
2

Verify GPU visibility

3

Set environment variables

These variables must be present before the Clojure process starts. JAVA_TOOL_OPTIONS is inherited by every worker JVM spawned by the launcher.
4

Confirm CUDA from the REPL


Running the Training Example

The shipped example in examples/distributed_training.clj demonstrates a complete DDP training loop with AMP and checkpointing.
Use :float16 to enable dynamic loss scaling. Use :bfloat16 for autocast without a scaler. :accumulation 2 performs two micro-batches per optimizer step and suppresses DDP synchronization until the final micro-batch.

Launching Workers with dist/launch!

For your own training namespace, call dist/launch! directly with a config map:

Job Management

Worker Entrypoint Map

Clorch invokes your :main function with a single map argument:

Environment Variables Set by the Launcher

The launcher sets the following for each child JVM:

Process Groups and Collectives

Workers launched by Clorch receive an initialized process group. For a custom launcher, initialize from environment variables:
dist/with-process-group initializes the group, runs the body, and always destroys the group, even on error.

Available Collective Operations

Every rank must call collectives in the same order with compatible shapes, dtypes, and split sizes. Collectives operate on CUDA tensors in place.
Reduces tensors across every rank in place.
Supported op values: :sum, :avg, :min, :max, :band, :bor, :bxor.

DistributedDataParallel (DDP)

DDP replicates a model across all ranks, synchronizes gradients via all-reduce after each backward pass, and averages them before the optimizer step.
The constructor broadcasts parameters from rank zero, verifying signature consistency across ranks. Backward hooks bucket local gradients and initiate asynchronous all-reduce operations. ddp/optimizer-step! waits for pending reductions, commits averaged gradients to the model, and steps the optimizer.
:find-unused-parameters? true and :gradient-as-bucket-view? true are unsupported and will throw during construction. Every synchronized backward pass must produce a gradient for every trainable parameter.

Gradient Accumulation

Use ddp/no-sync around every micro-batch except the last to suppress expensive all-reduce on intermediate steps:

Distributed Sampling

Each rank needs a disjoint, deterministic subset of the dataset. data/distributed-sampler partitions indices across replicas using a seeded shuffle.
Always call data/set-epoch! before each epoch. Omitting it causes every epoch to reuse the same permutation, which can cause training to overfit to the same mini-batches.
data/distributed-sampler defaults :num-replicas and :rank from the WORLD_SIZE and RANK environment variables when omitted, so worker entrypoints do not need to thread these values explicitly.

Mixed Precision

For AMP within distributed training, see the AMP page for full details. A brief integration example:

Checkpoints

Only rank zero writes the checkpoint. All ranks participate in save and restore barriers to keep execution synchronized.

Saving

Restoring

What Gets Saved

The tensor archive and EDN metadata file are written to temporary files first and then moved atomically, so a crash during writing never leaves a corrupt checkpoint.
Clorch does not capture CUDA generator state. If your data pipeline performs CUDA random operations (e.g., CUDA augmentations), save and reapply your CUDA seed manually when exact replay matters.

Current Verification Scope

The release suite covers CPU behavior and CUDA execution paths including NCCL, DDP backward, AMP overflow handling, fused scaled-dot-product attention, checkpoints, worker failures, and process cleanup. Two-rank validation on two RTX A5000 GPUs covers NCCL gradient reduction, parameter synchronization, bfloat16 AMP, gradient accumulation, and rank-zero checkpoint creation.
Run the GPU release check after configuring the host to confirm your environment: