Skip to main content
Clorch is distributed as a git dependency and pulls its native PyTorch binaries through JavaCPP. There is no separate native install step for CPU usage — the JVM downloads and caches the right platform binaries on the first run. GPU support requires additional system packages and a compatible NVIDIA driver, which are covered below.

Multi-GPU Requirements

Distributed training uses one worker JVM per GPU and NCCL for inter-process communication. The full validated stack for multi-GPU work is:
Additional runtime requirements:
  • Two or more NVIDIA GPUs visible to the same Linux host
  • A working NVIDIA driver with CUDA 13 support
  • One distinct CUDA device per rank
  • Enough host RAM and CUDA VRAM for one model replica per rank
  • A writable checkpoint directory and an available local TCP port
Read the Distributed CUDA Training guide for worker code, NCCL collectives, DDP, AMP, gradient accumulation, and checkpoint restore.

nREPL Dev Server

The repository’s deps.edn includes a :dev alias that starts an nREPL server. Clone the repository and run:
The nREPL server listens on 127.0.0.1:7891 by default. Override the port or bind address with environment variables:
When working with the repository examples, use clj -M:dev rather than plain clj so that the examples/ directory is on the classpath and example namespaces like distributed-training and pytorch-basics-tutorial can be required directly.