defmodel to a Llama-3-style chat loop with KV caching and a multi-GPU distributed training harness with AMP and gradient accumulation. Each example is a runnable Clojure namespace; load it in a REPL or evaluate it top-to-bottom. The snippets below are taken directly from the source files.
Simple Model
Minimal
defmodel with one training stepPyTorch Basics Tutorial
Tensors, datasets, optimization, and model save/load
Autograd Tutorial
Gradient computation and computation-graph behavior
Synthetic Training
End-to-end training loop on generated data
Modern Llama
RoPE, GQA, SwiGLU, and incremental KV caching
NanoChat
Compact Llama-style training, checkpointing, and chat generation
Distributed CUDA Training
NCCL workers, DDP, AMP, and gradient accumulation
Bayesian Linear Regression MCMC
Posterior sampling with Metropolis-Hastings
Einsum eDSL
Declarative tensor contractions with
einSimple Model
Source:examples/simple.clj
The simplest possible end-to-end demonstration of Clorch. It defines a two-layer MLP using the defmodel macro, constructs a fake batch, runs a forward pass, computes MSE loss, and performs one Adam optimizer step — all inside a with-torch scope to manage native memory correctly.
defmodel accepts constructor arguments, a binding vector of registered sub-modules, and a forward form. Registered fields participate automatically in nn/parameters, nn/to, and state dictionaries.
PyTorch Basics Tutorial
Source:examples/pytorch_basics_tutorial.clj
A direct Clojure port of the official PyTorch “Learn the Basics” tutorial. It demonstrates tensor creation, a custom dataset backed by the data/dataset protocol, a multi-layer perceptron trained over multiple epochs with SGD and cross-entropy loss, and model serialization via torch/save and torch/load.
The training loop pattern — zero-grad → forward → loss → backward → step — matches PyTorch exactly and is idiomatic for all Clorch training code.
data/dataset with :size and :get-item callbacks, returning {:data … :target …} maps that the dataloader batches automatically:
Autograd Tutorial
Source:examples/autograd_tutorial.clj
A port of Sebastian Raschka’s Automatic Differentiation Made Easy tutorial. It covers scalar and tensor gradients, calling backward to populate .grad fields, reading gradients with autograd/grad, and iterating over tensor slices with torch/tseq.
Synthetic Training
Source:examples/synthetic.clj
An end-to-end training demonstration on procedurally generated multi-class data. It creates Gaussian clusters in 10-dimensional space, wraps them in a tensor-dataset, runs a two-layer ReLU MLP with SGD and cross-entropy loss for several epochs, and prints the epoch-averaged loss. It also demonstrates explicit cleanup of native resources with data/cleanup-data!.
Modern Llama
Source:examples/modern_llama.clj
Demonstrates a single Llama-style transformer block built from Clorch’s LLM primitives: nn/GroupedQueryAttention for grouped-query attention (fewer K/V heads than Q heads), nn/SwiGLU for the feed-forward block, nn/rmsnorm for pre-norm, and torch/precompute-rope-freqs + torch/apply-rope for rotary position embeddings. The example shows both a regular forward pass and an incremental KV-cache forward pass where a new token is appended to a prefix.
atom as :kv-cache and slice pre-computed RoPE frequencies to the exact token positions being processed:
NanoChat
Source:examples/nanochat.clj
A compact, single-device Llama-3-style chat demo inspired by Karpathy’s NanoChat. It tokenizes a text corpus with jtokkit, trains a small Llama model (2 layers, 4 attention heads, 1 K/V head, 128-dimensional embeddings) using Adam with gradient clipping, saves a checkpoint, and supports interactive chat with streaming token generation.
The Llama model is built with defmodel nesting LlamaBlock records:
F/softmax probabilities, maintains a sliding context window, and optionally retains KV-cache tensors across steps using torch/retain!:
nanochat requires the-verdict.txt in the working directory for training. It falls back gracefully to untrained weights for the chat function if no checkpoint exists.Distributed CUDA Training
Source:examples/distributed_training.clj
A production-pattern multi-GPU training example demonstrating NCCL process groups, synchronous DistributedDataParallel, clorch.amp autocast with both :float16 (dynamic scaling) and :bfloat16, gradient accumulation with ddp/no-sync, and rank-zero checkpoint saving via dist/save-checkpoint!.
The train-worker function is the per-rank entrypoint. It is passed rank, world-size, process-group, and args by the launcher:
run-local!, passing device indices and training hyperparameters:
Bayesian Linear Regression MCMC
Source:examples/bayesian_linear_regression_mcmc.clj
Demonstrates probabilistic inference using clorch.distributions. It generates 120 synthetic observations from a known linear model, defines a log-posterior combining Normal priors on weights and bias with a Gaussian likelihood, and samples the posterior using random-walk Metropolis-Hastings. Multiple independent chains are run, with R-hat convergence diagnostics computed across chains.
log-alpha, and accepts or rejects based on a uniform draw:
examples/out/ as CSV and EDN.
Einsum eDSL
Source:examples/einsum_edsl.clj
Shows how to use clorch.einsum’s ein macro — a declarative Clojure eDSL for tensor contractions. Index variables are declared with declare, then used directly in ein expressions without string notation. The macro supports matrix-vector products, outer products, traces, and scalar-scaled contractions.