Skip to main content
The clorch.optim namespace provides thin Clojure wrappers over LibTorch’s C++ optimizer implementations. Because the heavy lifting is done in native code, there is minimal JVM overhead per update step even for large models.
All constructors accept a TensorVector of parameters as their first argument. Obtain this by calling nn/parameters on your model.

SGD — Stochastic Gradient Descent

The classic first-order optimizer. With momentum it approximates gradient averaging, which dampens oscillations and speeds convergence.

Adam

Adaptive Moment Estimation. Maintains per-parameter first and second moment estimates, making it robust to sparse gradients and noisy loss surfaces.

AdamW

AdamW decouples weight decay from the gradient update, applying it directly to the parameters rather than adding it to the gradient. This is the recommended optimizer for transformer and LLM training.
AdamW’s default :weight-decay is 0.01 rather than 0 (unlike plain adam). This reflects its intended use as a regularizer rather than an optimizer modification. A common value for transformer fine-tuning is 0.1.
For transformer and LLM training, use AdamW with lr 3e-4, weight-decay 0.1, and a cosine learning-rate schedule. This combination appears in most modern language model training recipes.

RMSprop

Divides the learning rate by an exponentially decaying average of squared gradients. Originally proposed for non-stationary objectives and recurrent networks.

Adagrad

Accumulates all past squared gradients, giving a larger effective learning rate to infrequent parameters. Useful for sparse feature problems (NLP with bag-of-words representations).

Lifecycle: zero-grad and step

Every optimizer follows the same two-call update cycle.

Clearing Gradients

Gradient tensors accumulate by addition across calls to autograd/backward. Always zero them out at the start of each training iteration:

Applying the Update

After backpropagation, advance all parameters one step in the direction of the negative gradient:

Full Training Step Example

Call train-step from inside a t/with-torch scope so that intermediate tensors are released after each iteration:
Do not hold a reference to the loss tensor beyond the with-torch scope. Extract any values you need (e.g. with t/item-float) before the scope closes. See the Memory Management page for the complete canonical training loop.