clorch.optim namespace provides thin Clojure wrappers over LibTorch’s C++ optimizer implementations. Because the heavy lifting is done in native code, there is minimal JVM overhead per update step even for large models.
TensorVector of parameters as their first argument. Obtain this by calling nn/parameters on your model.
SGD — Stochastic Gradient Descent
The classic first-order optimizer. With momentum it approximates gradient averaging, which dampens oscillations and speeds convergence.Adam
Adaptive Moment Estimation. Maintains per-parameter first and second moment estimates, making it robust to sparse gradients and noisy loss surfaces.AdamW
AdamW decouples weight decay from the gradient update, applying it directly to the parameters rather than adding it to the gradient. This is the recommended optimizer for transformer and LLM training.AdamW’s default
:weight-decay is 0.01 rather than 0 (unlike plain adam). This reflects its intended use as a regularizer rather than an optimizer modification. A common value for transformer fine-tuning is 0.1.RMSprop
Divides the learning rate by an exponentially decaying average of squared gradients. Originally proposed for non-stationary objectives and recurrent networks.Adagrad
Accumulates all past squared gradients, giving a larger effective learning rate to infrequent parameters. Useful for sparse feature problems (NLP with bag-of-words representations).Lifecycle: zero-grad and step
Every optimizer follows the same two-call update cycle.Clearing Gradients
Gradient tensors accumulate by addition across calls toautograd/backward. Always zero them out at the start of each training iteration:
Applying the Update
After backpropagation, advance all parameters one step in the direction of the negative gradient:Full Training Step Example
train-step from inside a t/with-torch scope so that intermediate tensors are released after each iteration: