> ## Documentation Index
> Fetch the complete documentation index at: https://antlobach-clorch-182a83cb.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Optimizers: SGD, Adam, AdamW, RMSprop, and Adagrad

> Clorch optimizer reference — SGD, Adam, AdamW, RMSprop, Adagrad constructors, key options, zero-grad/step lifecycle, and a complete train-step example.

The `clorch.optim` namespace provides thin Clojure wrappers over LibTorch's C++ optimizer implementations. Because the heavy lifting is done in native code, there is minimal JVM overhead per update step even for large models.

```clojure theme={null}
(require '[clorch.optim :as optim]
         '[clorch.nn :as nn])
```

All constructors accept a `TensorVector` of parameters as their first argument. Obtain this by calling `nn/parameters` on your model.

***

## SGD — Stochastic Gradient Descent

The classic first-order optimizer. With momentum it approximates gradient averaging, which dampens oscillations and speeds convergence.

```clojure theme={null}
(def opt (optim/sgd (nn/parameters model)
                    :lr           0.01
                    :momentum     0.9
                    :dampening    0
                    :weight-decay 0
                    :nesterov     false))
```

| Option | Default | Description |
| - | - | - |
| `:lr` | `0.01` | Learning rate |
| `:momentum` | `0` | Momentum factor |
| `:dampening` | `0` | Dampening for momentum |
| `:weight-decay` | `0` | L2 regularization coefficient |
| `:nesterov` | `false` | Use Nesterov momentum |

***

## Adam

Adaptive Moment Estimation. Maintains per-parameter first and second moment estimates, making it robust to sparse gradients and noisy loss surfaces.

```clojure theme={null}
(def opt (optim/adam (nn/parameters model)
                     :lr           0.001
                     :betas        [0.9 0.999]
                     :eps          1e-8
                     :weight-decay 0
                     :amsgrad      false))
```

| Option | Default | Description |
| - | - | - |
| `:lr` | `0.001` | Learning rate |
| `:betas` | `[0.9 0.999]` | Coefficients for moving averages of gradient and squared gradient |
| `:eps` | `1e-8` | Numerical stability constant added to denominator |
| `:weight-decay` | `0` | L2 regularization coefficient |
| `:amsgrad` | `false` | Use AMSGrad variant |

***

## AdamW

AdamW decouples weight decay from the gradient update, applying it directly to the parameters rather than adding it to the gradient. This is the recommended optimizer for transformer and LLM training.

```clojure theme={null}
(def opt (optim/adamw (nn/parameters model)
                      :lr           3e-4
                      :betas        [0.9 0.999]
                      :eps          1e-8
                      :weight-decay 0.01
                      :amsgrad      false))
```

| Option | Default | Description |
| - | - | - |
| `:lr` | `0.001` | Learning rate |
| `:betas` | `[0.9 0.999]` | Coefficients for moving averages of gradient and squared gradient |
| `:eps` | `1e-8` | Numerical stability constant added to denominator |
| `:weight-decay` | `0.01` | Weight decay coefficient (decoupled from gradient update) |
| `:amsgrad` | `false` | Use AMSGrad variant |

<Note>
  AdamW's default `:weight-decay` is `0.01` rather than `0` (unlike plain `adam`). This reflects its intended use as a regularizer rather than an optimizer modification. A common value for transformer fine-tuning is `0.1`.
</Note>

<Tip>
  For transformer and LLM training, use AdamW with `lr 3e-4`, `weight-decay 0.1`, and a cosine learning-rate schedule. This combination appears in most modern language model training recipes.
</Tip>

***

## RMSprop

Divides the learning rate by an exponentially decaying average of squared gradients. Originally proposed for non-stationary objectives and recurrent networks.

```clojure theme={null}
(optim/rmsprop (nn/parameters model)
               :lr           0.01
               :alpha        0.99
               :eps          1e-8
               :weight-decay 0
               :momentum     0
               :centered     false)
```

| Option | Default | Description |
| - | - | - |
| `:lr` | `0.01` | Learning rate |
| `:alpha` | `0.99` | Smoothing constant for squared gradient average |
| `:eps` | `1e-8` | Numerical stability constant |
| `:weight-decay` | `0` | L2 regularization |
| `:momentum` | `0` | Momentum factor |
| `:centered` | `false` | Normalize gradient by estimated variance when `true` |

***

## Adagrad

Accumulates all past squared gradients, giving a larger effective learning rate to infrequent parameters. Useful for sparse feature problems (NLP with bag-of-words representations).

```clojure theme={null}
(optim/adagrad (nn/parameters model)
               :lr                        0.01
               :lr-decay                  0
               :weight-decay              0
               :initial-accumulator-value 0
               :eps                       1e-10)
```

| Option | Default | Description |
| - | - | - |
| `:lr` | `0.01` | Learning rate |
| `:lr-decay` | `0` | Learning rate decay |
| `:weight-decay` | `0` | L2 regularization coefficient |
| `:initial-accumulator-value` | `0` | Starting value for the sum of squared gradients |
| `:eps` | `1e-10` | Numerical stability constant added to denominator |

***

## Lifecycle: zero-grad and step

Every optimizer follows the same two-call update cycle.

### Clearing Gradients

Gradient tensors accumulate by addition across calls to `autograd/backward`. Always zero them out at the start of each training iteration:

```clojure theme={null}
(optim/zero-grad opt)
```

### Applying the Update

After backpropagation, advance all parameters one step in the direction of the negative gradient:

```clojure theme={null}
(optim/step opt)
```

***

## Full Training Step Example

```clojure theme={null}
(require '[clorch.torch :as t]
         '[clorch.nn :as nn]
         '[clorch.nn.functional :as F]
         '[clorch.optim :as optim]
         '[clorch.autograd :as autograd])

(defn train-step [model optimizer batch]
  ;; Clear accumulated gradients
  (optim/zero-grad optimizer)
  (let [pred (nn/forward model (:data batch))
        loss (F/mse-loss pred (:target batch))]
    ;; Compute gradients via reverse-mode autodiff
    (autograd/backward loss)
    ;; Update parameters
    (optim/step optimizer)
    ;; Return loss as a plain JVM float
    (t/item-float loss)))
```

Call `train-step` from inside a `t/with-torch` scope so that intermediate tensors are released after each iteration:

```clojure theme={null}
(doseq [batch dataloader]
  (t/with-torch
    (let [loss-val (train-step model optimizer batch)]
      (println "loss:" loss-val))))
```

<Warning>
  Do **not** hold a reference to the loss tensor beyond the `with-torch` scope. Extract any values you need (e.g. with `t/item-float`) before the scope closes. See the [Memory Management](/concepts/memory) page for the complete canonical training loop.
</Warning>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.