> ## Documentation Index
> Fetch the complete documentation index at: https://antlobach-clorch-182a83cb.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Automatic Differentiation with Clorch's Autograd Engine

> How Clorch wraps LibTorch's native reverse-mode AD engine to compute gradients, control graph tracking, and wire backward passes into a training loop.

Automatic differentiation is the mechanism that makes gradient-based learning practical. Clorch's `clorch.autograd` namespace is a thin, idiomatic wrapper around LibTorch's native reverse-mode autograd engine — the same engine that powers PyTorch. Every tensor operation that touches a gradient-tracked tensor is recorded in a dynamic computation graph. Calling `autograd/backward` on any scalar node in that graph walks the graph in reverse and accumulates gradients at all leaf tensors. The entire process is transparent to the caller; you write forward math as ordinary function calls and gradients appear automatically.

```clojure theme={null}
(require '[clorch.torch    :as t]
         '[clorch.autograd :as autograd])
```

***

## Enabling Gradient Tracking

A tensor only participates in the computation graph when created with `{:requires-grad true}`. Tensors created without this flag are treated as constants — they pass values through without recording any operations.

```clojure theme={null}
;; Leaf tensor — gradients will accumulate here
(def x (t/tensor [2.0] {:requires-grad true}))

;; Intermediate node — graph edge is created automatically
(def y (t/pow x 2))  ;; y = x²
```

You can also enable tracking on an existing tensor retroactively:

```clojure theme={null}
(autograd/set-requires-grad x true)
(autograd/set-requires-grad x false) ;; disable again
```

<Note>
  Only floating-point tensors support `requires-grad`. Setting it on an integer tensor throws a LibTorch error.
</Note>

***

## Computing Gradients

### `backward`

Call `autograd/backward` on any scalar-valued tensor to trigger reverse-mode accumulation. After the call, the accumulated gradient is available at every tracked leaf via `autograd/grad`.

```clojure theme={null}
;; y = x², dy/dx = 2x
(def x (t/tensor [2.0] {:requires-grad true}))
(def y (t/pow x 2))

(autograd/backward y)
(autograd/grad x)  ;; → [4.0]  (2 × 2.0)
```

### `grad`

`autograd/grad` simply reads the `.grad` field of the underlying tensor. It returns a tensor (not a scalar), so use `t/item-float` when you need a JVM number:

```clojure theme={null}
(t/item-float (autograd/grad x))  ;; 4.0
```

### Worked examples

<Tabs>
  <Tab title="x²">
    ```clojure theme={null}
    ;; f(x) = x², f'(x) = 2x
    (def x (t/tensor [3.0] {:requires-grad true}))
    (def y (t/pow x 2))

    (autograd/backward y)
    (autograd/grad x)
    ;; → [6.0]   (2 × 3.0)
    ```
  </Tab>

  <Tab title="x³">
    ```clojure theme={null}
    ;; f(x) = x³, f'(x) = 3x²
    (def x (t/tensor [2.0] {:requires-grad true}))
    (def y (t/pow x 3))

    (autograd/backward y)
    (autograd/grad x)
    ;; → [12.0]  (3 × 2.0²)
    ```
  </Tab>

  <Tab title="Composition">
    ```clojure theme={null}
    ;; f(x) = sin(x²), f'(x) = 2x·cos(x²)
    (def x  (t/tensor [1.0] {:requires-grad true}))
    (def y  (t/sin (t/pow x 2)))

    (autograd/backward y)
    (autograd/grad x)
    ;; ≈ [1.0806]
    ```
  </Tab>
</Tabs>

***

## Detaching from the Graph

`autograd/detach` returns a new tensor that shares storage with the original but has no gradient history. Use it when you need the current numeric value of a tensor for a side-effect (logging, a metric, a threshold check) without polluting the computation graph.

```clojure theme={null}
(def x (t/tensor [2.0] {:requires-grad true}))
(def y (t/pow x 2))

;; Detach before computing statistics that must not be differentiable
(def y-detached (autograd/detach y))

;; y-detached has the same values as y
(t/item-float y-detached)  ;; 4.0

;; But gradients do not flow through y-detached
```

<Warning>
  `detach` creates a view of the original storage. Mutating the detached tensor in-place will affect the original. Prefer `(t/clone (autograd/detach y))` if you need an independent copy.
</Warning>

***

## Disabling Gradient Tracking with `no-grad`

The `autograd/no-grad` macro wraps a block of code in a LibTorch `NoGradGuard`, which prevents any tensor operation inside the block from being recorded in the computation graph. This is essential for:

* **Inference / validation** — avoids allocating intermediate activation nodes, cutting memory by roughly half for a typical forward pass.
* **Metric and loss logging** — you want numbers, not graph nodes.
* **Parameter updates** — weight tensors should not track optimizer arithmetic.

```clojure theme={null}
(autograd/no-grad
  (let [pred (nn/forward model x)]
    (calculate-accuracy pred targets)))
```

Nested `no-grad` blocks are safe. The guard is re-entrant and restored correctly when the block exits, even on exception.

<Tip>
  Every inference call in production should be wrapped in `no-grad`. Forgetting it can cause memory to grow unboundedly because the graph accumulates across calls.
</Tip>

***

## Manual Gradient Control

`autograd/set-requires-grad` gives you fine-grained control over which tensors are tracked. The most common use case is freezing part of a model for transfer learning:

```clojure theme={null}
;; Freeze every parameter in a sub-module
(doseq [p (nn/parameters encoder)]
  (autograd/set-requires-grad p false))

;; Later, unfreeze for fine-tuning
(doseq [p (nn/parameters encoder)]
  (autograd/set-requires-grad p true))
```

***

## Training Loop Integration

A canonical gradient update has three steps: zero accumulated gradients from the previous iteration, run the forward pass and call `backward`, then apply the optimizer step. All three belong in a single `with-torch` scope per batch so that intermediate tensors are released deterministically.

```clojure theme={null}
(require '[clorch.torch          :as t]
         '[clorch.nn             :as nn]
         '[clorch.nn.functional  :as F]
         '[clorch.autograd       :as autograd]
         '[clorch.optim          :as optim])

(let [model     (create-model)
      optimizer (optim/adam (nn/parameters model))]
  (doseq [{:keys [data target]} dataloader]
    (let [loss-value
          (t/with-torch
            ;; 1. Zero gradients from the previous batch
            (optim/zero-grad optimizer)
            ;; 2. Forward pass + backward
            (let [prediction (nn/forward model data)
                  loss       (F/cross-entropy prediction target)]
              (autograd/backward loss)
              ;; 3. Parameter update
              (optim/step optimizer)
              ;; Return a JVM scalar so no tensor escapes the scope
              (t/item-float loss)))]
      (println "Loss:" loss-value))))
```

<Steps>
  <Step title="Zero gradients">
    `optim/zero-grad` clears `.grad` on every tracked parameter. Without this step, gradients accumulate across batches.
  </Step>

  <Step title="Backward pass">
    `autograd/backward` traverses the computation graph from the scalar loss to every leaf, accumulating `∂loss/∂param` at each parameter tensor.
  </Step>

  <Step title="Optimizer step">
    `optim/step` reads the accumulated gradients and updates each parameter according to the optimizer rule (SGD, Adam, AdamW, etc.).
  </Step>

  <Step title="Return a scalar">
    The `with-torch` scope sees `(t/item-float loss)` — a plain JVM `Float` — as its final result. No tensors are retained across iterations, keeping native memory bounded. See the [Memory Management](/concepts/memory) page for a complete explanation of why this matters.
  </Step>
</Steps>

***

## Autograd API Reference

<CardGroup cols={2}>
  <Card title="autograd/backward" icon="arrow-right-arrow-left">
    Computes reverse-mode gradients for all tracked leaves reachable from the scalar tensor argument. Modifies `.grad` fields in place.
  </Card>

  <Card title="autograd/grad" icon="magnifying-glass">
    Returns the accumulated gradient tensor at a leaf. Returns `nil` before `backward` is called or if the tensor has no gradient.
  </Card>

  <Card title="autograd/detach" icon="link-slash">
    Returns a view of the tensor with no graph history. Safe for metrics and logging. Does not copy storage.
  </Card>

  <Card title="autograd/no-grad" icon="ban">
    Macro that disables graph recording for the duration of its body. Mandatory for inference loops and validation steps.
  </Card>

  <Card title="autograd/set-requires-grad" icon="toggle-on">
    Enables or disables gradient accumulation on an existing tensor. Use to freeze/unfreeze model parameters.
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.