> ## Documentation Index
> Fetch the complete documentation index at: https://antlobach-clorch-182a83cb.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Clorch Activation Functions: Module and Functional

> All Clorch activations in nn/ and F/ forms: ReLU, GeLU, SiLU, ELU, SELU, softmax, shrinkage, Mish, Hardswish, and more with output range comparison table.

Activation functions introduce the non-linearity that lets neural networks approximate complex functions. Clorch provides every standard activation in two interchangeable forms: a stateful module form under `clorch.nn` (for use inside `defmodel` or `nn/sequential`) and a purely functional form under `clorch.nn.functional` (for direct call in a forward body). Both delegate to the same LibTorch C++ kernels.

```clojure theme={null}
(require '[clorch.nn :as nn]
         '[clorch.nn.functional :as F])
```

***

## Standard Activations

| Activation | `clorch.nn` module | `clorch.nn.functional` |
| - | - | - |
| ReLU | `(nn/relu)` | `(F/relu x)` |
| Sigmoid | `(nn/sigmoid)` | `(F/sigmoid x)` |
| Tanh | `(nn/tanh)` | `(F/tanh x)` |
| GeLU | `(nn/gelu)` | `(F/gelu x)` |
| SiLU (Swish) | `(nn/silu)` | `(F/silu x)` |
| Softmax | `(nn/softmax dim)` | `(F/softmax x dim)` |

<Tabs>
  <Tab title="ReLU">
    The most widely-used activation for hidden layers. Computationally cheap and avoids the vanishing gradient problem for positive activations.

    ```clojure theme={null}
    ;; Module form (place in sequential or defmodel)
    (nn/relu)

    ;; Functional form (call directly in forward)
    (F/relu x)
    ```
  </Tab>

  <Tab title="GeLU">
    Gaussian Error Linear Unit. Preferred in transformer architectures (BERT, GPT) because it provides a smooth approximation of ReLU with non-zero gradients for slightly negative inputs.

    ```clojure theme={null}
    (nn/gelu)
    (F/gelu x)
    ```
  </Tab>

  <Tab title="SiLU / Swish">
    Sigmoid Linear Unit (also called Swish). Used in modern vision and language models. Closely related to GeLU but computed as `x * sigmoid(x)`.

    ```clojure theme={null}
    (nn/silu)
    (F/silu x)
    ```
  </Tab>

  <Tab title="Softmax">
    Converts a vector of logits into a probability distribution. Always pass the class dimension explicitly.

    ```clojure theme={null}
    ;; Module: dim is baked in at construction time
    (nn/softmax -1)

    ;; Functional: dim supplied at call time
    (F/softmax logits -1)
    ```
  </Tab>
</Tabs>

***

## Enhanced and Specialized Activations

### LeakyReLU

Passes negative inputs through with a small, fixed slope instead of zeroing them. Useful when dead neurons are a concern.

<Tabs>
  <Tab title="Module">
    ```clojure theme={null}
    ;; Default negative slope: 0.01
    (nn/leaky-relu)

    ;; Custom slope
    (nn/leaky-relu 0.1)
    ```
  </Tab>

  <Tab title="Functional">
    ```clojure theme={null}
    (F/leaky-relu x 0.1)
    ```
  </Tab>
</Tabs>

### PReLU — Parametric ReLU

The negative slope is a **learnable parameter** rather than a fixed constant. Because `PReLURecord` stores the slope as an `nn/parameter`, it participates in `nn/parameters` and gradient descent automatically.

<Tabs>
  <Tab title="Module">
    ```clojure theme={null}
    ;; Single shared slope (default)
    (nn/prelu)

    ;; One slope per channel, initialised to 0.25
    (nn/prelu {:num-parameters 64 :init 0.25})
    ```
  </Tab>

  <Tab title="Functional">
    ```clojure theme={null}
    ;; You manage the weight tensor yourself
    (F/prelu x weight)
    ```
  </Tab>
</Tabs>

### ELU — Exponential Linear Unit

Smooth for negative values (`alpha * (exp(x) - 1)`), which can help learning speed compared to ReLU.

<Tabs>
  <Tab title="Module">
    ```clojure theme={null}
    (nn/elu)         ;; alpha = 1.0
    (nn/elu 0.5)     ;; alpha = 0.5
    ```
  </Tab>

  <Tab title="Functional">
    ```clojure theme={null}
    (F/elu x 1.0)
    ```
  </Tab>
</Tabs>

### SELU — Scaled Exponential Linear Unit

Self-normalizing variant of ELU with fixed scale and alpha parameters. Designed for use with `alpha-dropout`.

```clojure theme={null}
(nn/selu)
(F/selu x)
```

### CELU — Continuously Differentiable ELU

A variant of ELU that is continuously differentiable everywhere. The `alpha` parameter controls the saturation value for negative inputs.

```clojure theme={null}
(nn/celu)           ;; alpha = 1.0
(nn/celu 0.5)
(F/celu x 1.0)
```

### GLU — Gated Linear Unit

Splits the input in half along `dim` and uses one half as a sigmoid gate on the other. Foundation of modern gated feed-forward blocks (SwiGLU, GeGLU).

<Tabs>
  <Tab title="Module">
    ```clojure theme={null}
    (nn/glu)     ;; default dim = -1
    (nn/glu 1)   ;; gate along dimension 1
    ```
  </Tab>

  <Tab title="Functional">
    ```clojure theme={null}
    (F/glu x -1)
    ```
  </Tab>
</Tabs>

### Softplus

A smooth approximation of ReLU defined as `log(1 + exp(beta * x)) / beta`. The `threshold` parameter switches to a linear function for large values for numerical stability.

```clojure theme={null}
(nn/softplus)              ;; beta = 1.0, threshold = 20.0
(nn/softplus 2.0 20.0)     ;; custom beta and threshold
(F/softplus x)
(F/softplus x 2.0 20.0)
```

### Log-Softmax and Softmin

```clojure theme={null}
;; Log of the softmax — numerically more stable than log(softmax(x))
(nn/log-softmax -1)         ;; dim = -1
(F/log-softmax x -1)

;; Softmax of negated input
(nn/softmin -1)
(F/softmin x -1)
```

### Log-Sigmoid and Softsign

```clojure theme={null}
;; Logarithm of the sigmoid function
(nn/log-sigmoid)
(F/log-sigmoid x)

;; x / (1 + |x|), smooth bounded activation
(nn/softsign)
(F/softsign x)
```

### ReLU6

ReLU clamped to a maximum value of 6. Used in MobileNet-style architectures.

```clojure theme={null}
(nn/relu6)
(F/relu6 x)
```

### RReLU — Randomized Leaky ReLU

Uses a random slope drawn from a uniform distribution during training and the midpoint during evaluation.

```clojure theme={null}
(nn/rrelu :lower 0.125 :upper 0.333)
(F/rrelu x :lower 0.125 :upper 0.333)
```

### Hardtanh

Clamps values to `[min_val, max_val]` with unit slope in between. Defaults to `[-1, 1]`.

```clojure theme={null}
(nn/hardtanh)           ;; [-1.0, 1.0]
(nn/hardtanh -2.0 2.0)
(F/hardtanh x -1.0 1.0)
```

### Threshold

Sets all values below `threshold-val` to `value`.

```clojure theme={null}
(nn/threshold 0.0 -1.0)         ;; values < 0 become -1
(F/threshold x 0.0 -1.0)
```

***

## Shrinkage Functions

Shrinkage activations set small-magnitude values to zero, encouraging sparse representations.

<Tabs>
  <Tab title="Hardshrink">
    Sets values in `(-λ, λ)` to zero; passes everything else through unchanged.

    ```clojure theme={null}
    (nn/hardshrink)       ;; λ = 0.5
    (nn/hardshrink 0.3)
    (F/hardshrink x 0.5)
    ```
  </Tab>

  <Tab title="Softshrink">
    Shifts values toward zero by `λ` rather than hard-zeroing them.

    ```clojure theme={null}
    (nn/softshrink)       ;; λ = 0.5
    (nn/softshrink 0.3)
    (F/softshrink x 0.5)
    ```
  </Tab>

  <Tab title="Tanhshrink">
    Returns `x - tanh(x)`. No threshold parameter.

    ```clojure theme={null}
    (nn/tanhshrink)
    (F/tanhshrink x)
    ```
  </Tab>
</Tabs>

***

## Modern Activations

<Tabs>
  <Tab title="Mish">
    `x * tanh(softplus(x))`. Often outperforms ReLU and Swish in image models without any hyperparameter tuning.

    ```clojure theme={null}
    (nn/mish)
    (F/mish x)
    ```
  </Tab>

  <Tab title="Hardswish">
    Piecewise-linear approximation of Swish. Computationally cheaper and hardware-friendly.

    ```clojure theme={null}
    (nn/hardswish)
    (F/hardswish x)
    ```
  </Tab>

  <Tab title="Hardsigmoid">
    Piecewise-linear approximation of Sigmoid. Maps to `[0, 1]` range with zero gradient outside `(-3, 3)`.

    ```clojure theme={null}
    (nn/hardsigmoid)
    (F/hardsigmoid x)
    ```
  </Tab>
</Tabs>

***

## Comparison Table

| Activation | Output Range | Primary Use Case |
| - | - | - |
| `ReLU` | \[0, ∞) | Default hidden-layer activation, CNNs |
| `Sigmoid` | (0, 1) | Binary classification output, gating |
| `Tanh` | (−1, 1) | Centered output, RNN hidden states |
| `Softmax` | (0, 1) per class | Multi-class classification output |
| `GeLU` | (−0.17, ∞) | Transformer encoder/decoder layers |
| `SiLU` | (−0.28, ∞) | Modern vision and language models |
| `LeakyReLU` | (−∞, ∞) | GANs, models prone to dead neurons |
| `ELU` | (−α, ∞) | Faster convergence vs ReLU in deep nets |
| `SELU` | (−λα, ∞) | Self-normalizing networks |
| `GLU` | (−∞, ∞) | Gated feed-forward blocks |
| `Softplus` | (0, ∞) | Smooth approximation of ReLU |
| `Mish` | (−0.31, ∞) | Image classification, drop-in ReLU swap |
| `Hardswish` | \[0, ∞) | Mobile / edge inference (efficient) |
| `Hardsigmoid` | \[0, 1] | Mobile / edge inference (efficient) |
| `Hardtanh` | \[min, max] | Quantization-aware training |

<Info>
  The `nn/silu` and `nn/hardswish` module constructors return Clojure function wrappers (not native `Module` objects) because LibTorch does not expose separate module classes for these. They are fully compatible with `nn/sequential` and `defmodel` forward bodies.
</Info>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.