Skip to main content
Activation functions introduce the non-linearity that lets neural networks approximate complex functions. Clorch provides every standard activation in two interchangeable forms: a stateful module form under clorch.nn (for use inside defmodel or nn/sequential) and a purely functional form under clorch.nn.functional (for direct call in a forward body). Both delegate to the same LibTorch C++ kernels.

Standard Activations

The most widely-used activation for hidden layers. Computationally cheap and avoids the vanishing gradient problem for positive activations.

Enhanced and Specialized Activations

LeakyReLU

Passes negative inputs through with a small, fixed slope instead of zeroing them. Useful when dead neurons are a concern.

PReLU — Parametric ReLU

The negative slope is a learnable parameter rather than a fixed constant. Because PReLURecord stores the slope as an nn/parameter, it participates in nn/parameters and gradient descent automatically.

ELU — Exponential Linear Unit

Smooth for negative values (alpha * (exp(x) - 1)), which can help learning speed compared to ReLU.

SELU — Scaled Exponential Linear Unit

Self-normalizing variant of ELU with fixed scale and alpha parameters. Designed for use with alpha-dropout.

CELU — Continuously Differentiable ELU

A variant of ELU that is continuously differentiable everywhere. The alpha parameter controls the saturation value for negative inputs.

GLU — Gated Linear Unit

Splits the input in half along dim and uses one half as a sigmoid gate on the other. Foundation of modern gated feed-forward blocks (SwiGLU, GeGLU).

Softplus

A smooth approximation of ReLU defined as log(1 + exp(beta * x)) / beta. The threshold parameter switches to a linear function for large values for numerical stability.

Log-Softmax and Softmin

Log-Sigmoid and Softsign

ReLU6

ReLU clamped to a maximum value of 6. Used in MobileNet-style architectures.

RReLU — Randomized Leaky ReLU

Uses a random slope drawn from a uniform distribution during training and the midpoint during evaluation.

Hardtanh

Clamps values to [min_val, max_val] with unit slope in between. Defaults to [-1, 1].

Threshold

Sets all values below threshold-val to value.

Shrinkage Functions

Shrinkage activations set small-magnitude values to zero, encouraging sparse representations.
Sets values in (-λ, λ) to zero; passes everything else through unchanged.

Modern Activations

x * tanh(softplus(x)). Often outperforms ReLU and Swish in image models without any hyperparameter tuning.

Comparison Table

The nn/silu and nn/hardswish module constructors return Clojure function wrappers (not native Module objects) because LibTorch does not expose separate module classes for these. They are fully compatible with nn/sequential and defmodel forward bodies.