Skip to main content
The initial values assigned to a network’s weights have a profound effect on whether training converges, how quickly it does so, and how deep the network can be. Poorly initialised weights lead to either vanishing gradients (weights shrink toward zero) or exploding gradients (signals amplify out of control). The clorch.nn.init namespace provides a set of in-place initializers that modify a tensor’s values directly — wrapping the corresponding LibTorch C++ routines.
All init/ functions modify the tensor in place and return the same tensor. They run inside a no-grad scope so the initialization operation itself is not recorded in the autograd graph.

Basic Initializers

These functions are useful for quick experiments, zeroing bias tensors, or setting a known baseline before applying a more sophisticated scheme.

Xavier (Glorot) Initialization

Xavier initialization scales the random weights based on the number of input and output connections (fan-in and fan-out) so that the variance of activations and gradients stays roughly constant across layers. It is designed for layers followed by symmetric activations such as tanh or sigmoid.

Xavier Uniform

Draws weights from a uniform distribution bounded by ± sqrt(6 / (fan_in + fan_out)) scaled by an optional gain.

Xavier Normal

Draws weights from a normal distribution with standard deviation sqrt(2 / (fan_in + fan_out)) scaled by gain.
Use Xavier initialization when your network uses tanh, sigmoid, or softmax activations. Avoid it with ReLU — use Kaiming instead, because ReLU zeroes half the activations and changes the effective fan.

Kaiming (He) Initialization

Kaiming initialization accounts for the fact that ReLU-like activations zero out half their inputs. It scales weights so that the variance is preserved through the forward pass (:fan-in mode) or the backward pass (:fan-out mode).

Kaiming Uniform

Kaiming Normal

Kaiming Options

The default :non-linearity in clorch.nn.init is :leaky-relu (matching PyTorch’s default). Always pass :non-linearity :relu explicitly when initializing layers before a standard ReLU activation.

Usage Pattern Inside defmodel

The recommended place to initialize custom weights is immediately after constructing the model, operating directly on the tensor fields exposed by the record.
For a model with multiple layers, initialize each layer in sequence:
Bias tensors are almost always initialised to zero regardless of the weight scheme. Kaiming and Xavier initializers operate only on the weight matrix and have no meaningful interpretation for a 1-D bias vector.

Choosing the Right Initializer