Skip to main content
Clorch’s clorch.data namespace provides a composable data loading pipeline built around the IDataset protocol. A dataset knows its size and how to produce one item; a dataloader sequences items into batches, handles shuffling, and optionally parallelizes item loading across worker threads or subprocess workers. For distributed training, the DistributedSampler ensures each rank receives a disjoint, deterministic partition of the data.

The IDataset Protocol

Any Clojure value that implements IDataset can be passed to a dataloader.

Creating a Dataset with data/dataset

The data/dataset function builds a minimal dataset from two keyword arguments:
For process-worker support, also provide :process-spec:

Defining Datasets with data/defdataset

defdataset generates a named constructor and a record that implements IDataset. Fields in the binding vector become record slots, so the dataset state is immutable.

Tensor Datasets

For simple in-memory supervised learning, data/tensor-dataset creates a dataset from two tensors directly:

Dataloaders

data/dataloader wraps a dataset and produces lazy sequences of batches.

Dataloader Options

When :num-workers is zero, batches are built synchronously on the calling thread. This is fine for small datasets or when items are already tensor-backed in memory.

Worker Backends

Thread workers share the same JVM heap and can access in-memory datasets directly. They are simpler but do not isolate failures.

Distributed Sampling

In multi-rank training, every rank must receive a disjoint, deterministic subset of the dataset so the same item is never processed twice in the same epoch.

Creating a DistributedSampler

distributed-sampler also accepts the dataset itself instead of a plain integer — it will call get-size to determine the count.

Sampler Options

data/set-epoch!

Call set-epoch! at the start of every training epoch to advance the shuffle seed. Each rank applies the same permutation independently, which guarantees the partition remains disjoint.
If you omit set-epoch!, the sampler reuses the same shuffle every epoch. This causes every epoch to train on identical mini-batches, which can cause overfitting and poor convergence.

data/sample-indices

data/sample-indices returns a vector of integer indices for this rank and epoch. Pass it to partition-all to form batches:

Integrating with a Dataloader

Pass the sampler to data/dataloader directly. The loader reads indices from the sampler instead of shuffling internally.

Collation

The default collate function (data/default-collate) handles three cases:
  • Tensors: stacks a list of tensors along a new batch dimension using torch/stack.
  • Maps: recursively collates each key, producing a map of batched tensors.
  • Other values: wraps them in a Clojure vector.
To override collation, provide a :collate-fn to the dataloader:

Resource Cleanup

Tensors stacked by the default collate function are retained so they survive past the native-memory scope of the worker. If your training loop wraps each batch in t/with-torch, call data/cleanup-data! on the batch after you have extracted all JVM-scalar results: