PORTING_STATUS.md). These are separate measurements and should not be conflated.
The 100% figure applies only to the 40 cross-language numerical scenarios exercised by
tests_comparison/compare_torch.py. Every currently paired Clorch/Python scenario passes within its configured numerical tolerance. It does not mean Clorch implements the full PyTorch API surface. A version-pinned breadth percentage will be published after generating a PyTorch 2.10 symbol inventory.Current Parity Summary
The release suite verifies CPU behavior and a single-GPU CUDA path, including CUDA discovery, NCCL world size one, DDP backward, AMP overflow handling, fused scaled-dot-product attention, checkpoints, worker failures, and process cleanup. Multi-rank validation requires a host with at least two visible NVIDIA GPUs.
Capability Comparison
This table maps each major PyTorch area to its current status in Clorch and the remaining work needed to reach full parity.LLM-Relevant Architecture
Clorch ships purpose-built primitives for large language model construction. The table below lists each capability, its implementation status, and the corresponding Clorch surface or example file.Roadmap to 100%
The path to full, version-pinned parity with PyTorch 2.10 requires both closing functional gaps and building the measurement infrastructure to prove completeness. The eight steps below are in dependency order.Step 1 — Freeze the target
Step 1 — Freeze the target
Select an exact upstream PyTorch release and generate a machine-readable inventory of its supported public Python and C++ APIs. Without a pinned inventory, no breadth percentage is meaningful.
Step 2 — Turn inventory into contracts
Step 2 — Turn inventory into contracts
Map every upstream symbol to one of: implemented, partial, intentionally different, or missing. Attach behavior tests to every implemented mapping so regressions surface automatically.
Step 3 — Close core tensor gaps
Step 3 — Close core tensor gaps
Finish dense overloads and edge cases, then sparse, nested, meta, quantized, and named-tensor behavior. This is the largest volume of work in the roadmap.
Step 4 — Complete autograd and modules
Step 4 — Complete autograd and modules
Add custom autograd functions, grad checking, hooks, inference controls, native transformer/container modules, and remaining functional/loss APIs.
Step 5 — Complete optimization and data
Step 5 — Complete optimization and data
Add public optimizer state, schedulers, remaining algorithms, iterable data loading, DataPipes, and worker/pinning parity.
Step 6 — Complete accelerator support
Step 6 — Complete accelerator support
Add CUDA streams, events, graphs, allocator controls, Gloo, FSDP, tensor parallelism, elastic execution, and broader multi-GPU conformance tests.
Step 7 — Complete export and ecosystem surfaces
Step 7 — Complete export and ecosystem surfaces
Add compile/export/ONNX/package behavior and explicitly scope domain libraries (TorchVision, TorchAudio, TorchText, TorchData).
Step 8 — Prove the result
Step 8 — Prove the result
Run generated cross-language conformance tests across Linux, macOS, Windows, CPU, CUDA, supported dtypes, and error/edge-case behavior. Publish the version-pinned breadth percentage.