← PRA / report example

Tracel AI / Burn: engineering review, September 2026

tracel-ai/burn · main · work merged 1–30 September 2026 (UTC−4)

Previous report: August 2026.

Executive read

The August review established several lines to follow: migration toward Flex and CubeCL, model export through graph capture, quantized fine-tuning, and reliability of storage and training. September moved some unfinished integrations into main and deepened several of those lines. Two open outcomes did not advance: graph capture gained no export integration, and the deprecated NdArray and LibTorch backends were not removed.

Burn also spent September preparing version 0.22. A migration guide was added to main, followed by the fourth prerelease on 22 September. The project changed how backends are configured, separated several components, and corrected failures in model storage, numerical operations, and training.

Version 0.22.0 itself was not released by the end of the month. Version 0.21.0 from 7 May remained the latest stable release, so the changes below were not yet available to its users.

Key developments

1. Backend selection is now explicit

In August, Burn removed Candle, deprecated NdArray and LibTorch, and recommended Flex or CubeCL for new configurations. The August backend section described that as a migration rather than a completed removal. September moved from recommendation to explicit configuration.

A backend determines where Burn runs computations. Flex targets CPUs, while CubeCL underpins variants for CUDA, ROCm, WGPU, and other devices. Burn’s 0.22 prereleases had temporarily selected Flex when no backend was enabled. Now Device::default() requires an enabled backend and reports which Cargo feature to add. Builds that only capture a computation graph can still work without an execution backend.

Linear algebra and signal processing moved into the separate burn-linalg and burn-signal crates. They must be enabled explicitly, and their former API paths were removed without a deprecation period. Explicit backend selection is not a regression for stable-version users because backends were already optional in version 0.21.

Sources: #5722, #5520, #5572, #5720. Technical review.

2. The PyTorch reader became a separate component

August brought targeted fixes for PyTorch tensor strides and missing or truncated data. September turned that work into a broader component rather than another isolated loader fix. The earlier state is described in the August storage and loading section.

pytorch-reader reads .pt and .pth files so trained weights can be moved from PyTorch into Burn. In September it became independent of both Burn and PyTorch. The reader can now extract weights from files created with torch.save(model), find tensors inside lists and tuples, and accept archives whose CRC checks were disabled.

The feature remains a weight-transfer mechanism. It does not restore the Python class, model architecture, or forward method, so the model structure must still be defined in Burn. Version 0.22.0-pre.4 of the separate crate was published on crates.io on 22 September.

Sources: #5656, #5766, #5664, #5756, #5749, crates.io. Technical review.

3. A failed save no longer destroys the previous checkpoint

This closes a specific item from the August open considerations. Atomic SafeTensors writing was still WIP at the August cutoff. It entered main on 1 September and was then hardened during the month.

SafeTensors writing previously truncated the destination before all tensors had been read from the device. A failure partway through could leave a damaged file where the last working checkpoint had been. The new container is now assembled in a temporary file and replaces the old one only after completion.

The no-overwrite check also moved to publication time. On filesystems with hard-link support, this prevents another process from creating the destination between the check and publication. Atomic writing became the default burn-pack behavior on 28 September.

Sources: #5489, #5781, #5832. Technical review.

4. Numerical edge cases moved closer to PyTorch behavior

This is a new September theme, but it follows directly from the August backend direction. Making Flex and CubeCL the recommended paths also makes their agreement on edge cases more important.

Burn sometimes hid NaN, the value used to signal an invalid numerical result. Operations that select a maximum, discard negative values, or limit values to a range could replace it with an ordinary number. Incorrect handling of sign(NaN) could also produce a plausible finite gradient. September’s changes corrected these cases and made part of the behavior a common backend contract.

Other fixes covered a fully masked quiet_softmax, pooling output size with ceil_mode, and the direction of Tensor::roll. The roll correction changes the visible result of a public API, but it was not listed in the migration guide by the end of the month. The work is incomplete because tests still identify backend and data-type combinations with different behavior.

Sources: #5665, #5662, #5658, #5718, #5743, #5798, #5912. Technical review.

5. Training metrics and memory-saving modes returned to their stated behavior

August fixed when training decisions read their metrics, plus the aggregation of Dice and retrieval of ROUGE-L. September moved from timing to the formulas themselves. See the August metrics section.

Accuracy and top-k accuracy incorrectly weighted padded samples in the epoch average, and a fully padded batch could add NaN. Average precision differed from scikit-learn when model scores were tied. In the reference case, the expected value changed from 0.5918 to 0.5501. AUROC was rewritten without a pairwise table covering every example. These changes correct metric calculation, not model quality.

Two memory-saving training modes were also repaired. With gradient accumulation, an incomplete window at the end of an epoch was discarded and the learning-rate schedule advanced per batch instead of per parameter update. Gradient checkpointing, which saves memory by recomputing intermediate values, lost its selected strategy after the first step. Both modes now follow their stated behavior. A separate fix preserves fine-tuning adapters when a validation copy of the model is created, continuing the August QLoRA correctness work.

Sources: #5578, #5684, #5900, #5729, #5618, #5821. Metrics and training loop.

6. Operations preserve advantageous memory layouts

August’s measured performance work changed how one convolution gradient was computed. September addressed a different bottleneck: preserving useful memory layouts between operations. The measurements therefore describe separate optimizations and cannot be treated as one trend. See the August performance section.

A convolution can return a tensor in a memory order that suits the next computation, but later operations previously read it with a large stride and lost vectorization. Elementwise operations and type casts now follow the tensor’s physical order. A channels-last kernel was also added for transposed convolution. User code does not need to change.

In one PR author’s measurement on an NVIDIA L4 under CUDA, training-step time fell from 418.7 ms to 364.2 ms. In another measurement, the slowdown of a channels-last layout relative to a contiguous layout fell from 6.28x to 1.10x. The configurations and baselines differed, so the results cannot be combined or generalized to Burn as a whole.

Sources: #5625, #5620, #5796, #5760. Technical review.

7. The server releases resources after a remote session disconnects

August isolated a failed operation from independent work in the Fusion queue. September extended failure handling at a different layer: the lifetime of a remote connection and its server-side resources. The earlier behavior is described in the August execution-failure section.

Burn Remote runs tensor operations on another computer or accelerator. Previously, neither side checked the connection state. A disconnected client could wait indefinitely for a read, while the server left its session and tensors on the device. According to the fix description, several such sessions filled a card with 6 GB of memory and caused the next run to fail for lack of memory.

TCP keep-alive now detects a failure in about 30 seconds. Read errors and worker-thread panics pass through session cleanup. Separately, the Iroh server gained relay and port configuration plus token-based access. This reclaims resources after a failure, but it does not restore the interrupted session or its data.

Sources: #5825, #5828, #5878, #5910. Technical review.

Open considerations

The transition to 0.22 is incomplete. After the v0.22.0-pre.4 tag, CubeCL and Cubek briefly moved to published versions, then returned to Git revisions on 23 September. At month end, main was not configured for a stable release.

Deprecated backends remain part of the project. burn-ndarray and burn-tch, deprecated in the August backend transition, continued to be published and received new functionality. The public history does not yet show their removal.

Not every incompatibility is marked consistently. Ten September PR titles carried an explicit breaking-change marker. At least five more changed public APIs without one. In particular, Module::no_grad() now controls only gradients. Its former effect on Dropout and BatchNorm moved to freeze(). Such code can continue to compile while behaving differently.

Numerical parity across backends remains incomplete. Tests still restrict max pooling, clamp, and ReLU cases on some backends and data types. The graph capture work described in August as a foundation for model export received no substantive follow-up in September’s main.

CubeCL and Cubek changed frequently. Of 246 PRs, 36 changed only these dependency revisions, while another 18 combined a revision update with Burn’s own work. This shows the pace of integration, but it does not establish that external dependencies are constraining the project.

Coverage and limits

The window contains 246 PRs and no direct commits to main. This memo examines seven developments in detail. The technical review covers all 246 PRs exactly once and describes WIP separately.

Public Git history does not show private work, revenue, or adoption. Benchmark figures are reported by PR authors and were not independently reproduced. Tests were inspected but not run.

Privacy

Burn and Tracel AI are trademarks of their respective owners. PRA is not affiliated with Tracel AI.