← PRA / report example

Tracel AI / Burn: technical review, September 2026

Companion to the management memo. The previous month is covered by the August technical review.

Scope and sources

The repository reviewed is tracel-ai/burn, branch main, window [2026-09-01T00:00:00−04:00, 2026-10-01T00:00:00−04:00). The UTC−4 zone is used, as in the August review. Date of review: 1 October 2026. Snapshot of main: commit 91d892a4ad1e928866b1950b3f99e8a4061191e2 of 30 September, 16:11:23 −04:00.

The set was determined by the committer date of the squash commit in the first-parent history of main. The window contains 246 PRs and no direct commits. Counting by UTC would additionally pull in #5510, which belongs to August in UTC−4, so it is not counted here.

The review covers all work merged during the period. The memo’s themes are examined in detail; every other change is expected to appear as a verified entry in the full-coverage section. Tests were read from the diff and not run independently, so there are no claims about them passing. Code references are pinned to revision 91d892a4. The behaviour of a later main may differ.

Four crates appeared during the period: burn-einsum, burn-linalg, burn-signal and pytorch-reader. The number of directories in crates/ grew from 34 on 31 August to 38 on 30 September.

1. The backend is chosen explicitly, and operations have moved into separate crates

Specialised operations and PyTorch reading were separated from the former modules

Linear algebra was moved out of burn-tensor into the new burn-linalg crate in #5572, squash commit 47ceefde21. The former crates/burn-tensor/src/tensor/linalg/ directory, present in the 31 August snapshot, was recast as the LinalgOps backend extension carrying the #[backend_extension] attribute and moved into crates/burn-linalg/. In the 30 September snapshot the old directory is gone. The paths burn_tensor::linalg and burn_core::tensor::linalg were removed with no deprecation period. The compatible path burn::linalg is available only with the linalg feature enabled, and that feature is not part of default for the facade crate in crates/burn/Cargo.toml. A host fallback through svd_host::svd_host_data was kept for LinalgOps::svd, so a backend without native SVD moves data to the CPU.

Signal processing was moved out to burn-signal in the same manner in #5720, squash commit 1c337e0d80. The new crates/burn-signal/ defines SignalOps as a backend extension. rfft and irfft are implemented as custom operations, and remote and capture require register_fft_ops. The old burn_tensor::signal path was removed. The signal feature in crates/burn/Cargo.toml is likewise not enabled by default. The same PR removed the stub for NdArray, since that backend provides no FFT.

Reading PyTorch files was extracted from burn-store into the independent pytorch-reader crate in #5656, squash commit 0fdac31a3d. The new crates/pytorch-reader/ does not depend on Burn. It uses its own DType and Tensor with lazy byte reading. Its direct dependencies are limited to byteorder, crc32fast, num-traits, serde, tar, thiserror and zip. burn-store re-exports the reader as burn_store::pytorch_reader and provides the bridge::from_pytorch bridge in crates/burn-store/. The former public path burn_store::nested was removed.

burn-einsum is of a different nature. #5583, squash commit 10b74fbb3d, does not move an existing module out. It adds new equation parsing, tensor-contraction planning, a procedural implementation in crates/burn-derive/, an API in crates/burn-tensor/ and a separate crates/burn-einsum/. The einsum dependency in crates/burn-tensor/Cargo.toml is neither optional nor behind a feature gate, so it is always compiled. The follow-up #5727, squash commit da2209a9de, changed CI only. Publishing of burn-einsum was added to .github/workflows/publish.yml, and the crate was included in the NO_STD_CRATES set.

All four new crates are registered in .github/workflows/publish.yml as separate published artifacts.

Extensions became part of the backend routing mechanism

The foundation for this new split of operations landed in #5520, squash commit 0e2dc6b631. burn-backend-extension was split into catalog, derive, dispatch, extension, ir and routing in crates/burn-backend-extension/. The former large set of macros in crates/burn-dispatch/src/macros.rs was replaced by generation through extensions. It is this #[backend_extension] mechanism that burn-linalg and burn-signal then used.

After the rebuild, promotion of the autodiff context had to be restored. That was done in #5551, squash commit cf1b6d2a60. The test file crates/burn/tests/autodiff_context.rs checks that calling an operation through the new extension mechanism preserves and correctly promotes the autodiff context. In the 30 September snapshot the file no longer exists at that path: #5647 moved it to crates/burn-backend-tests/tests/autodiff/context.rs. The link above points at the revision in which the test was added.

In #5570, squash commit 54d36fa936, burn-dispatch stopped depending on the facade crates of the CubeCL runtimes. The change is fixed in crates/burn-dispatch/Cargo.toml.

In #5694, squash commit 8116dff453, the new burn-linalg stopped pulling in cubecl-backend through its default features. The dependency boundary is set in crates/burn-linalg/Cargo.toml.

The execution backend is no longer chosen implicitly

#5722, squash commit 596d9acfbd, made burn-flex an optional dependency and replaced the default_backend configuration with backend_enabled in crates/burn-dispatch/build.rs. For a configuration with no execution backend, the uninhabited type NoBackend was introduced. Calling DispatchDevice::default() in such a build panics with an exact message:

No execution backend is enabled. Enable a Burn backend feature such as `flex`, `wgpu`, or `cuda`.
To record a graph without executing it, enable `capture` and use Device::capture()

This change does not rule out builds intended only for recording a graph. For those, the combination of the capture feature and Device::capture() remains.

Implicit Flex existed only in the 0.22 pre-releases, from pre.1 through pre.3. In stable v0.21.0, burn-dispatch and all backends were already optional. So #5722 breaks the configuration of users of the 0.22 pre-releases, but is not a regression relative to stable 0.21.

The module API separated gradient control from layer mode

#5721, squash commit 8f08fdad05, removed the AutodiffModule trait entirely. The valid() method moved into Module, and AutodiffModule::from_inner was removed. The train() method was already in Module; the change lifted its where Self: AutodiffModule bound and its implementation through from_inner. A single Rust type could represent both the training and the validation model before this PR as well. The substantive new boundary is a different one: a Module bound alone is no longer enough to guarantee that autodiff is enabled. The main implementation is in crates/burn-core/src/module/base.rs.

#5557, squash commit f7f08fd57c, removed the public Tensor::no_grad() and replaced it with Tensor::without_autodiff() in crates/burn-tensor/src/tensor/. inner() became an alias of without_autodiff() and no longer panics for a tensor without autodiff. The reverse conversion through from_inner became idempotent thanks to the new Tensor::autodiff(). This change did not touch the identically named Module::no_grad() method.

A more dangerous change landed in #5537, squash commit af55cf4432. Previously the Module implementation for Param<Flag> overrode no_grad() through with_value(false). As a result, model.no_grad() not only disabled gradients but also switched off dropout and the updating of accumulated BatchNorm statistics. After that override was removed, no_grad() controls tensor gradients only. The former behaviour moved to the new freeze(). unfreeze() and set_require_grad_group() were added alongside it. The change runs through crates/burn-core/src/module/base.rs and crates/burn-core/src/module/lora.rs.

The no_grad_does_not_disable_a_flag test in flag.rs checks that no_grad() no longer clears a module’s mode flag. The no_grad_only_disables_tensor_gradients test in dropout.rs checks that the method disables tensor gradients without putting dropout into validation mode. These tests pin down the new semantics. They do not confirm compatibility with the old behaviour. The compiler gives no warning about a change of this kind, so partial fine-tuning code using Dropout or BatchNorm may keep compiling while behaving differently.

#5571, squash commit d16f7ba2ed, complements this rebuild but is not itself a breaking change. Tensor::is_autodiff() and Tensor::is_tracked() were added to the tensor API. grad() and grad_remove() now return None when autodiff is off. #[must_use] was added to inner(), without_autodiff() and autodiff(). The changes are in crates/burn-tensor/src/tensor/.

The tensor API gained new option structs and stricter data

#5521, squash commit 7b815b01bd, changed how padding is represented in ConvOptions. Instead of a single symmetric value it stores (before, after) pairs. new() is intended for the symmetric case and new_with_padding() for the asymmetric one. PaddedConvOptions was deprecated. The change runs through crates/burn-backend/src/backend/ops/. As of the end of September, Conv3d in the IR still accepts symmetric padding only, so support for asymmetry across the whole stack is unfinished.

#5853, squash commit 9adff02644, moved max_pool1d, max_pool2d, avg_pool1d, avg_pool2d and the max_pool*_with_indices variants onto the MaxPoolOptions and AvgPoolOptions structs. Padding is given as start and end pairs through with_padding_pairs(). The defaults are aligned with PyTorch: stride equals kernel, dilation equals one, ceil_mode is off, count_include_pad is on. Asymmetric padding was moved from the nn modules down to the functional level in crates/burn-tensor/src/tensor/. There are no avg_pool*_with_indices variants. Migration examples were added to burn-book/src/migrating-to-0.22.md.

#5849, squash commit 84b9669a59, replaced the interpolate(x, output_size, options) call with interpolate(x, options). The target size now lives in options.output_size, or is computed from options.scale_factor. The public tensor API changed in crates/burn-tensor/src/tensor/, but the backend trait and the IR were not touched.

#5539, squash commit 163b7f442d, added multi-axis variants of the vector norm. At the same time it changed the narrow max_abs_dims(&[]) case. An empty axis list previously returned the original tensor; it now returns self.abs(), which matches ONNX with noop_with_empty_axes=true. At the time of the change the code was in crates/burn-tensor/src/tensor/linalg/vector_norm.rs, before linear algebra was moved out into burn-linalg.

#5838, squash commit e201487165, made the fields of TensorData private in crates/burn-std/src/data/tensor/base.rs. Access is provided through shape(), dtype(), bytes(), into_parts() and with_bytes_mut(). try_from_bytes returns DataError::InvalidByteLength when the buffer length does not match the shape and type. from_bytes panics on such input. Deserialisation was moved onto try_from_bytes, and the constructors check for overflow when counting elements. The length of quantised data is not yet validated, which was filed as the open #5836.

#5846, squash commit af8b3306db, removed into_tiled from Tensor, the backend trait, the IR, fusion and the router. The file crates/burn-std/src/tensor/tiled.rs was deleted. The practical significance of this removal is low. The API was added in #5839, squash commit 9ef294317c, on the morning of 25 September and removed roughly two and a half hours later. It was present in no tag.

The default dependency set was reduced

The most visible reduction comes from #5722, squash commit 596d9acfbd. Building the facade burn without a backend feature no longer compiles the Flex CPU backend. The configuration is set in crates/burn/Cargo.toml.

#5675, squash commit faa20a1f4a, removed tracing from the default features of burn-autodiff and burn-fusion. Their default = ["std", "tracing"] was replaced with default = ["std"]. For dependencies in burn-backend, burn-dataset and burn-vision, the weak syntax cubecl?/tracing and burn-std?/tracing is used. The change is concentrated in the Cargo manifests of the crates concerned and is described in crates/burn/src/lib.rs. No instrumentation code was added or removed.

#5806, squash commit 0b6f7cbf04, removed rl from the default features of burn and burn-train in crates/burn/Cargo.toml and crates/burn-train/Cargo.toml. This is a fix for regression #5769, not a new subsystem. In burn-train the feature was already part of default in v0.21.0. In the facade burn it became default in May 2026 in #5012, squash commit 6b721d887c.

#5544, squash commit ef7c981590, dropped unused dependencies from the Cargo manifests of a number of crates and examples. In crates/burn-cubecl-fusion/Cargo.toml, rmp-serde was replaced with ciborium.

#5680, squash commit 11f9ff07d0, moved the dataframe feature in crates/burn-dataset/Cargo.toml from the full polars to polars-core. The default set for that capability no longer includes the lazy engine or the file formats.

#5715, squash commit f8534af9e8, disabled the default features of the zip workspace dependency in the root Cargo.toml. By the PR author’s measurement, the dependency tree of pytorch-reader shrank from 76 crates to 33. That measurement was not independently verified. The same PR fixed the nlp feature of burn-dataset.

#5831, squash commit 28235bc19e, removed bincode from the workspace dependencies in the root Cargo.toml. The dependency remains in the lock file transitively through Polars.

The migration guide appeared before the 0.22 changes were finished

The guide burn-book/src/migrating-to-0.22.md was created in #5762, squash commit 9f1bb5112f. The same PR updated the README, the building-blocks sections, the contributor book and the notebooks.

The guide was extended by #5768, squash commit faec398324, and #5800, squash commit f56efabd89. As of 30 September the document contained 17 thematic sections. It kept being changed after the v0.22.0-pre.4 tag. After the tag, new incompatibilities were introduced by #5838, #5849, #5869 and #5853. #5800 itself was a documentation rework, #5806 fixed features, and #5821 fixed validation.

#5731, squash commit 28bfe74d5f, replaced standalone code fragments in the book with mdBook includes from maintained examples. As a result, the examples included in the book compile together with the original examples. The change is in burn-book/.

#5845, squash commit f5bd4c075c, added the page burn-book/src/advanced/runtime-configuration.md about burn.toml. This is documentation of an existing mechanism. Support for burn.toml was added back on 23 April 2026 in #4864.

v0.22.0-pre.4 did not complete the release transition

The lightweight tag v0.22.0-pre.4 was placed on 22 September on commit 9147c11a from #5777, which fixed the check of published versions through cargo info. The tag points not at the version-bump commit but at its direct descendant.

The bump itself from 0.22.0-pre.3 to 0.22.0-pre.4 was done in #5776, squash commit d3c8d7c615. That PR also replaced the git-rev dependencies on CubeCL and Cubek in the root Cargo.toml with the published versions =0.11.0-pre.4 and =0.3.0-pre.4. The workspace was thereby brought into a state fit for publishing by tag through .github/workflows/publish.yml.

As early as 23 September, #5797, squash commit 0fa63b3e45, returned main to dependencies by git revision. In the Cargo.toml of snapshot 91d892a4, CubeCL is again specified through git and rev. So main at the end of September was not in a publishable configuration.

The last stable tag remained v0.21.0 of 7 May 2026. The 0.22 pre-releases came out on 29 July, 10 August, 25 August and 22 September. No stable 0.22.0 was released in September.

Breaking-change marking is incomplete

Ten September PRs carried the ! marker in their title: #5521, #5539, #5572, #5647, #5675, #5720, #5721, #5849, #5853 and #5869.

At least five further PRs changed the public API without such a marker: #5557, #5537, #5838, #5846 and #5722. Of these, #5537 is particularly substantive, because it silently changes the semantics of Module::no_grad() and causes no compilation error. #5722 affects users of the 0.22 pre-releases but does not worsen behaviour relative to stable 0.21. #5846 formally removed a public method, but into_tiled lived for about two and a half hours on 25 September and reached no tag.

2. Reading PyTorch checkpoints

A standalone reader and a bridge into Burn

Reading of .pt and .pth was extracted from burn-store into the standalone pytorch-reader crate in #5656, squash commit 0fdac31a3d. The new crates/pytorch-reader/ depends on neither Burn nor PyTorch. The full list of its direct dependencies is limited to byteorder, crc32fast, num-traits, serde, tar, thiserror and zip. The crate name carries no burn- prefix, and its description positions the component outright as a reader independent of both systems.

The crate defines its own DType and Tensor. Tensor::read() materialises bytes only on request. burn-store pulls the reader in as an optional dependency behind the pytorch feature and re-exports it as burn_store::pytorch_reader. The bridge crates/burn-store/src/bridge.rs provides bridge::from_pytorch, which wraps a read tensor in a deferred burn_pack::Tensor. The former public path burn_store::nested was removed in the same #5656, squash commit 0fdac31a3d.

A separate publish-pytorch-reader job appeared in .github/workflows/publish.yml. The publish-burn-store job depends on it, so that the new standalone dependency is published first. Version pytorch-reader 0.22.0-pre.4 was published on 22 September 2026. The crate’s documentation was then synchronised between crates/pytorch-reader/README.md and crates/pytorch-reader/src/lib.rs in #5747, squash commit b3984918b5.

Rework of the formats and of pickle

Before the crate was extracted, the reader was substantially reworked in #5593, squash commit 33df298f84. The 570-line lazy_data.rs was removed and work with deferred data moved into storage.rs. The ZIP is now opened once, entries are looked up relative to the directory containing data.pkl, and allocations whose sizes come from an untrusted file are bounded. The changes are concentrated in crates/pytorch-reader/src/.

Two outcomes must not be credited to #5593 as new capabilities. TAR support and the rejection of big-endian existed before squash commit 33df298f84. This is visible in its parent revision of crates/burn-store/src/pytorch/reader.rs, lines 461–477 and 873. The PR rewrote these paths and pinned them with fixtures, but did not introduce the corresponding behaviour for the first time.

The claimed support for pickle protocols 0–5 likewise does not mean full coverage of every protocol. After #5593, squash commit 33df298f84, the opcodes EXT1, EXT2, EXT4, NEXT_BUFFER and READONLY_BUFFER are still rejected. The practical improvement is that protocol 4 no longer fails immediately on FRAME, not that the whole pickle specification is implemented.

Hardening the reader

#5700, squash commit 96e2ec1e12, removed clone_unsafely from enum deserialisation. The function copied a Serde visitor through ptr::copy_nonoverlapping even though its type required neither Copy nor Clone. A visitor holding heap data could be freed twice. This is not a September regression. The code appeared in March 2024 in #1436, commit 0138e16af6, and in June 2026 moved into burn-store/src/nested/de.rs. Ahead of the standalone reader’s publication the fix landed in crates/pytorch-reader/src/nested/de.rs. About ten Deserializer method implementations that called unimplemented!() were replaced with returned errors.

The file crates/pytorch-reader/tests/nested_de_miri.rs was added in #5700, squash commit 96e2ec1e12. It is intended to be run under Miri and checks that enum deserialisation no longer copies or frees the visitor in an unsafe way.

#5737, squash commit d79dab2d12, changed ownership of the source file in crates/pytorch-reader/src/storage.rs. LegacySource no longer stores only a path and no longer reopens the file for every storage. An open descriptor is now kept, a duplicate of it is created for reading, and bytes are read at explicit offsets. Deleting the file after it has been opened no longer breaks subsequent reads. Replacing the file through rename no longer mixes the directory of the new archive with the bytes of the old file. The PR #5752, closed without merging, was superseded by this fix, so it has no squash commit in main.

#5728, squash commit e68583e42f, stopped turning every uninterpreted pickle object into PickleValue::None. That behaviour let load_config silently substitute defaults such as 0, an empty string or false for numpy scalars, torch.dtype, torch.device and tensors. NestedValue::Unsupported(String), which preserves the name of the unsupported Python type, appeared in crates/pytorch-reader/src/nested/.

#5749, squash commit 4d26deac29, moved PickleError and PytorchError onto thiserror with #[from] and #[source]. A nested io::Error is now reachable through the standard Error::source() chain. The detect_format function in crates/pytorch-reader/src/reader.rs requires a valid pickle header and rejects SafeTensors separately. This removes the ambiguity in which the first length byte of a SafeTensors JSON header could coincide with a pickle opcode.

#5756, squash commit a36db79bc0, added compatibility with torch.save(compute_crc32=False). In that mode PyTorch writes a zero CRC-32 into the ZIP headers, and the zip crate treated it as an incorrect checksum. The archive-source code in crates/pytorch-reader/src/ now treats zero as the absence of a checksum, which matches the behaviour of torch.load.

#5664, squash commit 317fb474fc, extended extract_tensors in crates/pytorch-reader/src/. The traversal previously descended only through dictionaries, so tensors inside a list or tuple were lost without an error. Elements now receive indexed names. For example, {"weights": [w1, w2]} becomes weights.0 and weights.1. This matches the naming of Vec<Module> in Burn and of nn.ModuleList in a state_dict. The PR closes the regression from issue #5595.

#5714, squash commit b2faf4f336, narrowed the locking scope of ZipSource in crates/pytorch-reader/src/storage.rs. The Mutex<ZipArchive> was previously held for the whole duration of a storage read, so parallel Tensor::read() calls effectively ran sequentially. The lock is now needed only to look up an entry. In the author’s measurement, for 64 tensors of one million f32 each on eight threads, the time fell from 24 ms to 6.3 ms. The single-threaded result did not change. This measurement was not independently reproduced.

Two regression chains

The first chain: #5714, squash commit b2faf4f336, → #5737, squash commit d79dab2d12. For the sake of parallel reading, #5714 left a second open-by-path of the file on Windows. If a rename over the file happened between the two opens, the ZIP directory belonged to the new file while positional reads could address the old one. #5737 replaced the repeated File::open with duplication of the already open descriptor through try_clone(). The defect was introduced and fixed within three days. Both changes run through crates/pytorch-reader/src/storage.rs.

The second chain: #5728, squash commit e68583e42f, → #5700, squash commit 96e2ec1e12, → #5764, squash commit 03e2da00c5. #5728 added NestedValue::Unsupported. #5700 then replaced unimplemented!("deserialize_any") with an exhaustive match, but added no arm for the new variant. The result in crates/pytorch-reader/src/nested/de.rs was compilation error E0004, not merely an unreachable runtime error. The nested module is included unconditionally, there was no wildcard, the type was not #[non_exhaustive], and there was no feature gate either. #5764 added the missing arm 84 minutes later.

Reading full-model saves

#5766, squash commit 855fe4e812, added extraction of weights from pickles created through torch.save(model). In crates/pytorch-reader/src/ the BUILD operation recognises an object whose state contains the _parameters, _buffers and _modules tables and interprets it as a module’s state_dict(). Parameter names keep their familiar form such as layer1.weight.

The same #5766, squash commit 855fe4e812, added reading of protocol 2 set and frozenset as lists and fixed BUILD memoisation. A repeated reference to a module could previously return its state from before BUILD, which made nested tensors disappear.

The limits of the capability are recorded in crates/pytorch-reader/src/lib.rs, crates/pytorch-reader/README.md and burn-book/src/saving-and-loading.md:

So #5766, squash commit 855fe4e812, removes the mandatory prior export through state_dict() but remains a weight-transfer mechanism. It is not loading of an executable Python model. In line with the new boundary, the section declaring whole-model saving inherently wrong was removed from the book.

3. Writing and publishing checkpoints

Atomic writes and protection of an existing file

#5489, squash commit 137b73a5ae, fixed a dangerous ordering in SafetensorsStore::collect_from in crates/burn-store/. The former path used safetensors::serialize_to_file, which truncated the target file at the start of the write. Tensors were materialised from the device only after that. A readback error left a truncated file where the last valid checkpoint had been.

After #5489 the new container is first built in full in an adjacent temporary file and only then renamed into the destination. The shared implementation was moved into the public burn_pack::AtomicFile. The process-level guarantee is that after an error, a panic or a stop, either the complete new container or the previous file remains.

Review of #5489, squash commit 137b73a5ae, surfaced three further problems in crates/burn-pack/src/atomic.rs. First, sync_all on Windows requires a handle with the GENERIC_WRITE right; without it every file save failed with Access is denied. Second, commit no longer accepts an arbitrary path, which makes it impossible to publish a scratch file outside the predetermined destination. Third, rename_onto carries the access mode of the existing destination over to the temporary file; without that, a checkpoint with mode 0600 could be replaced by a file with mode 0644. Ownership and hard links are not preserved across the rename, since publication creates a new inode. That boundary is documented.

#5781, squash commit d17a1ff8b8, closed a race in overwrite(false) mode in crates/burn-pack/src/atomic.rs. The absence of the destination was previously checked before tensors were materialised. A file created by another process during the save was then silently overwritten by the rename.

The new AtomicFile::commit_new publishes the scratch file by creating a hard link at the target path. An already existing path returns Error::AlreadyExists, so the check and the publication become a single file operation. On a file system without hard-link support, a fallback to rename after confirming the destination’s absence is used. That is a weaker guarantee with a race window. The preliminary check was kept as a fast path, but now uses symlink_metadata so that a dangling symbolic link too is detected before tensors are read.

#5832, squash commit 18ac73b70b, made atomicity the default behaviour in crates/burn-pack/src/writer.rs. Writer::write_to_file now uses the atomic path. The former direct behaviour is available through write_to_file_in_place. The separate write_to_file_atomic method was removed.

This is a source-breaking change, but it affects an unpublished API. burn-pack appeared only on 16 June 2026 in commit eda9663665 and is absent from stable v0.21.0. Besides, #5832, squash commit 18ac73b70b, landed on 28 September. That is later than the v0.22.0-pre.4 tag of 22 September, so atomicity by default is not in even that pre-release.

Checking types and data integrity

#5490, squash commit c1d6874cb9, fixed Applier::apply_tensor in crates/burn-store/src/. The shape was checked previously, but not the kind of dtype. Float data could, for example, be applied to an Int parameter, after which the report announced a successful application and the error arose later in try_to_vec. The ApplyError::DTypeMismatch variant already existed, but no path constructed it.

The kind of dtype is now checked before materialisation, so a rejected tensor is not read from disk at all. The width of the type is deliberately not required to match. Half precision may be loaded as F16, and PyTorch int64 as I64. For a Bool parameter only U8, U32 and Bool are allowed, whereas the former code accepted any integer dtype.

#5576, squash commit a7be13739c, added a check of each tensor’s byte length through validate_tensor_byte_len to crates/burn-pack/. The expected size is computed from shape and dtype. Before this the reader checked the maximum offset in the container but could accept an individual tensor with an inconsistent size.

#5838, squash commit e201487165, carried the same check through to the burn-store bridge. The bridge::tensor_data function in crates/burn-store/src/bridge.rs now calls TensorData::try_from_bytes from crates/burn-std/src/data/tensor/base.rs. An inconsistent length returns PackError::TensorBytesSizeMismatch instead of panicking. For quantised data this check is not yet performed, which was filed as open issue #5836.

Filters, name mapping and streaming writes

#5558, squash commit ef521c0896, changed PathFilter::with_regex in crates/burn-store/src/. The former if let Ok(regex) silently ignored an invalid pattern, so the filter carried on with a different set of paths. Construction now calls expect("Invalid regex pattern"). The boundary of the decision is that this is an explicit panic at configuration time, not a returned Result.

#5750, squash commit bd9a529bbc, added map_indices_contiguous_except(regex) to crates/burn-store/src/. The former map_indices_contiguous renumbered lists on an all-or-nothing basis. The new variant allows matching paths to be excluded. The PR’s third commit fixed an adjacent defect: builders that change path names now clear an already populated tensor cache, so the result does not depend on call order.

#5666, squash commit 52c0535842, rewrote the streaming-write check in crates/burn-pack/. Instead of a global counting allocator, each test tensor holds bytes::Bytes and registers its own index on release. The check records that individual tensors’ data is released as the streaming write proceeds, rather than held until the whole container is finished. Coverage was extended to the then-existing write_to_file_atomic and to into_bytes. Global state and unsafe were removed from the test code.

Connection to the August review

#5489 appeared in the August review as WIP with branch commit 7263ca5cb6 of 31 August and integration expected on 1 September. The PR did indeed land on 1 September as squash commit 137b73a5ae. In September this line continued with no-clobber publication in #5781, squash commit d17a1ff8b8, and atomicity by default in #5832, squash commit 18ac73b70b. The main implementation of the whole chain is in crates/burn-pack/src/atomic.rs and crates/burn-pack/src/writer.rs.

August’s check for missing and truncated tensor data also received two continuations. #5576, squash commit a7be13739c, checks the size inside burn-pack. #5838, squash commit e201487165, moves bridge::tensor_data from a panic to a typed error. The adjacent #5490, squash commit c1d6874cb9, closes the dtype-kind mismatch that previously passed as a successful application.

Adjacent theme: dataset storage

#5546, squash commit e80e7981b9, concerns not model checkpoints but dataset storage. SqliteDataset in crates/burn-dataset/ was moved from rusqlite, r2d2_sqlite and serde_rusqlite to the Turso Rust engine. The public names were kept, and database files remain compatible in both directions.

The sqlite feature was removed from the default set of burn-dataset. In crates/burn-dataset/Cargo.toml the configuration changed from default = ["sqlite-bundled"] to an empty default. The facade feature burn/dataset no longer provides SqliteDataset on its own; burn/sqlite is required for it. The change is recorded as a migration entry in burn-book/src/migrating-to-0.22.md.

The Turso dependency in #5546, squash commit e80e7981b9, is pinned to the exact pre-release version =0.8.0-pre.8. At the same time turso::Error is present in the public SqliteDatasetError, so an external pre-release is part of the API surface. The move also did not eliminate the dependency on C: turso_core transitively uses simsimd, which compiles C code unconditionally.

The new path uses Turso in WAL mode, groups inserts into transactions and performs an explicit WAL checkpoint in set_completed, checking the returned result row. The max(row_id) query was unwrapped from coalesce, since the former form degraded into a full scan. In the author’s measurement on a split of 560 thousand rows, the time fell from 86 ms to 5 ms. These changes are in crates/burn-dataset/src/, and the figures quoted were not independently verified.

4. Edge-case numeric semantics

At least twenty changes landed in main during September that bring Burn’s behaviour on NaN, infinities and range boundaries into line with what PyTorch does. Twelve of them concern NaN and signed infinities directly. The practical point of such work is that before it a corrupted number could disappear along the way: relu, clamp and max pooling returned the boundary instead of NaN, and abs() gave a plausible finite gradient through an incorrect sign(NaN). Training carried on, and there was no sign of a failure.

This is a series, not a set of coincidences

The claim that this is a series rests on three independent observations, each visible in the repository itself.

First. The changes reference shared tracking issues with numbered items. In #5662, squash commit 11f8c072e1, the body states Refs #5615 (item 1). In #5663, squash commit c0a8b0d13d, it states Refs #5615 (item 6). The remaining links go to individual issues: Fixes #5609 in #5658, Fixes #5659 in #5718, Fixes #5610 in #5652, Fixes #5602 in #5661 and #5808, Fixes #5791 in #5798, Fixes #5788 in #5790, Fixes #5789 in #5827.

Second. In a single test file, crates/burn-backend-tests/tests/tensor/float/ops/clamp.rs, four different contributors passed coverage to one another in sequence through compilation conditions:

Third. The requirement is fixed in the code as a backend contract, not only in a test. #5665, squash commit 66a8a5ff8a, added to crates/burn-backend/src/backend/ops/tensor.rs the statement that sign(NaN) == 0 is part of the operation’s contract and that every backend implementation is obliged to honour it. The commit body states how the bug was found: by differential comparison against LibTorch, where x.log().abs().backward() gave -2.0 instead of -0.0. It also names the source of the divergence — a mismatched trait bound in an earlier refactoring, not a deliberate decision.

In most cases the reference is named outright in the code or in the commit body. In crates/burn-flex/src/ops/pool.rs the condition was supplemented with an is_nan check and a note about matching PyTorch. #5652 speaks of using the remainder with a single modulus operation, as PyTorch does. #5912 refers to torch.roll, and #5798 to PyTorch’s pooling output size.

What exactly was fixed

The disappearance of NaN in Flex operations was closed in #5662 for max pooling and in #5658, squash commit 52c101f455, for relu, clamp_min and clamp_max. The test test_max_pool2d_with_indices_nan_propagation in crates/burn-backend-tests/tests/tensor/float/module/maxpool2d.rs checks both the value itself and the index of the last NaN for F32, F64, F16 and BF16.

#5717 closed a case that was not an incorrect result but a process crash: f32::clamp panics if a boundary is NaN. #5718 and #5909 extended the correct clamp behaviour to the CubeCL backends, and #5733 to NdArray.

#5743, squash commit 752e82b252, fixed quiet_softmax on a fully masked slice, where subtracting minus infinity from minus infinity gave NaN. In crates/burn-tensor/src/tensor/activation/base.rs the maximum over a slice is now replaced with zero if all values equal minus infinity. The state of the former coverage deserves separate mention: the existing test_quiet_softmax_grad test was replaced in this PR, because it did not call quiet_softmax at all.

#5585, squash commit 5871c72ef6, removed a NaN in the cosine similarity of two zero vectors. The lower bound was previously applied to the product of the norms, so the square of a small epsilon went to zero in f32. Each vector is now normalised separately.

#5754, squash commit ab84b8e69f, fixed fmod_scalar, which returned the input unchanged for an infinite divisor, so that the remainder of infinity divided by infinity gave infinity instead of NaN.

Neighbouring edge cases were closed in #5629 (overflow of an eight-bit type in a boolean select_assign, replaced with a bitwise OR), #5661 and #5808 (masking the shift amount by the type’s width), #5738 (unit axes in the contiguity check, which sent data down the copy path and returned incorrect results on the CPU backend), #5790 (the result type of empty slice, repeat, cat and one_hot_fill), #5652 (the remainder, with tests in crates/burn-backend-tests/tests/tensor/float/ops/remainder.rs).

#5912, squash commit 91d892a4ad, stands out from the rest in that it changes the observable result of a public API. unchecked_roll_dim now cuts the tensor at size - shift rather than at shift. The correspondence table in burn-book/src/building-blocks/tensor.md promised that Tensor::roll matches torch.roll before this change too, which means the code diverged from the project’s own documentation. The former expected values in crates/burn-backend-tests/tests/tensor/int/ops/roll.rs encoded the opposite direction and were rewritten in the same PR, which also added a one-dimensional case that breaks the symmetry. In burn-book/src/migrating-to-0.22.md at the 30 September snapshot, roll is not mentioned, so a user following the migration guide will not learn of this change.

#5798, squash commit 9eadada8ef, brought the pooling output size under ceil_mode in line with PyTorch’s behaviour: the last window that starts in the trailing padding is discarded. calculate_pool_output_size in crates/burn-backend/src/backend/ops/modules/conv.rs and pool_output_size in crates/burn-flex/src/ops/pool.rs were changed, and an overflowing subtraction on an unsigned type was eliminated along the way.

The superficially similar #5653, squash commit d580fbca80, cannot be counted as part of PyTorch parity. Neither PyTorch nor a tracking issue is mentioned in its body. In substance it reconciles the forward and backward passes with each other: the backward pass of average pooling divided by the constant kernel volume, whereas the forward pass divided by the actual window size. The correct formulation is that average-pooling gradients under ceil_mode were incorrect.

What is missing at the end of the period

The line is unfinished, and that is visible from the compilation conditions remaining in the tests at the 30 September snapshot. In crates/burn-backend-tests/tests/tensor/float/module/maxpool2d.rs a comment is kept stating that the CubeCL and NdArray backends still skip a NaN if it is not the first value. The test clamp_min_max_nan_propagation_f64 in crates/burn-backend-tests/tests/tensor/float/ops/clamp.rs remains under any(flex, ndarray). The same gate sits on the double-precision variant in crates/burn-backend-tests/tests/tensor/float/activation/relu.rs. The observable confirmation that this line is finished will be the lifting of those gates and the closing of issues #5615 and #5602.

Early diagnostics at the public API boundary

A parallel line moves the failure from the backend to the public API boundary, where the message is the same for every implementation.

#5580, squash commit 64a69aca89, added a rank check in matrix multiplication. Before it the behaviour depended on the backend: NdArray failed with an overflowing subtraction, LibTorch returned a rank-zero scalar product, and for CubeCL no valid input existed. #5555, squash commit 38f3d8a7d1, added a compatibility check of batch dimensions under broadcasting and references an issue opened long before this period. #5564, squash commit a181e614aa, extended the checks to remainder, powi, powf, hypot and atan2 and fixed the operation identifier in the error text.

#5554, squash commit 5bd0170384, goes further and moves some of these checks to compile time. The macros assert_shape!, debug_assert_shape! and unpack_shape! are applied in cross-entropy, cosine embedding, CTC, RNNT, positional encoding and the attention mechanisms, so a rank mismatch becomes a build error.

#5860, squash commit c0878a0ab0, added a finiteness check for scalar parameters in the CELU, Hardtanh, Dropout and GaussianNoise configurations. #5827, squash commit e814443c76, replaced an opaque message about an unexpected data type with an explicit rejection of quantised input in *_like-style operations and in one_hot. The byte-length check from #5838 is described in the section on writing checkpoints.

5. Training metrics

Three metrics that can be used when saving a model and for early stopping were computed incorrectly. This is not noise but a systematic bias, and in two cases it is reproduced in the tests’ reference values.

#5578, squash commit 2646751f3b, closed two layers of the same bug in accuracy. The value inside a batch was divided by the number of valid examples, but the full batch size, padding included, was passed into the epoch accumulation. The epoch average came out biased. Separately, a fully padded batch produced a division of zero by zero and spoiled the whole epoch’s result with NaN. crates/burn-train/src/metric/acc.rs and crates/burn-train/src/metric/top_k_acc.rs were fixed, and crates/burn-train/src/metric/state.rs pins down the rule that an update with a zero counter records the current value and does not change the accumulated one. The tests test_accuracy_epoch_aggregation_excludes_padding and test_fully_padded_batch_does_not_poison_epoch_accuracy check, respectively, the weight of padding in the epoch average and the absence of epoch spoiling by a fully padded batch.

#5684, squash commit fb90882e3e, fixed average precision in the presence of tied scores. A threshold includes every example with the same score, so each positive example within a group must receive the precision computed at the end of the group. This is implemented through a reversal, a cumulative minimum and a reverse reversal. Direct confirmation of the scale of the divergence is in the PR itself: the expected value of the multilabel_micro reference case was changed from 0.5918017848017848 to 0.550135118149824. That is, the former numbers diverged from scikit-learn’s average_precision_score whenever the scores contained ties. The new test checks that the result is independent of the order of examples and of the split into batches.

#5900, squash commit 58cc57d767, rewrote AUROC. The former implementation built a pairwise comparison of all examples with materialisation of an examples-by-examples-by-classes tensor. The new one sorts the scores and walks groups of equal values in widened integer arithmetic. A test_auroc_large_epoch test on twenty thousand examples was added. It matters here not to confuse the rewrite with a change of result: the numeric semantics of NaN are preserved exactly, since the counts of positive and negative examples are computed before NaN values are discarded, and pairs containing NaN remain in the denominator, granting neither a win nor a tie. That is precisely what the test_auroc_nan_keeps_pairwise_semantics test pins down. Ties are now determined by exact equality, so positive and negative zero fall into the same group.

Alongside, two loss functions were fixed where padding and smoothing produced incorrect normalisation. #5569, squash commit 1bceb880ec, eliminated the dilution of cross-entropy: the mask removed padding tokens from the numerator, but the denominator remained the full batch. A shared averaging function that applies the mask to the normaliser too was introduced in crates/burn-nn/src/loss/cross_entropy.rs. A fully padded batch returns NaN by documented design. The assert_padding_invariant test compares a padded batch with an equivalent unpadded one across all combinations of weights, smoothing and logits. #5892, squash commit 9e2eefff24, added clamping of probabilities in the label-smoothing path, which previously took a logarithm without a bound and gave infinity at zero probability.

#5875, squash commit 440a9e2299, moved dropout in multi-head attention from the scores before softmax to the weights after it. A warning was added to the returned struct that during training the rows of weights may not sum to one. The attention_dropout_applies_after_softmax test checks that with equal logits every weight after inverted dropout equals zero or one.

The connection with August here is direct and changes the character of the theme. August fixed the moment at which a metric is read: model saving and early stopping consulted metrics before the computation had finished, plus, separately, Dice aggregation and a crash when obtaining the final ROUGE-L. September fixes the formula itself. There is no continuation on Dice or ROUGE-L in this window.

6. The training loop, autodiff and optimisers

Gradient accumulation and checkpointing

#5729, squash commit 288b697e89, eliminated two independent defects. The first was that an unfinished gradient-accumulation window at the end of an epoch was simply discarded. A progress-completion state and a condition under which the accumulated value is applied were added. The second defect concerned the learning-rate schedule: the scheduler was stepped per batch rather than per optimiser update. In the ordinary and multi-device strategies this meant advancing the schedule N times faster when accumulating over N batches. In the distributed strategy a loop over participants called the scheduler step once per participant, so a single optimiser update took N times the number of participants schedule steps. The loop was replaced by incrementing the iteration counter by the number of participants, and the single schedule-step call was moved immediately before the optimiser step in all three strategies. A check that the number of accumulations is positive was added. A new test in crates/burn-train/tests/gradient_accumulation.rs uses a counting scheduler and checks the number of its calls and that the partial window is applied exactly once.

#5618, squash commit 2440750e19, closed a defect that prevented training with gradient checkpointing in balanced mode from reaching a second step. The module optimiser took a parameter off the autodiff tape and returned it, assigning the default checkpointing strategy. After the first step the updated parameters ended up on the default strategy while the inputs and frozen parameters stayed on balanced, and the next forward pass refused to combine them. The same showed up in L-BFGS when unrolling the vector back. Following review, a parameter context was introduced that preserves the require-gradient flag, the distributed flag and the strategy.

#5821, squash commit b9dec8235c, continues August’s reparameterisation theme. Obtaining a validation copy of a model previously folded the adapters into the weights and lost their structure. The validation copy now preserves the reparameterisation, quantised bases included, and the folding was moved into an explicit materialisation operation. Export of adapters only, through a parameter group selected by a regular expression, was added. The changes affect crates/burn-core/src/module/lora.rs, and the test is in crates/burn-core/tests/reparameterization.rs. The corresponding paragraph in the project’s book, stating that a snapshot folds adapters and discards checkpointing strategies, was replaced.

#5799, squash commit 2e5d2152be, changed how restoration of optimiser state is checked. Instead of a byte-for-byte comparison of the record, functional equivalence is checked, that is, that the step after restoration matches the step without it. The check was applied to Adagrad, Adam, AdamW, Adan, Lion, Muon, RMSProp and SGD.

Autodiff

#5647, squash commit 6212b88fe0, is marked as breaking. The fields of an autodiff tensor became crate-internal, a node guard object was introduced in crates/burn-autodiff/src/ops/base.rs, and preparation of a backward operation takes such guard objects and releases them only after the child has been registered. The effect is that an abandoned graph correctly gives up its buffers. Users of custom backends need to move to accessor methods instead of touching fields directly, and that is recorded in the migration guide. The test test_mm_reclaims_abandoned_graph_buffers_after_unrelated_backward in crates/burn-backend-tests/tests/autodiff/memory_management.rs is marked as a best-effort check: it allows up to sixty-four attempts and is excluded for NdArray. It follows that the release may be deferred, and the test gives no strict guarantee about when memory is returned.

#5645, squash commit 68a92fdfc9, closed a silent bug. Before it, a repeated backward pass over an already consumed tape could silently produce an incorrect parameter update. The name of the new test says exactly that: consumed_graph_is_rejected_before_an_incorrect_parameter_update in crates/burn-backend-tests/tests/autodiff/graph_reuse.rs. Such a call is now rejected, and the provenance check happens before additional steps are consumed, so fresh branches remain usable. A leaf flag appeared on the parent node, thanks to which parameters remain valid points at which to obtain a gradient. Documentation was added in crates/burn-tensor/src/tensor/api/autodiff.rs stating outright that retaining the graph for a repeated backward pass is not supported and that losses should be combined or the forward pass repeated instead.

#5692, squash commit 98e48ddbdd, split out a separate operation for raising to an integer scalar power and registers an explicit zero gradient for the zeroth power, adding fast paths for powers one, two, minus one and minus two. The tests are in crates/burn-backend-tests/tests/autodiff/pow.rs and cover non-finite inputs among other cases.

#5837, squash commit 28c871ce18, takes one line but fixes a genuine bug. Insertion of a dimension and the subsequent summation used an index one greater than required. Because repetition along a dimension lays copies out in blocks, with a non-uniform incoming gradient the result was not merely permuted but outright incorrect. The test should_diff_repeat_non_uniform_grad in crates/burn-backend-tests/tests/autodiff/repeat_dim.rs supplies a non-uniform gradient and expects values the former code could not produce.

#5547, squash commit 708573a449, fixed the backward pass of the product, where global reductions give rank one while binary operations require matching ranks. There is no test of its own in this PR. Coverage of empty axes came separately in #5598, squash commit 0d26e3f83a, and #5641, squash commit ee94f264ca, with tests in crates/burn-backend-tests/tests/autodiff/aggregation.rs.

#5624, squash commit 7ecb912cab, added transfer of a tensor between backends while preserving the graph. Moving to a device preserves the graph, branching off a copy does not, and moving to another device does not apply that device’s autodiff settings.

#5889, squash commit 697ae01af7, deserves separate mention because it shows the state of the former coverage. A test that had lost its expected-panic attribute, and therefore checked nothing, was removed. Three tests sharing an expected panic were tied to specific messages. The full-precision check previously only made sure gradients were present; it now verifies the type and the values. Autodiff tests of the attention mechanism, which did not exist at all, were added, along with tests of index retrieval in minimum, maximum and max pooling. This sharpens August’s caveat that tests check specific scenarios: some of them did not check even that.

Optimisers

#5775, squash commit f9faec9a66, added Adafactor in crates/burn-optim/src/optim/adafactor.rs. Second moments are stored in factored form by rows and columns for matrices and for tensors above rank two, so optimiser state does not grow as the number of parameters. The learning rate is passed to the step rather than to the configuration, by default the relative step is bounded by the inverse square root of the step number, and scaling by the root-mean-square value of the parameter is enabled. Tests in the same directory check the vector and matrix paths against a scalar reference implementation over two steps, check a round trip through serialisation, check type preservation at reduced precision with state in single precision, and check rejection of invalid epsilons and of the clipping threshold.

#5741, squash commit c7c72164e3, added Lion, which stores one moment instead of two. It matters here not to retell the tests as claiming more than they do. Both tests are marked as smoke tests: in the first, which trains a tiny linear regression, the comment says outright that this is not a reproduction of the paper’s accuracy. The second compares not memory at run time but the length of serialised state on a single linear layer, and the comment calls it a stable, backend-independent indirect indicator.

#5843, squash commit 81b5a3b315, fixed gradient clipping by norm at half precision. The sum of squares overflowed the half-precision range, the norm became infinite, the clipping coefficient went to zero and the gradient was zeroed out entirely. The norm, the coefficient and the scaling are now computed in single precision with the result cast back. The tests test_clip_by_norm_f16_does_not_overflow and a variant for a norm above the half-precision limit are in crates/burn-optim/src/grad_clipping/base.rs.

7. Physical memory layout at execution time

Operands are read in the order in which they lie in memory

The most coherent performance line of September is not about individual kernels but about the fact that Burn used to force data into contiguous form instead of working with its actual layout. A convolution may return a tensor physically laid out so that the channel varies fastest, even though the API presents it as a tensor with channels in the second dimension. Any following operation read such a tensor with a large stride and did not vectorise.

#5625, squash commit 5533260720, introduced launching in memory order for binary, scalar and unary element-wise operations. Operands are fed in the order in which they are laid out, the order is chosen by a byte-weighted vote, and the result is permuted back. The helper functions that determine dimension order moved from burn-cubecl-fusion into crates/burn-std/src/tensor/layout.rs. The PR author’s measurement: CUDA, one L4 card, median of twenty runs. The training step fell from 418.7 to 364.2 ms, of which the backward pass went from 225.7 to 181.5 ms. The forward pass did not change, 166.7 against 166.0 ms.

#5620, squash commit 265bb944d9, lifted a restriction that could make this whole mechanism not work at all. A tensor with alignment gaps in its allocation is now admitted to the vote on a block’s layout, with a gap allowed but overlap not. Before that, on a backend with an aligning allocator the layout propagation never triggered, which from the outside looked like an unexplained difference between backends. The author’s measurement: L4, median of ten runs. On element-wise normalisation with an activation over a channels-last representation, shapes 4 by 48 by 384 by 384, the gap behind the contiguous layout narrowed from 6.28× to 1.10×. The convolutional encoder sped up by 1.58×, the full forward pass by 1.6×, the training step by 1.11×.

#5796, squash commit 0e38b76bf9, extended the same approach to type casting. The author’s measurement on wgpu: casting from half to single precision for a channels-last tensor sped up from 8.6 to 2.2 ms, against 2.1 ms on a contiguous layout.

#5760, squash commit 2e626a85d7, replaced the transposed two-dimensional convolution kernel with a channels-last variant. The author’s measurement on an Apple M4 under wgpu showed that the ratio of the data-gradient time to the forward pass for one-dimensional convolution fell from 6.9 to 0.3. A substantive bug was fixed along the way: the end of the window was derived from the start plus a fixed width, so with a stride greater than the dilated kernel size the output collapsed to a single offset.

An important caveat applies to all four numbers. These are PR authors’ measurements, taken on different configurations and against different baselines, so the percentages cannot be added up. On the forward pass, which fuses into a single operation from end to end, there is no measurable effect, and that is visible in #5625’s own data.

The operation-fusion scheduler

#5622, squash commit 00215d4092, changed the moment at which a block is committed: only the part of it that cannot be fused is now committed. The description contains the picture before the change, namely 3,588 single-element segments per step and 3,238 unfused element-wise operations, of which 2,504 came in runs of two or more. The author’s measurement on an L4 at batch size four, medians of twenty runs: from 366.1 to 320.0 ms, and again from 370.8 to 324.8 ms.

#5850, squash commit 219077efa1, removed quadratic complexity in the router: the fusion estimate was recomputed over the whole accumulated graph after every operation. The count is now kept incrementally, and the full recomputation was kept as a test reference in crates/burn-router/src/fusion.rs. There are no numbers in the PR, so the size of the gain cannot be stated.

#5901, squash commit 7bc1dd6204, eliminated the case where a tensor released on another thread before its producer had run forced the fusion server to flush the entire operation stream. The motivation is taken straight from the commit body and relates to remote execution: the training event stream releases each step’s outputs, and on top of a remote backend every step was cut in a new place. A release now waits in a deferred queue.

#5621, squash commit 1fd8d24bee, sped up scatter-add by allocating one worker per value and using atomic addition. The limits are stated explicitly: the addition operation only, and only types with atomic addition, whereas assignment, multiplication and logical OR remained sequential. There is a cost here that must be named next to the gain: the order of additions is not fixed, so the rounding of a floating-point sum may differ between runs. The gate is in crates/burn-cubecl/src/kernel/index/scatter.rs. The author’s measurement on an L4, a bank of one-dimensional convolutions, 1,600,000 values into 8,192 slots: the training step from 95.9 to 81.2 ms.

Fusion fixes over the month produced two cases with visible consequences. #5871, squash commit d9b8377fae, eliminated the choice of vector size along an unsuitable axis for a transposed operand, which the trailing part of a fused kernel also reads. The consequence was observable: an expression with a transposed operand returned garbage, which made the Newton–Schulz iteration in the Muon optimiser diverge into NaN. #5847, squash commit 96425f5959, fixed a mismatch in which the shape was taken from the output arguments and the strides from the input ones, which gave either foreign strides or an out-of-bounds index. #5725, squash commit 18b3132c68, removed autotuning for empty reductions, and #5902, squash commit f328596411, a double remapping of the input offsets of a fused matrix multiplication.

#5822, squash commit 80d3a024f3, factored out a shared policy for bounding autotuning by a roofline model and removed the duplicate table in burn-cubecl. The fused matrix-multiplication and reduction tables got bounds for the first time. There are no numbers in the PR.

#5673, squash commit d40f99ea69, added generation of fusion implementations for backend extensions in the new crates/burn-backend-extension/src/fusion.rs. This is a direct consequence of the extension mechanism from section 1: operations that have been moved out must be able to fuse.

Tiled storage went through three steps in a row: #5631, squash commit df92ec0750, carried such a tensor through all the transformations, #5839, squash commit 9ef294317c, added the operation that converts to tiled form, and #5894, squash commit a2edbd2663, renamed the operations in the wake of the external dependency. That is, the line moves at upstream’s pace and changed names once in a single month.

Convolutions

#5545, squash commit 7eaab04251, is often described as speeding up convolution in general, and that is wrong. A single file, crates/burn-cubecl/src/kernel/conv/direct.rs, was changed, and the optimisation is enabled by a condition requiring the hardware’s maximum plane size to equal one. That is, it works only on the CubeCL CPU runtime, while on a GPU the former kernel is compiled. The author’s measurements on a 5700X processor under the CPU runtime, medians of four runs: ResNet50 at batch size one from 165.1 to 145.1 ms, at batch size eight from 818.7 to 604.5 ms, MobileNet unchanged. The reason for the gate is also stated in the commit body: on three NVIDIA cards under CUDA the new shape was nine and twelve per cent slower in single precision and up to twice as slow in half precision.

Eight days later this kernel left the repository. #5646, squash commit da9e4224e4, replaced it with a call to the direct convolution of an external dependency, shrinking the same file from 422 lines to 87, and the commit body says the kernel moved verbatim.

#5811, squash commit d85aa0ab39, moved the folding of four-dimensional blocks onto grouped convolution, thanks to which the unrolling weight shrank along the input channel. There are no measurements in the PR, so the conclusion about a gain follows from the shape of the weight and not from a measurement.

Flex

Twenty PRs landed in Flex over the month, of which five concern performance, fourteen fixes and one documentation. The performance list: #5617, squash commit 78159b741e, on views, copies, in-place operations and parallelism; #5636, squash commit 6e8932ca2a, on parallelising attention over batch-and-head pairs in crates/burn-flex/src/ops/attention.rs; #5603, squash commit 89bb954882, on in-place slice assignment in crates/burn-flex/src/ops/slice.rs; #5761, squash commit 1dafe9f4a2, on the time and memory of transposed convolution in crates/burn-flex/src/ops/conv_transpose.rs; #5865, squash commit 943b90eb20, on updating the vectorisation library and hoisting the SIMD path choice out of hot loops, with changes in crates/burn-flex/src/simd/kernels.rs among other places.

A direct caveat is needed here, otherwise the conclusion will outrun the data. This line contains no published measurements. #5617, #5636, #5761 and #5603 contain no numbers at all, only added benchmarks. The single number for Flex over the whole month appears in #5865, where a matrix multiplication of shape one million by eight sped up from roughly 650 to 290 microseconds, with the hardware not named. So it cannot be claimed on September’s data that Flex became faster by any particular fraction. This differs from August, where a measurement was given with a specific configuration.

It is worth noting separately that #5603 is less a speed-up than the removal of a hidden cost: conversion to contiguous form returned a clone, a second owning pointer kept the reference count at two, and an in-place change copied the whole destination on every slice assignment. A consuming variant of the operation was added.

Correctness fixes in Flex matter more than the above for a user who was recommended this backend for CPU in August. #5634, squash commit 3d72c99c84, closed a summation fast path that accepted a zero-stride view, so that an expanded tensor gave an inflated sum. #5650, squash commit b7946a625c, restored accounting for trailing padding in three-dimensional convolution. #5663, squash commit c0a8b0d13d, replaced round-half-away-from-zero with round-half-to-even in grid sampling, which changes the selected pixel relative to PyTorch. The remaining Flex fixes are described in the section on numeric semantics.

New tensor capabilities

#5259, squash commit 5f74b3938a, closed August’s WIP on batched SVD, but not in August’s form. The power method was replaced with the one-sided Jacobi method with a tournament scheme for traversing pairs, which gives a linear rather than quadratic number of kernel launches per pass. The right singular vectors are obtained by an inverse transformation, and wide matrices are handled through transposition. The public entry point is in crates/burn-linalg/src/functions/svd.rs, that is, already in the new crate from section 1.

#5583, squash commit 10b74fbb3d, added einsum with two interfaces, a runtime call on an equation string and a macro. #5582, squash commit 766436c4ef, deserves separate attention in connection with August’s deprecation theme: before it every backend failed on the minimum and maximum variants of scatter and select-assign, and the implementation was added at once in NdArray, Flex, CubeCL and LibTorch together with the backward pass. That is, new functionality is still being written for backends declared as deprecating too. The matches in routing became exhaustive, so a missed combination is now a compilation error.

#5372, squash commit 4e8380f15e, added adaptive three-dimensional average pooling in CubeCL, and #5616, squash commit d38cb74d6b, integer exponentiation.

#5688, squash commit ac79057b48, introduced device identity and enumeration of physical cards, under which the same card reachable through CUDA and through Vulkan gives a single entry. Grouping is by bus address, or by device identifier on Windows. #5739, squash commit bbea651c02, arranged for a lazily initialised parameter to be created on the device it moved to, rather than on the original one followed by a copy. #5746, squash commit b9175d2225, made it possible to count a model’s parameters from shapes without allocating anything in memory. The purpose of the latter is stated outright in the commit body: planning how to split a model across devices. These are the only two primitives that reached main from the theme discussed in the WIP section.

The pace of the external dependency

Of the window’s 246 PRs, 36 change only the revision of the external CubeCL and Cubek dependencies in the root manifest together with the lock file. A further 18 PRs move the same pin but also carry their own code changes, so they cannot be called dependency maintenance. In all, 54 PRs touch the revision pin, that is 22 per cent of the month. Some fixes come precisely from there, for example the rejection of invalid kernel shapes in #5890 and the reading of an empty tensor in #5704. The limit of inference matters here: the public history shows the frequency of integration but not the reason for each revision, so it cannot be concluded from this data that Burn’s development is constrained by the external dependency. The only direct fact in that direction for September is the move of the convolution kernel from #5545 into Cubek eight days later.

8. Remote execution

Twelve PRs with scope remote landed over the month, and eleven of them were merged between 24 and 30 September. The twelfth, #5724, squash commit 07575391ff, left installation of a process-wide log subscriber to the program rather than the library. In effect all work on this theme is concentrated in the last week of the period.

Before examining the content, two different notions that are easily conflated need separating. The file crates/burn-tensor/src/tensor/distributed.rs was not changed once in September. That is, distributed training with data replication and collective operations did not advance in this window. Everything discussed below concerns the transport and the session lifecycle of burn-remote.

A client failure stopped leaving resources occupied

Three changes close one class of failure and work only together.

#5825, squash commit e0319c6de7, made a failure detectable at all. Before it neither side sent probe messages and nothing timed out, so a client could wait on a read indefinitely. A TCP-level liveness check on the connection was added: the first probe after ten seconds, then every five, with a reset after four unanswered. As a result a vanished peer shows up as a read error in about thirty seconds, that is, the same way the Iroh transport does it through its own idle timeout.

#5828, squash commit 905c1954f6, closed a leak on the server. The session service loop exited through the error-propagation operator on any read error, bypassing session teardown. A client that disconnected without an explicit close — and that is every client, including an orderly endpoint close and an idle timeout — left the session registered forever. The worker thread did not exit and tensors stayed on the device. The consequence is named in the description outright: on a card with six gigabytes, several sessions in a row filled the memory, after which the next run failed for want of it. The read loop was moved into a named function, the handshake response is encoded before the session is bound, and the thread binding is performed under a single lock. The changes run through crates/burn-remote/src/server/pump.rs.

#5878, squash commit d2412e60df, closed the same path for a panic. A panic on a session’s worker thread unwound past teardown, objects exposed for transfer within the host stayed in the server’s registry, and the memory was not returned to the device. Unwind interception and response tasks with abort handles were added in crates/burn-remote/src/server/worker.rs.

There is a citation trap here worth knowing about. The body of squash commit d2412e60df mentions #5903 and #5904, but their changes are not in its own diff: both commits are already its ancestors. The diff of #5878 is limited to the worker, startup and local-exchange files. Citing this PR by the text of its message credits it with someone else’s work.

Independent fixes outside this line

#5829, squash commit d373a5f90a, eliminated a stall of blocking reads inside an asynchronous task after roughly 128 operations with the processor fully loaded. The cause is that every poll spent the task’s cooperative budget, which is only replenished when control is yielded. The remedy is to lift the limit explicitly for such a region.

#5903, squash commit 5dac07a9ed, raised the WebSocket frame size limit, which defaulted to sixteen megabytes and therefore made any upload of a larger tensor drop the session. The new limit is aligned with what the Iroh transport uses.

#5898, squash commit 52bb097cbf, eliminated an overlap of device addresses. Remote devices occupied type identifier zero, so a remote device with a given number received the same identifier as a local CUDA card with that number and shared its runtime worker thread. The remote device type was moved to the end of the range next to the graph-capture type in crates/burn-backend/src/backend/router_device_type.rs.

#5851, squash commit 9b5402e9f1, closed a leak during data loading: an initialisation operation captured by the fusion scheduler in the router waited on the execution of a stream that does not exist on a loader thread. The description quotes consumption growing over an epoch from 74 to 1,865 megabytes, but that figure was obtained on a remote-training example that is not in main, so it cannot be reproduced from the repository.

#5678, squash commit dec3046c39, does not concern burn-remote at all, although it is thematically close. A transfer between two devices of the same wgpu, ROCm or Metal runtime panicked on the device threads, the panics were swallowed, and the destination tensor was left unwritten. The path through the collective-operations library for CUDA was kept.

A configurable server

#5910, squash commit d97e804890, is the theme’s largest PR of the month and the only one that adds a capability rather than removing a failure. Channel and peer builders for the Iroh transport appeared, along with token authorisation with constant-time digest comparison, loading or creation of a secret in a file readable only by its owner — written through a temporary file and a rename that does not replace an existing one — a choice of relay mode from three variants, a configurable port, and connection to a peer by its identifier. Offloading of segmentation was moved into the configuration file and is off by default after #5819, squash commit 017464b63d. Debug output of the peer and session-initialisation structs does not print the secret. The update of the transport library itself to version 1.2.0 landed separately in #5830, squash commit 757438aa98, and consists of one line in the manifest.

The limit of inference

What is described above is the release of resources on failure, not resilience to it. Reconnection of an already established session is not in main at the 30 September snapshot: the ensure-connection function in crates/burn-remote/src/client/service.rs returns immediately if the writing half is present, and it remains present after the connection has died. The retries from #5824, squash commit 740cc4f810, with delays from a quarter of a second to eight seconds, apply only to the first opening of the channels. A client’s tensors are not restored after a disconnect.

The client’s behaviour on a disconnect is not uniform. Synchronisation and asynchronous reading of a tensor return an execution error, writes that do not wait for a response are logged and discarded, but a query of the data types in use and the handshake panic.

In that light #5904, squash commit e501c844bf, moves the panic rather than removing it. The server stopped crashing on a request for a non-existent device index and now logs and closes the stream instead, but the client in that same scenario crashes while awaiting a successful handshake. There is no typed session failure in main: the corresponding enumeration does not exist, and the connection error type that does exist covers only a missing address, a bind error and an abort.

Session teardown is also not bounded in time. An ordinary disconnect does not abort the response tasks forcibly — that is done only after a panic — the executor synchronisation runs without a timeout, and the wait in the local exchange in crates/burn-remote/src/server/local_comm.rs is performed in an unbounded loop. A hung read can delay the closing of a session.

WIP at the end of the period

This section was reconstructed retrospectively from the Git history and current pull-request metadata. PR state was taken on 1 October 2026 rather than on 30 September, so no particular PR can be credited with draft status or with having an approval as of the period’s closing date: the available data contains no state-change history. A subsequent merge does not count towards September’s results. The size of a branch and the number of commits in it are not an estimate of readiness, and reasons for a delay cannot be inferred from them either. The dates of branches’ last commits are given in UTC−4, since in their original zone some of them look like 1 October although in the series’ zone they belong to 30 September.

Splitting a model by layers across devices. #5702 adds a pipeline trait, a layout by stages and a matching example. As of 30 September the branch holds eleven commits outside main, the last of them from 23 September. None of the types this PR introduces are present in snapshot 91d892a4. Only two primitives from this theme reached main, described in the section on memory layout: creation of a lazily initialised parameter on the target device in #5739 and counting the number of parameters from shapes without allocating memory in #5746. The purpose of the second is stated outright in its commit body, namely planning how to split a model across devices. So the repository contains groundwork for this capability, but the capability itself is not in main.

Returning execution errors to the training loop. #5896 replaces panics in the training loop with structured error types. It is the period’s largest open initiative, twenty-four commits outside main, the last from 30 September. The theme directly continues August’s change in which a compute failure stopped breaking independent work in the queue: there the error was tied to the tensors affected, here it is meant to reach the calling training code. The profiling that is sometimes associated with this branch is present in main, but came from a different and merged #5726, squash commit d0365962df.

A typed failure for connecting to a remote device. #5922 makes connecting return a result instead of panicking and introduces an enumeration of session failure causes. Alongside it are #5723 on serving streams a peer opens on a connection we dialed, #5920 on builds with a partial feature set, #5919 with a regression test for tensors dropped on another thread, and #5906 with an example of training on another machine’s GPU. None of these changes are in main, and the corresponding enumeration of failure causes is not there either. This matters for reading section 8: typed handling of session failure remains unfinished as of the end of September.

Retaining the graph for a repeated backward pass. #5043 moves the step and backward-operation traits onto shared ownership so as to allow several backward passes. Ten commits outside main, the last from 22 September. The contrast with what was merged in the window matters here: #5645 added documentation to main stating outright that retaining the graph is not supported and that losses should be combined or the forward pass repeated instead. That is, the absence of the capability is pinned down in the repository as of the end of the period, while the work to add it proceeds in an open branch.

The learning rate as a device tensor for graph-captured steps. #5872 moves the learning rate onto an enumeration that admits a tensor on the device and replaces a process-wide counter of active graph captures with a per-device query to the backend. Ten commits outside main, the last from 28 September. This is September’s only substantive contact with the computation-graph capture theme, and it is not merged. The graph-capture crate itself was touched by only two PRs tangentially in the window.

The remaining substantive initiatives, grouped by area. Reading a checkpoint from memory without a file continues in #5755, which belongs to section 2. Three-dimensional pooling primitives in #5801 and differentiable trilinear grid sampling in #5803 complement the breaking change to the pooling interface from section 1. The data gradient for convolutions in #5813 and the folding path with autotuning for three-dimensional transposed convolution in #5793 form a symmetric continuation of August’s weight-gradient work, and neither of them landed in main. An arbitrary-size Fourier transform in #5757 and a complex inverse transform in #5840 extend the new signal-processing crate, while a linear solve in #5856 and a Cholesky decomposition in #5858 extend the new linear-algebra crate. That is, both extracted crates immediately acquired a work queue of their own. The Matthews correlation metric in #5818 belongs to section 5 and is absent from main. Alongside are #5862 on an unused gate in the coupled LSTM implementation, #5908 on connected components in the CubeCL backend, #5911 on dropping an incomplete batch in the data loader, and #5674 on exposing tolerance fields.

Long-lived open PRs and a correction to the August material. #5504 with the GroupNorm primitive has been open since August, its branch’s last commit dating from 29 August. The internal August reference report described it, incorrectly, as merged. As of 30 September it is still open, as are six further PRs attributed to August’s work in that document: #3608, #5190, #5309, #5367, #5387 and #5541. Not one of the seven was integrated during the month.

PR Purpose Branch’s last commit (UTC−4) Commits outside main
#5896 refactor(train): Re-surface execution errors in training loop 2026-09-30 24
#5702 feat(core): pipeline trait to split a model by layers across devices 2026-09-23 11
#5906 feat(examples): train and run MNIST on another machine’s GPU with remote-mnist 2026-09-30 11
#5043 Retain graph feature for partial derivatives 2026-09-22 10
#5872 feat(optim): accept a device learning rate for graph-captured steps 2026-09-28 10
#5504 feat: add GroupNorm backend primitive and CubeCL kernel 2026-08-29 8
#5920 fix(remote): gate what partial feature builds leave unused 2026-09-30 8
#5723 fix(remote): serve the streams a peer opens on a connection we dialed 2026-09-28 7
#5801 feat: add 3D spatial pooling primitives (avg_pool3d, max_pool3d) (#5785) 2026-09-30 5
#5755 add from bytes 2026-09-26 3
#5757 feat(signal): support arbitrary-size rfft/irfft via Bluestein’s algorithm 2026-09-30 3
#5922 feat(remote)!: connecting to a remote device returns a Result 2026-09-30 3
#5674 Expose Tolerance fields 2026-09-14 2
#5803 feat: add differentiable trilinear grid_sample_3d 2026-09-23 2
#5919 test(remote): feed a queued reader from tensors dropped on another thread 2026-09-30 2
#5921 fix(metal): require native MSL for explicit Metal devices 2026-09-30 2
#5541 Validate CubeCL complex cast lowering 2026-08-31 1
#5793 feat(cubecl): add a col2im path and autotune for conv_transpose3d 2026-09-23 1
#5813 perf(conv): add im2col data-gradient path for dense convolutions 2026-09-24 1
#5818 feat(train): add Matthews correlation coefficient metric 2026-09-24 1
#5840 feat(signal): add complex inverse FFT (ifft) 2026-09-30 1
#5856 feat(linalg): add batched linear solve 2026-09-26 1
#5858 Cholesky implementation 2026-09-26 1
#5862 fix(nn): omit unused forget gate in coupled LSTM 2026-09-26 1
#5908 fix(vision): call the CPU connected components directly in the cube backend 2026-09-29 1
#5911 Add drop_last to the data loader 2026-09-29 1

Full coverage of the period

Below are all 246 PRs merged into main in the window, each exactly once. The grouping is by area of change, not by importance: the themes examined above appear here on equal terms with the rest, so that the list stays complete and checkable. For each PR the merge date in UTC−4 and the squash commit in main are given, by which the statement can be verified directly.

There is not a single direct commit in the window, so no separate section for them is needed. This differs from August, where the window contained one direct commit.

Updates of the external cubecl and cubek dependencies (41)

PR Merged Squash commit Content
#5548 2026-09-02 435b354974 update cube
#5550 2026-09-02 ab0439eef6 Update cubek
#5560 2026-09-03 7f5736d248 update cube
#5561 2026-09-03 6fe858268a chore(deps): update cubecl and cubek for fallible throughput probes
#5556 2026-09-04 6e3ada50d0 Chore/update cubecl runtime erasure
#5565 2026-09-04 530f68153e update cube
#5566 2026-09-04 1f96ddaacc Update/cube3
#5586 2026-09-06 05ac3f79c1 chore: update cubecl and cubek revs
#5587 2026-09-06 3073dba638 Update cubek
#5589 2026-09-07 abb50388bc update cubek
#5591 2026-09-07 b9a1563d9f Update CubeCL & CubeK
#5611 2026-09-08 5ecfb38965 update cubek
#5627 2026-09-08 8b35c753cc update cubek
#5632 2026-09-09 d4cd33666b update cube
#5644 2026-09-10 eb03c79803 update cube
#5649 2026-09-10 09f9b3f78b update cube
#5668 2026-09-14 e523ad3c3f update cube
#5670 2026-09-14 cfe65e1cfc Update/cubecl crate layout
#5691 2026-09-15 c5aa775981 update cube
#5697 2026-09-16 f621d6edb6 update cube
#5701 2026-09-16 d01bb5499a chore: update cubek
#5719 2026-09-17 3b8fb7387d update cube
#5703 2026-09-17 577fece293 chore: update cubek
#5732 2026-09-18 1d6f50330d update cube
#5763 2026-09-21 e13601d3c7 Chore/cubecl device capacity
#5770 2026-09-21 28550733d5 update cube
#5771 2026-09-22 c6f9fd4421 Chore/cubecl environment records
#5797 2026-09-23 0fa63b3e45 Chore/cubecl adaptive pool
#5810 2026-09-23 fafc9d7a92 Update cubecl and cubek to the tune plan record
#5841 2026-09-25 86781f0887 chore(deps): bump cubek to a066bdd
#5848 2026-09-25 c240f0011c chore(deps): bump cubek to b40b522
#5854 2026-09-25 474f46b73a chore(deps): bump cubecl to a1bb768 and cubek to 9db95ba
#5870 2026-09-28 49dcd5ba9e chore(deps): bump cubek to 7d60e30
#5887 2026-09-28 3d605097e2 chore(deps): bump cubek to 41a4ab0 and cubecl to 33b6dfb
#5895 2026-09-29 53ebdce405 chore: bump cubecl to f0cf8383 and cubek to 9f35e2ab
#5905 2026-09-29 5f3cc76826 chore: bump cubek to 6927503
#5907 2026-09-29 98ca0bdac8 update cube
#5913 2026-09-30 9fa49d9cdd chore: bump cubek to bb461fb
#5915 2026-09-30 f5d57e66b9 chore: bump cubek to 5753e85
#5917 2026-09-30 a6de085337 chore: bump cubecl to 48301b5 and cubek to 2e07d52
#5918 2026-09-30 903f6897e5 chore: bump cubek to e84f83d and cubecl to 48301b5

Reading PyTorch checkpoints (12)

PR Merged Squash commit Content
#5593 2026-09-10 33df298f84 fix(store): harden and restructure the PyTorch reader
#5664 2026-09-15 317fb474fc fix(store): collect PyTorch tensors nested in lists/tuples with indexed names
#5656 2026-09-16 0fdac31a3d refactor(store): extract the PyTorch reader into a burn-free pytorch-reader crate
#5728 2026-09-18 e68583e42f fix(pytorch-reader): report values load_config cannot represent instead of defaulting
#5714 2026-09-18 b2faf4f336 perf(pytorch-reader): read stored ZIP entries outside the archive lock
#5700 2026-09-21 96e2ec1e12 fix(pytorch-reader): remove unsafe visitor cloning and deserialization panics
#5737 2026-09-21 d79dab2d12 fix(pytorch-reader): open a checkpoint once and never reopen it by path
#5747 2026-09-21 b3984918b5 docs(pytorch-reader): sync README with lib.rs
#5764 2026-09-21 03e2da00c5 fix(pytorch-reader): handle unsupported values in deserialize_any
#5749 2026-09-22 4d26deac29 fix(pytorch-reader): thiserror source chains + reject non-checkpoint files
#5756 2026-09-22 a36db79bc0 fix(pytorch-reader): accept checkpoints saved with compute_crc32=False
#5766 2026-09-23 855fe4e812 feat(pytorch-reader): load torch.save(model) full-model pickles

Model storage: burn-store and burn-pack (8)

PR Merged Squash commit Content
#5489 2026-09-01 137b73a5ae fix(store): write safetensors files atomically
#5490 2026-09-02 c1d6874cb9 fix(store): reject a loaded tensor whose dtype is the wrong kind
#5558 2026-09-04 ef521c0896 fix(store): reject invalid path filter regex
#5576 2026-09-08 a7be13739c fix(pack): validate tensor byte lengths when reading
#5666 2026-09-14 52c0535842 test(pack): check streaming memory with a drop hook, not a global allocator
#5750 2026-09-21 bd9a529bbc fix(burn-store): scope contiguous index mapping per prefix
#5781 2026-09-23 d17a1ff8b8 fix(store): enforce overwrite(false) at publish time, not just before the save
#5832 2026-09-28 18ac73b70b feat(pack): make atomic writes the default

Remote execution (13)

PR Merged Squash commit Content
#5724 2026-09-18 07575391ff fix(remote): leave the process-wide log subscriber to the program
#5819 2026-09-24 017464b63d fix(remote): turn off iroh GSO until iroh#4555 is fixed
#5830 2026-09-25 757438aa98 chore(deps): bump iroh to 1.2.0
#5824 2026-09-25 740cc4f810 fix(remote): try again while a server is not reachable yet
#5825 2026-09-25 e0319c6de7 fix(remote): notice a WebSocket peer that vanished without closing
#5829 2026-09-28 d373a5f90a fix(remote): blocking reads inside a tokio task stall after 128
#5828 2026-09-28 905c1954f6 fix(remote): end a server session when its client disconnects
#5873 2026-09-28 ccdfede767 test(remote): serve the tests on ports the OS picks, with burn-remote linked once
#5903 2026-09-29 5dac07a9ed fix(remote): let a WebSocket server read what its clients send
#5904 2026-09-29 e501c844bf fix(remote): refuse a session for a device the server does not host
#5898 2026-09-29 52bb097cbf fix(remote): keep remote devices off local GPUs’ runner threads
#5878 2026-09-30 d2412e60df fix(remote): clean up a session whose task panicked
#5910 2026-09-30 d97e804890 feat(remote): configure an Iroh server’s relays, port and authorizer, and dial one by its id

Autodiff (9)

PR Merged Squash commit Content
#5547 2026-09-02 708573a449 fix(autodiff): prod backward broadcasting
#5623 2026-09-09 04e0f2ded6 perf(autodiff): reduce a broadcast gradient over all dims at once
#5624 2026-09-09 7ecb912cab feat(autodiff): support graph-preserving cross-backend transfers
#5647 2026-09-11 6212b88fe0 fix(autodiff)!: retain input nodes until child registration
#5645 2026-09-11 68a92fdfc9 fix(autodiff): explicitly reject consumed graph reuse and preserve reusable leaves
#5679 2026-09-15 7d9e2be68c perf(autodiff): avoid redundant traversal in checkpoint topological sort
#5692 2026-09-17 98e48ddbdd fix(autodiff): preserve gradients for zero scalar exponents
#5837 2026-09-25 28c871ce18 fix(autodiff): correct repeat_dim gradient ordering
#5889 2026-09-29 697ae01af7 test(autodiff): tighten weak tests and cover untested backward paths

Training loop, metrics, TUI, RL (8)

PR Merged Squash commit Content
#5618 2026-09-08 2440750e19 fix(optim): keep a parameter’s checkpointing strategy across an update
#5574 2026-09-08 b6a4f30fc0 fix(rl): preserve deterministic mode in async batches
#5578 2026-09-08 2646751f3b fix(train): exclude padded samples from accuracy aggregation
#5684 2026-09-15 fb90882e3e fix(train): handle tied scores in average precision
#5729 2026-09-21 288b697e89 fix(train): flush partial gradient accumulation and step LR per optimizer update
#5805 2026-09-24 d26efc54a4 feat(train): add labels to training progress loggers
#5866 2026-09-28 c357872392 fix(tui): handle already-joined thread in manual close
#5900 2026-09-29 58cc57d767 fix(train): compute AUROC with sorted score groups

Optimisers (4)

PR Merged Squash commit Content
#5775 2026-09-22 f9faec9a66 feat(optim): add Adafactor optimizer
#5741 2026-09-22 c7c72164e3 feat(optim): add Lion optimizer
#5799 2026-09-23 2e5d2152be test(optim): strengthen optimizer save-load round-trip coverage
#5843 2026-09-25 81b5a3b315 fix(optim): compute gradient norm clipping in F32 for half precision

Layers and loss functions (6)

PR Merged Squash commit Content
#5569 2026-09-04 1bceb880ec fix(nn): exclude pad tokens from cross-entropy normalization
#5592 2026-09-09 986f0c1da1 fix(nn): honor count_include_pad for asymmetric average pooling
#5676 2026-09-16 55ee6c82f8 perf(nn): use a dedicated BatchNorm training op with closed-form backward
#5860 2026-09-28 c0878a0ab0 fix(nn): validate scalar module configurations
#5892 2026-09-29 9e2eefff24 fix(nn): clamp probabilities in smoothed cross-entropy
#5875 2026-09-29 440a9e2299 fix(nn): apply attention dropout after softmax

Flex, the CPU backend (20)

PR Merged Squash commit Content
#5619 2026-09-09 dd37d19950 docs(flex): fix stale statements in burn-flex docs and comments
#5603 2026-09-09 89bb954882 perf(flex): write slice_assign in place when the destination is uniquely owned
#5617 2026-09-09 78159b741e perf(flex): optimize views, copies, in-place ops, and Rayon parallelism (#5613)
#5629 2026-09-09 4a0d726530 fix(flex): avoid overflow in boolean select updates
#5630 2026-09-10 9cc7f63ac7 fix(flex): support unsigned dtypes in int_argmax/int_argmin
#5634 2026-09-10 3d72c99c84 fix(flex): guard sum fast path with layout_covers_storage_once
#5640 2026-09-14 18af6a8c8f fix(flex): validate gather_nd and scatter_nd coordinates
#5663 2026-09-14 c0a8b0d13d fix(flex): round nearest grid_sample ties to even
#5661 2026-09-14 7ef8198a48 fix(flex): make u64 right shift logical
#5650 2026-09-15 b7946a625c fix(flex): honor asymmetric padding in conv3d
#5636 2026-09-15 6e8932ca2a perf(flex): parallelize attention across (batch, head) pairs (#5612)
#5652 2026-09-15 e205024136 fix(flex): handle remainder edge cases
#5662 2026-09-15 11f8c072e1 fix(flex): propagate NaN through max pooling
#5658 2026-09-16 52c101f455 fix(flex): propagate NaN through relu and clamp_min/clamp_max
#5717 2026-09-18 dae47a34b8 fix(flex): handle NaN clamp bounds
#5780 2026-09-23 d918eff531 fix(flex): compute f32 layer_norm variance in two passes
#5761 2026-09-23 1dafe9f4a2 perf(flex): reduce ConvTranspose time and memory use
#5808 2026-09-24 ebad4f9d6d fix(flex): mask shift amounts to each int dtype’s own width
#5865 2026-09-28 943b90eb20 perf(flex): upgrade macerator to 0.5.0 and hoist SIMD dispatch out of hot loops
#5888 2026-09-28 3bf6e03948 fix(flex): enforce quantization block alignment

NdArray and shared primitives (6)

PR Merged Squash commit Content
#5543 2026-09-01 13637e9558 fix(no-std): use shared sync primitives across crates
#5553 2026-09-03 37f87ab37f fix(ndarray): preserve SIMD unary layout and recip precision
#5639 2026-09-10 d26b1efbb7 test(burn-std): import vec! macro in layout.rs tests for no-std
#5665 2026-09-15 66a8a5ff8a fix(burn-ndarray): sign(NaN) must be 0, not its hidden sign bit
#5733 2026-09-21 6ade1047f8 fix(ndarray): preserve NaN through float clamps
#5738 2026-09-23 9b08e4aba5 fix(std): ignore unit axes when checking contiguity

Fusion and router (13)

PR Merged Squash commit Content
#5620 2026-09-08 265bb944d9 perf(fusion): let a padded tensor vote for the block’s layout
#5622 2026-09-08 00215d4092 perf(fusion): settle only the unfusable head of a block
#5633 2026-09-09 ba1886d2cf test(fusion): change f32 expected value to oracle
#5725 2026-09-21 18b3132c68 fix(fusion): avoid autotuning empty reductions
#5850 2026-09-28 219077efa1 perf(router): avoid rescanning the graph when scoring fusion
#5871 2026-09-28 d9b8377fae fix(fusion): keep the default vectorization axis for matmul operands the epilogue reads
#5847 2026-09-28 96425f5959 fix(fusion): resolve reduce reference strides against the reference’s argument list
#5851 2026-09-28 9b5402e9f1 fix(router): close the fuser on an upload so the upload can be freed
#5877 2026-09-28 64e563fd9e fix(router): serve unsigned int tensors in the interpreter
#5893 2026-09-29 11a0757036 test(fusion): relax half-precision absolute tolerance for matmul epilogue view
#5890 2026-09-29 bd0f1dbc40 fix(fusion): update cubek to reject invalid VecMat kernel shapes
#5902 2026-09-29 f328596411 fix(fusion): avoid remapping fused matmul input offsets twice
#5901 2026-09-30 7bc1dd6204 perf(fusion): defer a cross-thread drop until its producer runs instead of draining the stream

CubeCL kernels and specific runtimes (18)

PR Merged Squash commit Content
#5372 2026-09-08 4e8380f15e feat(cubecl): support AdaptiveAvgPool3d
#5616 2026-09-08 d38cb74d6b feat(cubecl): implement integer powi operations
#5621 2026-09-09 1fd8d24bee perf(cubecl): scatter-add runs one unit per value with atomic adds
#5635 2026-09-09 20d53f5cfc fix(cubecl): gate storage_tiled tests behind runtime features
#5625 2026-09-09 5533260720 perf(cubecl): walk elementwise kernels in operands’ memory order
#5646 2026-09-11 da9e4224e4 refactor(cubecl): call cubek’s direct convolution routine
#5654 2026-09-11 1414c8a14e fix(cubecl): update cubek to fix sums over overlapping views
#5678 2026-09-16 dec3046c39 fix(cubecl): same-runtime device moves without a peer transport, and quantized moves
#5704 2026-09-16 7bc513b711 fix(cubecl): update dependencies to fix empty tensor readback
#5718 2026-09-18 274cb46cb9 fix(cubecl): propagate NaN through two-sided clamp
#5767 2026-09-21 18fa3bccae fix(cubecl): handle broadcasting in mask_where and mask_fill
#5760 2026-09-23 2e626a85d7 perf(cubecl): NHWC direct kernel for conv_transpose2d
#5796 2026-09-23 0e38b76bf9 perf(cubecl): walk cast in its input’s memory order
#5815 2026-09-24 caa39cf874 Metal/wgpu msl tests
#5867 2026-09-28 af3100e5af fix(deps): restore gpu-allocator’s windows version to match wgpu-hal
#5844 2026-09-28 bea4cdef84 fix(cubecl): reject reshape that splits a quantization block across rows
#5881 2026-09-29 baf6943408 fix(cubecl): skip the reduce launch when another axis leaves the output empty
#5909 2026-09-30 56584ee48b fix(cube): propagate NaNs in clamp and stabilize GPU tests

Backend selection, dispatch, core modules (20)

PR Merged Squash commit Content
#5520 2026-09-01 0e2dc6b631 refactor(dispatch): unify routing and strengthen autodiff contract
#5551 2026-09-02 cf1b6d2a60 fix(dispatch): restore autodiff context promotion
#5537 2026-09-03 af55cf4432 feat(module): separate gradient control from module freezing
#5521 2026-09-03 7b815b01bd feat(backend)!: support asymmetric padding in conv1d and conv2d
#5570 2026-09-04 54d36fa936 fix(dispatch): decouple CubeCL runtime facade crates
#5598 2026-09-08 0d26e3f83a test(backend): cover empty-axis autodiff reductions
#5641 2026-09-10 ee94f264ca test(backend): cover empty-axis autodiff product reductions
#5673 2026-09-16 d40f99ea69 feat(extension): generate Fusion implementations for backend extensions
#5705 2026-09-17 b9088436b6 fix(core): let an init_mapper parameter train on an autodiff device
#5721 2026-09-17 8f08fdad05 refactor(module)!: merge AutodiffModule into Module
#5722 2026-09-18 596d9acfbd fix(dispatch): require explicit backend selection and support backend-free builds
#5746 2026-09-21 b9175d2225 fix(core): count a module’s parameters without initializing them
#5739 2026-09-21 bbea651c02 feat(core): a lazy parameter initializes on the device it moves to
#5653 2026-09-22 d580fbca80 fix(backends): correct average-pooling gradients with ceil mode
#5778 2026-09-22 1726b8d47e fix(module): preserve lazy initialization when collecting devices
#5794 2026-09-23 ed0be6f22d test(backend): skip ue4m3 scale accuracy cases where no 8-bit type exists
#5798 2026-09-24 9eadada8ef fix(backend): match PyTorch pooling output size in ceil_mode
#5821 2026-09-25 b9dec8235c fix(module): preserve reparameterizations during validation
#5869 2026-09-28 1a5a4a35ee fix(backend)!: separate device settings queries from initialization
#5916 2026-09-30 8453f26d8e test(backend): run the empty select and roll tests

Tensor API, linalg, signal, einsum (29)

PR Merged Squash commit Content
#5259 2026-09-02 5f74b3938a feat(tensor): add batched SVD decomposition to linalg
#5555 2026-09-03 38f3d8a7d1 fix(tensor): validate matmul batch broadcast in TensorCheck
#5557 2026-09-03 f7f08fd57c feat(tensor): replace no_grad with explicit autodiff conversions
#5559 2026-09-03 a30a9214a7 docs(tensor): fix unfold window formula
#5539 2026-09-04 163b7f442d feat(tensor)!: add multi-axis vector norm variants and update empty max_abs_dims semantics
#5554 2026-09-04 5bd0170384 feat(tensor): add assert_shape! and debug_assert_shape! macros
#5564 2026-09-04 a181e614aa feat(tensor): apply TensorCheck to remainder, powi, powf, hypot, atan2
#5571 2026-09-04 d16f7ba2ed feat(tensor): add is_autodiff and is_tracked state inspection
#5580 2026-09-08 64a69aca89 feat(tensor): validate matmul rank in TensorCheck
#5572 2026-09-08 47ceefde21 feat(linalg)!: extract tensor linalg into burn-linalg extension crate
#5585 2026-09-09 5871c72ef6 fix(tensor): avoid cosine similarity denominator underflow
#5648 2026-09-11 545682e718 fix(linalg): add autotune feature propagation
#5582 2026-09-14 766436c4ef feat(tensor): implement Min and Max scatter/select_assign across backends
#5693 2026-09-16 d6883eb4fa fix(linalg): support negative even lp norm orders
#5583 2026-09-16 10b74fbb3d feat(tensor): add einsum with runtime and macro APIs
#5720 2026-09-17 1c337e0d80 feat(signal)!: extract tensor signal into burn-signal extension crate
#5688 2026-09-18 ac79057b48 feat(tensor): device identity and one entry per physical GPU
#5743 2026-09-21 752e82b252 fix(tensor): fix quiet softmax for negative infinity slices
#5745 2026-09-21 b683dc45cc fix(tensor): preserve f64 precision in degree/radian conversions
#5754 2026-09-21 ab84b8e69f fix(tensor): return NaN for infinite dividends in fmod_scalar
#5790 2026-09-23 1471422107 fix(tensor): keep input dtype on empty slice/repeat/cat results and one_hot_fill
#5804 2026-09-24 b930e39f47 fix(linalg): import Vec in fusion svd for no_std builds
#5827 2026-09-25 e814443c76 fix(tensor): reject quantized inputs in *_like ops and one_hot
#5846 2026-09-25 af8b3306db refactor(tensor): remove into_tiled from the public API
#5838 2026-09-25 e201487165 fix(tensor): make TensorData fields private and validate construction
#5849 2026-09-28 84b9669a59 feat(tensor)!: accept output_size or scale_factor in InterpolateOptions
#5868 2026-09-29 cdeccbcf6c fix(tensor): display quantized tensors as metadata without panicking
#5853 2026-09-30 9adff02644 refactor(tensor)!: move pooling args into options structs with asymmetric padding
#5912 2026-09-30 91d892a4ad fix(tensor): roll shift direction to match torch.roll

Dependencies, build and CI (16)

PR Merged Squash commit Content
#5544 2026-09-01 ef7c981590 chore(deps): reduce unnecessary dependencies
#5669 2026-09-14 8f7a1c435c ci: reduce redundant compilation in macOS tests
#5672 2026-09-14 7832f67a00 fix(deps): update rustls for cargo audit
#5671 2026-09-14 45734c0170 ci: reduce redundant compilation across test suites
#5675 2026-09-14 faa20a1f4a chore(deps)!: make backend tracing opt-in
#5680 2026-09-15 11f9ff07d0 chore(deps): reduce dataframe dataset dependencies with polars-core
#5681 2026-09-15 f6e8251512 ci: reduce examples overhead by skipping coverage setup and enabling caching
#5694 2026-09-16 8116dff453 fix(deps): avoid enabling CubeCL through linalg defaults and respect vision defaults
#5715 2026-09-17 f8534af9e8 chore(deps): trim zip default features and fix burn-dataset nlp feature
#5727 2026-09-18 da2209a9de fix(ci): publish burn-einsum and include it in no-std checks
#5735 2026-09-18 3bd1a6e91f fix(deps): restore the windows crate versions the cube update changed
#5736 2026-09-18 3a93fbfc0f chore: remove recursion limit workarounds and update CubeCL configs
#5776 2026-09-22 d3c8d7c615 chore: bump version to 0.22.0-pre.4
#5777 2026-09-22 9147c11a19 fix(ci): use cargo info to check published crate versions
#5786 2026-09-23 ee211424ac chore(deps): bump actions/checkout from 6 to 7
#5831 2026-09-25 28235bc19e chore(deps): drop unused bincode workspace dependency

Documentation (9)

PR Merged Squash commit Content
#5731 2026-09-18 28bfe74d5f docs(book): update backend extension guides and include maintained example sources
#5748 2026-09-21 70326cfc4b docs: fix typos and grammar in books and API comments
#5765 2026-09-21 f726002c08 docs: clarify guidelines for minor documentation fixes
#5762 2026-09-21 9f1bb5112f docs: update guides and examples for Burn 0.22
#5768 2026-09-21 faec398324 docs: cover remaining 0.22 migration points and refresh stale pages
#5773 2026-09-22 bc2b9832f8 docs: select the package when running examples from the repo root
#5800 2026-09-23 f56efabd89 clarify 0.22 migration guidance and update API examples
#5845 2026-09-25 f5bd4c075c docs(book): document burn.toml runtime configuration
#5863 2026-09-28 6352b35ef4 docs: fix dead links in README and the Burn Book

Other (14)

PR Merged Squash commit Content
#5540 2026-09-02 046a21c436 Refactor/fusion write scope
#5545 2026-09-03 7eaab04251 perf(conv): accumulate direct convolution channels in vector registers on CPU
#5631 2026-09-09 df92ec0750 Feat/storage tiled carrier
#5546 2026-09-09 e80e7981b9 refactor(dataset): back SqliteDataset with Turso instead of rusqlite
#5667 2026-09-14 5685f782e9 fix(burn-dataset): cap MNIST item counts at the split size
#5726 2026-09-18 d0365962df Feat/profiling
#5814 2026-09-24 a9de0f0487 fix(test): align dtype-support expectations with the runtimes
#5816 2026-09-24 ffc4fe0223 fix(examples): correctly format web inference probability labels
#5806 2026-09-24 0b6f7cbf04 fix: remove rl from default features
#5823 2026-09-24 b90efe979d fix(examples): correct WebGPU inference and keep live predictions responsive
#5811 2026-09-24 d85aa0ab39 perf: use grouped convolution for fold4d
#5839 2026-09-25 9ef294317c Qa/tiled storage
#5822 2026-09-28 80d3a024f3 Share the autotune roofline policy and bound fused matmul and reduce tuning
#5894 2026-09-29 a2edbd2663 Refactor/storage tile naming

Privacy

Burn and Tracel AI are trademarks of their respective owners. PRA is not affiliated with Tracel AI.