Companion to the management memo. The previous month is covered by the August technical review.
The repository reviewed is tracel-ai/burn, branch
main, window
[2026-09-01T00:00:00−04:00, 2026-10-01T00:00:00−04:00). The
UTC−4 zone is used, as in the August review. Date of review: 1 October
2026. Snapshot of main: commit
91d892a4ad1e928866b1950b3f99e8a4061191e2 of 30 September,
16:11:23 −04:00.
The set was determined by the committer date of the
squash commit in the first-parent history of main. The
window contains 246 PRs and no direct commits. Counting by UTC would
additionally pull in #5510, which belongs to August in UTC−4, so it is
not counted here.
The review covers all work merged during the period. The memo’s
themes are examined in detail; every other change is expected to appear
as a verified entry in the full-coverage section. Tests were read from
the diff and not run independently, so there are no claims about them
passing. Code references are pinned to revision 91d892a4.
The behaviour of a later main may differ.
Four crates appeared during the period: burn-einsum,
burn-linalg, burn-signal and
pytorch-reader. The number of directories in
crates/ grew from 34 on 31 August to 38 on 30
September.
Linear algebra was moved out of burn-tensor into the new
burn-linalg crate in #5572, squash
commit 47ceefde21.
The
former crates/burn-tensor/src/tensor/linalg/ directory,
present in the 31 August snapshot, was recast as the
LinalgOps backend extension carrying the
#[backend_extension] attribute and moved into crates/burn-linalg/.
In the 30 September snapshot the old directory is gone. The paths
burn_tensor::linalg and
burn_core::tensor::linalg were removed with no deprecation
period. The compatible path burn::linalg is available only
with the linalg feature enabled, and that feature is not
part of default for the facade crate in crates/burn/Cargo.toml.
A host fallback through svd_host::svd_host_data was kept
for LinalgOps::svd, so a backend without native SVD moves
data to the CPU.
Signal processing was moved out to burn-signal in the
same manner in #5720, squash
commit 1c337e0d80.
The new crates/burn-signal/
defines SignalOps as a backend extension. rfft
and irfft are implemented as custom operations, and remote
and capture require register_fft_ops. The old
burn_tensor::signal path was removed. The
signal feature in crates/burn/Cargo.toml
is likewise not enabled by default. The same PR removed the stub for
NdArray, since that backend provides no FFT.
Reading PyTorch files was extracted from burn-store into
the independent pytorch-reader crate in #5656, squash
commit 0fdac31a3d.
The new crates/pytorch-reader/
does not depend on Burn. It uses its own DType and
Tensor with lazy byte reading. Its direct dependencies are
limited to byteorder, crc32fast,
num-traits, serde, tar,
thiserror and zip. burn-store
re-exports the reader as burn_store::pytorch_reader and
provides the bridge::from_pytorch bridge in crates/burn-store/.
The former public path burn_store::nested was removed.
burn-einsum is of a different nature. #5583, squash
commit 10b74fbb3d,
does not move an existing module out. It adds new equation parsing,
tensor-contraction planning, a procedural implementation in crates/burn-derive/,
an API in crates/burn-tensor/
and a separate crates/burn-einsum/.
The einsum dependency in crates/burn-tensor/Cargo.toml
is neither optional nor behind a feature gate, so it is always compiled.
The follow-up #5727, squash
commit da2209a9de,
changed CI only. Publishing of burn-einsum was added to .github/workflows/publish.yml,
and the crate was included in the NO_STD_CRATES set.
All four new crates are registered in .github/workflows/publish.yml
as separate published artifacts.
The foundation for this new split of operations landed in #5520, squash
commit 0e2dc6b631.
burn-backend-extension was split into catalog,
derive, dispatch, extension,
ir and routing in crates/burn-backend-extension/.
The former large set of macros in crates/burn-dispatch/src/macros.rs
was replaced by generation through extensions. It is this
#[backend_extension] mechanism that
burn-linalg and burn-signal then used.
After the rebuild, promotion of the autodiff context had to be
restored. That was done in #5551, squash
commit cf1b6d2a60.
The test file crates/burn/tests/autodiff_context.rs
checks that calling an operation through the new extension mechanism
preserves and correctly promotes the autodiff context. In the 30
September snapshot the file no longer exists at that path: #5647 moved it to
crates/burn-backend-tests/tests/autodiff/context.rs.
The link above points at the revision in which the test was added.
In #5570,
squash commit 54d36fa936,
burn-dispatch stopped depending on the facade crates of the
CubeCL runtimes. The change is fixed in crates/burn-dispatch/Cargo.toml.
In #5694,
squash commit 8116dff453,
the new burn-linalg stopped pulling in
cubecl-backend through its default features. The dependency
boundary is set in crates/burn-linalg/Cargo.toml.
#5722,
squash commit 596d9acfbd,
made burn-flex an optional dependency and replaced the
default_backend configuration with
backend_enabled in crates/burn-dispatch/build.rs.
For a configuration with no execution backend, the uninhabited type
NoBackend was introduced. Calling
DispatchDevice::default() in such a build panics with an
exact message:
No execution backend is enabled. Enable a Burn backend feature such as `flex`, `wgpu`, or `cuda`.
To record a graph without executing it, enable `capture` and use Device::capture()
This change does not rule out builds intended only for recording a
graph. For those, the combination of the capture feature
and Device::capture() remains.
Implicit Flex existed only in the 0.22 pre-releases, from
pre.1 through pre.3. In stable
v0.21.0, burn-dispatch and all backends were
already optional. So #5722 breaks the
configuration of users of the 0.22 pre-releases, but is not a regression
relative to stable 0.21.
#5721,
squash commit 8f08fdad05,
removed the AutodiffModule trait entirely. The
valid() method moved into Module, and
AutodiffModule::from_inner was removed. The
train() method was already in Module; the
change lifted its where Self: AutodiffModule bound and its
implementation through from_inner. A single Rust type could
represent both the training and the validation model before this PR as
well. The substantive new boundary is a different one: a
Module bound alone is no longer enough to guarantee that
autodiff is enabled. The main implementation is in crates/burn-core/src/module/base.rs.
#5557,
squash commit f7f08fd57c,
removed the public Tensor::no_grad() and replaced it with
Tensor::without_autodiff() in crates/burn-tensor/src/tensor/.
inner() became an alias of without_autodiff()
and no longer panics for a tensor without autodiff. The reverse
conversion through from_inner became idempotent thanks to
the new Tensor::autodiff(). This change did not touch the
identically named Module::no_grad() method.
A more dangerous change landed in #5537, squash
commit af55cf4432.
Previously the Module implementation for
Param<Flag> overrode no_grad() through
with_value(false). As a result,
model.no_grad() not only disabled gradients but also
switched off dropout and the updating of accumulated BatchNorm
statistics. After that override was removed, no_grad()
controls tensor gradients only. The former behaviour moved to the new
freeze(). unfreeze() and
set_require_grad_group() were added alongside it. The
change runs through crates/burn-core/src/module/base.rs
and crates/burn-core/src/module/lora.rs.
The no_grad_does_not_disable_a_flag test in
flag.rs checks that no_grad() no longer clears
a module’s mode flag. The
no_grad_only_disables_tensor_gradients test in
dropout.rs checks that the method disables tensor gradients
without putting dropout into validation mode. These tests pin down the
new semantics. They do not confirm compatibility with the old behaviour.
The compiler gives no warning about a change of this kind, so partial
fine-tuning code using Dropout or BatchNorm may keep compiling while
behaving differently.
#5571,
squash commit d16f7ba2ed,
complements this rebuild but is not itself a breaking change.
Tensor::is_autodiff() and Tensor::is_tracked()
were added to the tensor API. grad() and
grad_remove() now return None when autodiff is
off. #[must_use] was added to inner(),
without_autodiff() and autodiff(). The changes
are in crates/burn-tensor/src/tensor/.
#5521,
squash commit 7b815b01bd,
changed how padding is represented in ConvOptions. Instead
of a single symmetric value it stores (before, after)
pairs. new() is intended for the symmetric case and
new_with_padding() for the asymmetric one.
PaddedConvOptions was deprecated. The change runs through
crates/burn-backend/src/backend/ops/.
As of the end of September, Conv3d in the IR still accepts
symmetric padding only, so support for asymmetry across the whole stack
is unfinished.
#5853,
squash commit 9adff02644,
moved max_pool1d, max_pool2d,
avg_pool1d, avg_pool2d and the
max_pool*_with_indices variants onto the
MaxPoolOptions and AvgPoolOptions structs.
Padding is given as start and end pairs through
with_padding_pairs(). The defaults are aligned with
PyTorch: stride equals kernel, dilation equals one,
ceil_mode is off, count_include_pad is on.
Asymmetric padding was moved from the nn modules down to
the functional level in crates/burn-tensor/src/tensor/.
There are no avg_pool*_with_indices variants. Migration
examples were added to burn-book/src/migrating-to-0.22.md.
#5849,
squash commit 84b9669a59,
replaced the interpolate(x, output_size, options) call with
interpolate(x, options). The target size now lives in
options.output_size, or is computed from
options.scale_factor. The public tensor API changed in crates/burn-tensor/src/tensor/,
but the backend trait and the IR were not touched.
#5539,
squash commit 163b7f442d,
added multi-axis variants of the vector norm. At the same time it
changed the narrow max_abs_dims(&[]) case. An empty
axis list previously returned the original tensor; it now returns
self.abs(), which matches ONNX with
noop_with_empty_axes=true. At the time of the change the
code was in crates/burn-tensor/src/tensor/linalg/vector_norm.rs,
before linear algebra was moved out into burn-linalg.
#5838,
squash commit e201487165,
made the fields of TensorData private in crates/burn-std/src/data/tensor/base.rs.
Access is provided through shape(), dtype(),
bytes(), into_parts() and
with_bytes_mut(). try_from_bytes returns
DataError::InvalidByteLength when the buffer length does
not match the shape and type. from_bytes panics on such
input. Deserialisation was moved onto try_from_bytes, and
the constructors check for overflow when counting elements. The length
of quantised data is not yet validated, which was filed as the open
#5836.
#5846,
squash commit af8b3306db,
removed into_tiled from Tensor, the backend
trait, the IR, fusion and the router. The file crates/burn-std/src/tensor/tiled.rs
was deleted. The practical significance of this removal is low. The API
was added in #5839, squash
commit 9ef294317c,
on the morning of 25 September and removed roughly two and a half hours
later. It was present in no tag.
The most visible reduction comes from #5722, squash
commit 596d9acfbd.
Building the facade burn without a backend feature no
longer compiles the Flex CPU backend. The configuration is set in crates/burn/Cargo.toml.
#5675,
squash commit faa20a1f4a,
removed tracing from the default features of
burn-autodiff and burn-fusion. Their
default = ["std", "tracing"] was replaced with
default = ["std"]. For dependencies in
burn-backend, burn-dataset and
burn-vision, the weak syntax cubecl?/tracing
and burn-std?/tracing is used. The change is concentrated
in the Cargo manifests of the crates concerned and is described in crates/burn/src/lib.rs.
No instrumentation code was added or removed.
#5806,
squash commit 0b6f7cbf04,
removed rl from the default features of burn
and burn-train in crates/burn/Cargo.toml
and crates/burn-train/Cargo.toml.
This is a fix for regression #5769, not a new subsystem. In
burn-train the feature was already part of default in
v0.21.0. In the facade burn it became default
in May 2026 in #5012, squash commit 6b721d887c.
#5544,
squash commit ef7c981590,
dropped unused dependencies from the Cargo manifests of a number of
crates and examples. In crates/burn-cubecl-fusion/Cargo.toml,
rmp-serde was replaced with ciborium.
#5680,
squash commit 11f9ff07d0,
moved the dataframe feature in crates/burn-dataset/Cargo.toml
from the full polars to polars-core. The
default set for that capability no longer includes the lazy engine or
the file formats.
#5715,
squash commit f8534af9e8,
disabled the default features of the zip workspace
dependency in the root Cargo.toml.
By the PR author’s measurement, the dependency tree of
pytorch-reader shrank from 76 crates to 33. That
measurement was not independently verified. The same PR fixed the
nlp feature of burn-dataset.
#5831,
squash commit 28235bc19e,
removed bincode from the workspace dependencies in the root
Cargo.toml.
The dependency remains in the lock file transitively through Polars.
The guide burn-book/src/migrating-to-0.22.md
was created in #5762, squash
commit 9f1bb5112f.
The same PR updated the README, the building-blocks sections, the
contributor book and the notebooks.
The guide was extended by #5768, squash
commit faec398324,
and #5800,
squash commit f56efabd89.
As of 30 September the document contained 17 thematic sections. It kept
being changed after the v0.22.0-pre.4 tag. After the tag,
new incompatibilities were introduced by #5838, #5849, #5869 and #5853.
#5800 itself
was a documentation rework, #5806 fixed features, and #5821 fixed
validation.
#5731,
squash commit 28bfe74d5f,
replaced standalone code fragments in the book with mdBook includes from
maintained examples. As a result, the examples included in the book
compile together with the original examples. The change is in burn-book/.
#5845,
squash commit f5bd4c075c,
added the page burn-book/src/advanced/runtime-configuration.md
about burn.toml. This is documentation of an existing
mechanism. Support for burn.toml was added back on 23 April
2026 in #4864.
v0.22.0-pre.4
did not complete the release transitionThe lightweight tag v0.22.0-pre.4 was placed on 22
September on commit 9147c11a
from #5777,
which fixed the check of published versions through
cargo info. The tag points not at the version-bump commit
but at its direct descendant.
The bump itself from 0.22.0-pre.3 to
0.22.0-pre.4 was done in #5776, squash
commit d3c8d7c615.
That PR also replaced the git-rev dependencies on CubeCL and Cubek in
the root Cargo.toml
with the published versions =0.11.0-pre.4 and
=0.3.0-pre.4. The workspace was thereby brought into a
state fit for publishing by tag through .github/workflows/publish.yml.
As early as 23 September, #5797, squash
commit 0fa63b3e45,
returned main to dependencies by git revision. In the Cargo.toml
of snapshot 91d892a4, CubeCL is again specified through
git and rev. So main at the end
of September was not in a publishable configuration.
The last stable tag remained v0.21.0 of 7 May 2026. The
0.22 pre-releases came out on 29 July, 10 August, 25 August and 22
September. No stable 0.22.0 was released in September.
Ten September PRs carried the ! marker in their title:
#5521, #5539, #5572, #5647, #5675, #5720, #5721, #5849, #5853 and
#5869.
At least five further PRs changed the public API without such a
marker: #5557,
#5537, #5838, #5846 and #5722. Of these,
#5537 is particularly substantive, because it silently changes the
semantics of Module::no_grad() and causes no compilation
error. #5722 affects users of the 0.22 pre-releases but does not worsen
behaviour relative to stable 0.21. #5846 formally removed a public
method, but into_tiled lived for about two and a half hours
on 25 September and reached no tag.
Reading of .pt and .pth was extracted from
burn-store into the standalone pytorch-reader
crate in #5656, squash
commit 0fdac31a3d.
The new crates/pytorch-reader/
depends on neither Burn nor PyTorch. The full list of its direct
dependencies is limited to byteorder,
crc32fast, num-traits, serde,
tar, thiserror and zip. The crate
name carries no burn- prefix, and its description positions
the component outright as a reader independent of both systems.
The crate defines its own DType and Tensor.
Tensor::read() materialises bytes only on request.
burn-store pulls the reader in as an optional dependency
behind the pytorch feature and re-exports it as
burn_store::pytorch_reader. The bridge crates/burn-store/src/bridge.rs
provides bridge::from_pytorch, which wraps a read tensor in
a deferred burn_pack::Tensor. The former public path
burn_store::nested was removed in the same #5656, squash
commit 0fdac31a3d.
A separate publish-pytorch-reader job appeared in .github/workflows/publish.yml.
The publish-burn-store job depends on it, so that the new
standalone dependency is published first. Version pytorch-reader 0.22.0-pre.4
was published on 22 September 2026. The crate’s documentation was then
synchronised between crates/pytorch-reader/README.md
and crates/pytorch-reader/src/lib.rs
in #5747,
squash commit b3984918b5.
Before the crate was extracted, the reader was substantially reworked
in #5593,
squash commit 33df298f84.
The 570-line lazy_data.rs was removed and work with
deferred data moved into storage.rs. The ZIP is now opened
once, entries are looked up relative to the directory containing
data.pkl, and allocations whose sizes come from an
untrusted file are bounded. The changes are concentrated in crates/pytorch-reader/src/.
Two outcomes must not be credited to #5593 as new
capabilities. TAR support and the rejection of big-endian existed before
squash commit 33df298f84. This is visible in its parent
revision of crates/burn-store/src/pytorch/reader.rs, lines
461–477 and 873. The PR rewrote these paths and pinned them with
fixtures, but did not introduce the corresponding behaviour for the
first time.
The claimed support for pickle protocols 0–5 likewise does not mean
full coverage of every protocol. After #5593, squash
commit 33df298f84,
the opcodes EXT1, EXT2, EXT4,
NEXT_BUFFER and READONLY_BUFFER are still
rejected. The practical improvement is that protocol 4 no longer fails
immediately on FRAME, not that the whole pickle
specification is implemented.
#5700,
squash commit 96e2ec1e12,
removed clone_unsafely from enum deserialisation. The
function copied a Serde visitor through
ptr::copy_nonoverlapping even though its type required
neither Copy nor Clone. A visitor holding heap
data could be freed twice. This is not a September regression. The code
appeared in March 2024 in #1436, commit
0138e16af6, and in June 2026 moved into
burn-store/src/nested/de.rs. Ahead of the standalone
reader’s publication the fix landed in crates/pytorch-reader/src/nested/de.rs.
About ten Deserializer method implementations that called
unimplemented!() were replaced with returned errors.
The file crates/pytorch-reader/tests/nested_de_miri.rs
was added in #5700, squash
commit 96e2ec1e12.
It is intended to be run under Miri and checks that enum deserialisation
no longer copies or frees the visitor in an unsafe way.
#5737,
squash commit d79dab2d12,
changed ownership of the source file in crates/pytorch-reader/src/storage.rs.
LegacySource no longer stores only a path and no longer
reopens the file for every storage. An open descriptor is now kept, a
duplicate of it is created for reading, and bytes are read at explicit
offsets. Deleting the file after it has been opened no longer breaks
subsequent reads. Replacing the file through rename no longer mixes the
directory of the new archive with the bytes of the old file. The PR #5752, closed
without merging, was superseded by this fix, so it has no squash commit
in main.
#5728,
squash commit e68583e42f,
stopped turning every uninterpreted pickle object into
PickleValue::None. That behaviour let
load_config silently substitute defaults such as
0, an empty string or false for numpy scalars,
torch.dtype, torch.device and tensors.
NestedValue::Unsupported(String), which preserves the name
of the unsupported Python type, appeared in crates/pytorch-reader/src/nested/.
#5749,
squash commit 4d26deac29,
moved PickleError and PytorchError onto
thiserror with #[from] and
#[source]. A nested io::Error is now reachable
through the standard Error::source() chain. The
detect_format function in crates/pytorch-reader/src/reader.rs
requires a valid pickle header and rejects SafeTensors separately. This
removes the ambiguity in which the first length byte of a SafeTensors
JSON header could coincide with a pickle opcode.
#5756,
squash commit a36db79bc0,
added compatibility with torch.save(compute_crc32=False).
In that mode PyTorch writes a zero CRC-32 into the ZIP headers, and the
zip crate treated it as an incorrect checksum. The
archive-source code in crates/pytorch-reader/src/
now treats zero as the absence of a checksum, which matches the
behaviour of torch.load.
#5664,
squash commit 317fb474fc,
extended extract_tensors in crates/pytorch-reader/src/.
The traversal previously descended only through dictionaries, so tensors
inside a list or tuple were lost without an
error. Elements now receive indexed names. For example,
{"weights": [w1, w2]} becomes weights.0 and
weights.1. This matches the naming of
Vec<Module> in Burn and of nn.ModuleList
in a state_dict. The PR closes the regression from issue
#5595.
#5714,
squash commit b2faf4f336,
narrowed the locking scope of ZipSource in crates/pytorch-reader/src/storage.rs.
The Mutex<ZipArchive> was previously held for the
whole duration of a storage read, so parallel
Tensor::read() calls effectively ran sequentially. The lock
is now needed only to look up an entry. In the author’s measurement, for
64 tensors of one million f32 each on eight threads, the
time fell from 24 ms to 6.3 ms. The single-threaded result did not
change. This measurement was not independently reproduced.
The first chain: #5714, squash
commit b2faf4f336,
→ #5737,
squash commit d79dab2d12.
For the sake of parallel reading, #5714 left a second open-by-path of
the file on Windows. If a rename over the file happened between the two
opens, the ZIP directory belonged to the new file while positional reads
could address the old one. #5737 replaced the repeated
File::open with duplication of the already open descriptor
through try_clone(). The defect was introduced and fixed
within three days. Both changes run through crates/pytorch-reader/src/storage.rs.
The second chain: #5728, squash
commit e68583e42f,
→ #5700,
squash commit 96e2ec1e12,
→ #5764,
squash commit 03e2da00c5.
#5728 added NestedValue::Unsupported. #5700 then replaced
unimplemented!("deserialize_any") with an exhaustive
match, but added no arm for the new variant. The result in
crates/pytorch-reader/src/nested/de.rs
was compilation error E0004, not merely an unreachable runtime error.
The nested module is included unconditionally, there was no
wildcard, the type was not #[non_exhaustive], and there was
no feature gate either. #5764 added the missing arm 84 minutes
later.
#5766,
squash commit 855fe4e812,
added extraction of weights from pickles created through
torch.save(model). In crates/pytorch-reader/src/
the BUILD operation recognises an object whose state
contains the _parameters, _buffers and
_modules tables and interprets it as a module’s
state_dict(). Parameter names keep their familiar form such
as layer1.weight.
The same #5766, squash
commit 855fe4e812,
added reading of protocol 2 set and frozenset
as lists and fixed BUILD memoisation. A repeated reference
to a module could previously return its state from before
BUILD, which made nested tensors disappear.
The limits of the capability are recorded in crates/pytorch-reader/src/lib.rs,
crates/pytorch-reader/README.md
and burn-book/src/saving-and-loading.md:
None are not read._parameters and
_buffers are not extracted.get_extra_state() is not executed.forward method and the model’s computational
behaviour are not restored.So #5766,
squash commit 855fe4e812,
removes the mandatory prior export through state_dict() but
remains a weight-transfer mechanism. It is not loading of an executable
Python model. In line with the new boundary, the section declaring
whole-model saving inherently wrong was removed from the book.
#5489,
squash commit 137b73a5ae,
fixed a dangerous ordering in
SafetensorsStore::collect_from in crates/burn-store/.
The former path used safetensors::serialize_to_file, which
truncated the target file at the start of the write. Tensors were
materialised from the device only after that. A readback error left a
truncated file where the last valid checkpoint had been.
After #5489 the new container is first built in full in an adjacent
temporary file and only then renamed into the destination. The shared
implementation was moved into the public burn_pack::AtomicFile.
The process-level guarantee is that after an error, a panic or a stop,
either the complete new container or the previous file remains.
Review of #5489, squash
commit 137b73a5ae,
surfaced three further problems in crates/burn-pack/src/atomic.rs.
First, sync_all on Windows requires a handle with the
GENERIC_WRITE right; without it every file save failed with
Access is denied. Second, commit no longer
accepts an arbitrary path, which makes it impossible to publish a
scratch file outside the predetermined destination. Third,
rename_onto carries the access mode of the existing
destination over to the temporary file; without that, a checkpoint with
mode 0600 could be replaced by a file with mode
0644. Ownership and hard links are not preserved across the
rename, since publication creates a new inode. That boundary is
documented.
#5781,
squash commit d17a1ff8b8,
closed a race in overwrite(false) mode in crates/burn-pack/src/atomic.rs.
The absence of the destination was previously checked before tensors
were materialised. A file created by another process during the save was
then silently overwritten by the rename.
The new AtomicFile::commit_new publishes the scratch
file by creating a hard link at the target path. An already existing
path returns Error::AlreadyExists, so the check and the
publication become a single file operation. On a file system without
hard-link support, a fallback to rename after confirming the
destination’s absence is used. That is a weaker guarantee with a race
window. The preliminary check was kept as a fast path, but now uses
symlink_metadata so that a dangling symbolic link too is
detected before tensors are read.
#5832,
squash commit 18ac73b70b,
made atomicity the default behaviour in crates/burn-pack/src/writer.rs.
Writer::write_to_file now uses the atomic path. The former
direct behaviour is available through
write_to_file_in_place. The separate
write_to_file_atomic method was removed.
This is a source-breaking change, but it affects an unpublished API.
burn-pack appeared only on 16 June 2026 in commit
eda9663665 and is absent from stable v0.21.0.
Besides, #5832, squash
commit 18ac73b70b,
landed on 28 September. That is later than the
v0.22.0-pre.4 tag of 22 September, so atomicity by default
is not in even that pre-release.
#5490,
squash commit c1d6874cb9,
fixed Applier::apply_tensor in crates/burn-store/src/.
The shape was checked previously, but not the kind of dtype. Float data
could, for example, be applied to an Int parameter, after which the
report announced a successful application and the error arose later in
try_to_vec. The ApplyError::DTypeMismatch
variant already existed, but no path constructed it.
The kind of dtype is now checked before materialisation, so a
rejected tensor is not read from disk at all. The width of the type is
deliberately not required to match. Half precision may be loaded as
F16, and PyTorch int64 as I64.
For a Bool parameter only U8, U32 and
Bool are allowed, whereas the former code accepted any
integer dtype.
#5576,
squash commit a7be13739c,
added a check of each tensor’s byte length through
validate_tensor_byte_len to crates/burn-pack/.
The expected size is computed from shape and dtype. Before this the
reader checked the maximum offset in the container but could accept an
individual tensor with an inconsistent size.
#5838,
squash commit e201487165,
carried the same check through to the burn-store bridge.
The bridge::tensor_data function in crates/burn-store/src/bridge.rs
now calls TensorData::try_from_bytes from crates/burn-std/src/data/tensor/base.rs.
An inconsistent length returns
PackError::TensorBytesSizeMismatch instead of panicking.
For quantised data this check is not yet performed, which was filed as
open issue #5836.
#5558,
squash commit ef521c0896,
changed PathFilter::with_regex in crates/burn-store/src/.
The former if let Ok(regex) silently ignored an invalid
pattern, so the filter carried on with a different set of paths.
Construction now calls expect("Invalid regex pattern"). The
boundary of the decision is that this is an explicit panic at
configuration time, not a returned Result.
#5750,
squash commit bd9a529bbc,
added map_indices_contiguous_except(regex) to crates/burn-store/src/.
The former map_indices_contiguous renumbered lists on an
all-or-nothing basis. The new variant allows matching paths to be
excluded. The PR’s third commit fixed an adjacent defect: builders that
change path names now clear an already populated tensor cache, so the
result does not depend on call order.
#5666,
squash commit 52c0535842,
rewrote the streaming-write check in crates/burn-pack/.
Instead of a global counting allocator, each test tensor holds
bytes::Bytes and registers its own index on release. The
check records that individual tensors’ data is released as the streaming
write proceeds, rather than held until the whole container is finished.
Coverage was extended to the then-existing
write_to_file_atomic and to into_bytes. Global
state and unsafe were removed from the test code.
#5489
appeared in the August review as WIP with branch commit
7263ca5cb6 of 31 August and integration expected on 1
September. The PR did indeed land on 1 September as squash commit 137b73a5ae.
In September this line continued with no-clobber publication in #5781, squash
commit d17a1ff8b8,
and atomicity by default in #5832, squash
commit 18ac73b70b.
The main implementation of the whole chain is in crates/burn-pack/src/atomic.rs
and crates/burn-pack/src/writer.rs.
August’s check for missing and truncated tensor data also received
two continuations. #5576, squash
commit a7be13739c,
checks the size inside burn-pack. #5838, squash
commit e201487165,
moves bridge::tensor_data from a panic to a typed error.
The adjacent #5490, squash
commit c1d6874cb9,
closes the dtype-kind mismatch that previously passed as a successful
application.
#5546,
squash commit e80e7981b9,
concerns not model checkpoints but dataset storage.
SqliteDataset in crates/burn-dataset/
was moved from rusqlite, r2d2_sqlite and
serde_rusqlite to the Turso Rust engine. The public names
were kept, and database files remain compatible in both directions.
The sqlite feature was removed from the default set of
burn-dataset. In crates/burn-dataset/Cargo.toml
the configuration changed from default = ["sqlite-bundled"]
to an empty default. The facade feature burn/dataset no
longer provides SqliteDataset on its own;
burn/sqlite is required for it. The change is recorded as a
migration entry in burn-book/src/migrating-to-0.22.md.
The Turso dependency in #5546, squash
commit e80e7981b9,
is pinned to the exact pre-release version =0.8.0-pre.8. At
the same time turso::Error is present in the public
SqliteDatasetError, so an external pre-release is part of
the API surface. The move also did not eliminate the dependency on C:
turso_core transitively uses simsimd, which
compiles C code unconditionally.
The new path uses Turso in WAL mode, groups inserts into transactions
and performs an explicit WAL checkpoint in set_completed,
checking the returned result row. The max(row_id) query was
unwrapped from coalesce, since the former form degraded
into a full scan. In the author’s measurement on a split of 560 thousand
rows, the time fell from 86 ms to 5 ms. These changes are in crates/burn-dataset/src/,
and the figures quoted were not independently verified.
At least twenty changes landed in main during September
that bring Burn’s behaviour on NaN, infinities and range boundaries into
line with what PyTorch does. Twelve of them concern NaN and signed
infinities directly. The practical point of such work is that before it
a corrupted number could disappear along the way: relu,
clamp and max pooling returned the boundary instead of NaN,
and abs() gave a plausible finite gradient through an
incorrect sign(NaN). Training carried on, and there was no
sign of a failure.
The claim that this is a series rests on three independent observations, each visible in the repository itself.
First. The changes reference shared tracking issues with numbered
items. In #5662, squash
commit 11f8c072e1,
the body states Refs #5615 (item 1). In #5663, squash
commit c0a8b0d13d,
it states Refs #5615 (item 6). The remaining links go to
individual issues: Fixes #5609 in #5658,
Fixes #5659 in #5718,
Fixes #5610 in #5652,
Fixes #5602 in #5661 and #5808,
Fixes #5791 in #5798,
Fixes #5788 in #5790,
Fixes #5789 in #5827.
Second. In a single test file, crates/burn-backend-tests/tests/tensor/float/ops/clamp.rs,
four different contributors passed coverage to one another in sequence
through compilation conditions:
52c101f455,
added clamp_nan_propagation under
#[cfg(any(feature = "flex", feature = "ndarray"))] and left
a comment noting that two-sided clamp on the CubeCL backends still turns
NaN into the boundary, with a note about checking on CUDA.274cb46cb9,
fixed float_clamp in burn-cubecl to use
comparison with selection and lifted that gate for ordinary precision.
The comment retained the note that f64 still fails.dae47a34b8,
added clamp_nan_bound_propagation and
clamp_keeps_negative_zero for Flex.6ade1047f8,
added variants for the NdArray SIMD paths.56584ee48b,
lifted the gate from the signed-zero test and extended the clamp
boundary check to any(flex, cube).Third. The requirement is fixed in the code as a backend contract,
not only in a test. #5665, squash
commit 66a8a5ff8a,
added to crates/burn-backend/src/backend/ops/tensor.rs
the statement that sign(NaN) == 0 is part of the
operation’s contract and that every backend implementation is obliged to
honour it. The commit body states how the bug was found: by differential
comparison against LibTorch, where x.log().abs().backward()
gave -2.0 instead of -0.0. It also names the
source of the divergence — a mismatched trait bound in an earlier
refactoring, not a deliberate decision.
In most cases the reference is named outright in the code or in the
commit body. In crates/burn-flex/src/ops/pool.rs
the condition was supplemented with an is_nan check and a
note about matching PyTorch. #5652 speaks of
using the remainder with a single modulus operation, as PyTorch does. #5912 refers to
torch.roll, and #5798 to
PyTorch’s pooling output size.
The disappearance of NaN in Flex operations was closed in #5662 for max
pooling and in #5658, squash
commit 52c101f455,
for relu, clamp_min and
clamp_max. The test
test_max_pool2d_with_indices_nan_propagation in crates/burn-backend-tests/tests/tensor/float/module/maxpool2d.rs
checks both the value itself and the index of the last NaN for F32, F64,
F16 and BF16.
#5717
closed a case that was not an incorrect result but a process crash:
f32::clamp panics if a boundary is NaN. #5718 and #5909 extended
the correct clamp behaviour to the CubeCL backends, and #5733 to
NdArray.
#5743,
squash commit 752e82b252,
fixed quiet_softmax on a fully masked slice, where
subtracting minus infinity from minus infinity gave NaN. In crates/burn-tensor/src/tensor/activation/base.rs
the maximum over a slice is now replaced with zero if all values equal
minus infinity. The state of the former coverage deserves separate
mention: the existing test_quiet_softmax_grad test was
replaced in this PR, because it did not call quiet_softmax
at all.
#5585,
squash commit 5871c72ef6,
removed a NaN in the cosine similarity of two zero vectors. The lower
bound was previously applied to the product of the norms, so the square
of a small epsilon went to zero in f32. Each vector is now normalised
separately.
#5754,
squash commit ab84b8e69f,
fixed fmod_scalar, which returned the input unchanged for
an infinite divisor, so that the remainder of infinity divided by
infinity gave infinity instead of NaN.
Neighbouring edge cases were closed in #5629 (overflow
of an eight-bit type in a boolean select_assign, replaced
with a bitwise OR), #5661 and #5808 (masking
the shift amount by the type’s width), #5738 (unit axes
in the contiguity check, which sent data down the copy path and returned
incorrect results on the CPU backend), #5790 (the result
type of empty slice, repeat, cat
and one_hot_fill), #5652 (the
remainder, with tests in crates/burn-backend-tests/tests/tensor/float/ops/remainder.rs).
#5912,
squash commit 91d892a4ad,
stands out from the rest in that it changes the observable result of a
public API. unchecked_roll_dim now cuts the tensor at
size - shift rather than at shift. The
correspondence table in burn-book/src/building-blocks/tensor.md
promised that Tensor::roll matches torch.roll
before this change too, which means the code diverged from the project’s
own documentation. The former expected values in crates/burn-backend-tests/tests/tensor/int/ops/roll.rs
encoded the opposite direction and were rewritten in the same PR, which
also added a one-dimensional case that breaks the symmetry. In burn-book/src/migrating-to-0.22.md
at the 30 September snapshot, roll is not mentioned, so a
user following the migration guide will not learn of this change.
#5798,
squash commit 9eadada8ef,
brought the pooling output size under ceil_mode in line
with PyTorch’s behaviour: the last window that starts in the trailing
padding is discarded. calculate_pool_output_size in crates/burn-backend/src/backend/ops/modules/conv.rs
and pool_output_size in crates/burn-flex/src/ops/pool.rs
were changed, and an overflowing subtraction on an unsigned type was
eliminated along the way.
The superficially similar #5653, squash
commit d580fbca80,
cannot be counted as part of PyTorch parity. Neither PyTorch nor a
tracking issue is mentioned in its body. In substance it reconciles the
forward and backward passes with each other: the backward pass of
average pooling divided by the constant kernel volume, whereas the
forward pass divided by the actual window size. The correct formulation
is that average-pooling gradients under ceil_mode were
incorrect.
The line is unfinished, and that is visible from the compilation
conditions remaining in the tests at the 30 September snapshot. In crates/burn-backend-tests/tests/tensor/float/module/maxpool2d.rs
a comment is kept stating that the CubeCL and NdArray backends still
skip a NaN if it is not the first value. The test
clamp_min_max_nan_propagation_f64 in crates/burn-backend-tests/tests/tensor/float/ops/clamp.rs
remains under any(flex, ndarray). The same gate sits on the
double-precision variant in crates/burn-backend-tests/tests/tensor/float/activation/relu.rs.
The observable confirmation that this line is finished will be the
lifting of those gates and the closing of issues #5615 and #5602.
A parallel line moves the failure from the backend to the public API boundary, where the message is the same for every implementation.
#5580,
squash commit 64a69aca89,
added a rank check in matrix multiplication. Before it the behaviour
depended on the backend: NdArray failed with an overflowing subtraction,
LibTorch returned a rank-zero scalar product, and for CubeCL no valid
input existed. #5555, squash
commit 38f3d8a7d1,
added a compatibility check of batch dimensions under broadcasting and
references an issue opened long before this period. #5564, squash
commit a181e614aa,
extended the checks to remainder, powi,
powf, hypot and atan2 and fixed
the operation identifier in the error text.
#5554,
squash commit 5bd0170384,
goes further and moves some of these checks to compile time. The macros
assert_shape!, debug_assert_shape! and
unpack_shape! are applied in cross-entropy, cosine
embedding, CTC, RNNT, positional encoding and the attention mechanisms,
so a rank mismatch becomes a build error.
#5860,
squash commit c0878a0ab0,
added a finiteness check for scalar parameters in the CELU, Hardtanh,
Dropout and GaussianNoise configurations. #5827, squash
commit e814443c76,
replaced an opaque message about an unexpected data type with an
explicit rejection of quantised input in *_like-style
operations and in one_hot. The byte-length check from #5838 is
described in the section on writing
checkpoints.
Three metrics that can be used when saving a model and for early stopping were computed incorrectly. This is not noise but a systematic bias, and in two cases it is reproduced in the tests’ reference values.
#5578,
squash commit 2646751f3b,
closed two layers of the same bug in accuracy. The value inside a batch
was divided by the number of valid examples, but the full batch size,
padding included, was passed into the epoch accumulation. The epoch
average came out biased. Separately, a fully padded batch produced a
division of zero by zero and spoiled the whole epoch’s result with NaN.
crates/burn-train/src/metric/acc.rs
and crates/burn-train/src/metric/top_k_acc.rs
were fixed, and crates/burn-train/src/metric/state.rs
pins down the rule that an update with a zero counter records the
current value and does not change the accumulated one. The tests
test_accuracy_epoch_aggregation_excludes_padding and
test_fully_padded_batch_does_not_poison_epoch_accuracy
check, respectively, the weight of padding in the epoch average and the
absence of epoch spoiling by a fully padded batch.
#5684,
squash commit fb90882e3e,
fixed average precision in the presence of tied scores. A threshold
includes every example with the same score, so each positive example
within a group must receive the precision computed at the end of the
group. This is implemented through a reversal, a cumulative minimum and
a reverse reversal. Direct confirmation of the scale of the divergence
is in the PR itself: the expected value of the
multilabel_micro reference case was changed from
0.5918017848017848 to 0.550135118149824. That
is, the former numbers diverged from scikit-learn’s
average_precision_score whenever the scores contained ties.
The new test checks that the result is independent of the order of
examples and of the split into batches.
#5900,
squash commit 58cc57d767,
rewrote AUROC. The former implementation built a pairwise comparison of
all examples with materialisation of an examples-by-examples-by-classes
tensor. The new one sorts the scores and walks groups of equal values in
widened integer arithmetic. A test_auroc_large_epoch test
on twenty thousand examples was added. It matters here not to confuse
the rewrite with a change of result: the numeric semantics of NaN are
preserved exactly, since the counts of positive and negative examples
are computed before NaN values are discarded, and pairs containing NaN
remain in the denominator, granting neither a win nor a tie. That is
precisely what the test_auroc_nan_keeps_pairwise_semantics
test pins down. Ties are now determined by exact equality, so positive
and negative zero fall into the same group.
Alongside, two loss functions were fixed where padding and smoothing
produced incorrect normalisation. #5569, squash
commit 1bceb880ec,
eliminated the dilution of cross-entropy: the mask removed padding
tokens from the numerator, but the denominator remained the full batch.
A shared averaging function that applies the mask to the normaliser too
was introduced in crates/burn-nn/src/loss/cross_entropy.rs.
A fully padded batch returns NaN by documented design. The
assert_padding_invariant test compares a padded batch with
an equivalent unpadded one across all combinations of weights, smoothing
and logits. #5892, squash
commit 9e2eefff24,
added clamping of probabilities in the label-smoothing path, which
previously took a logarithm without a bound and gave infinity at zero
probability.
#5875,
squash commit 440a9e2299,
moved dropout in multi-head attention from the scores before softmax to
the weights after it. A warning was added to the returned struct that
during training the rows of weights may not sum to one. The
attention_dropout_applies_after_softmax test checks that
with equal logits every weight after inverted dropout equals zero or
one.
The connection with August here is direct and changes the character of the theme. August fixed the moment at which a metric is read: model saving and early stopping consulted metrics before the computation had finished, plus, separately, Dice aggregation and a crash when obtaining the final ROUGE-L. September fixes the formula itself. There is no continuation on Dice or ROUGE-L in this window.
#5729,
squash commit 288b697e89,
eliminated two independent defects. The first was that an unfinished
gradient-accumulation window at the end of an epoch was simply
discarded. A progress-completion state and a condition under which the
accumulated value is applied were added. The second defect concerned the
learning-rate schedule: the scheduler was stepped per batch rather than
per optimiser update. In the ordinary and multi-device strategies this
meant advancing the schedule N times faster when accumulating over N
batches. In the distributed strategy a loop over participants called the
scheduler step once per participant, so a single optimiser update took N
times the number of participants schedule steps. The loop was replaced
by incrementing the iteration counter by the number of participants, and
the single schedule-step call was moved immediately before the optimiser
step in all three strategies. A check that the number of accumulations
is positive was added. A new test in crates/burn-train/tests/gradient_accumulation.rs
uses a counting scheduler and checks the number of its calls and that
the partial window is applied exactly once.
#5618,
squash commit 2440750e19,
closed a defect that prevented training with gradient checkpointing in
balanced mode from reaching a second step. The module optimiser took a
parameter off the autodiff tape and returned it, assigning the default
checkpointing strategy. After the first step the updated parameters
ended up on the default strategy while the inputs and frozen parameters
stayed on balanced, and the next forward pass refused to combine them.
The same showed up in L-BFGS when unrolling the vector back. Following
review, a parameter context was introduced that preserves the
require-gradient flag, the distributed flag and the strategy.
#5821,
squash commit b9dec8235c,
continues August’s reparameterisation theme. Obtaining a validation copy
of a model previously folded the adapters into the weights and lost
their structure. The validation copy now preserves the
reparameterisation, quantised bases included, and the folding was moved
into an explicit materialisation operation. Export of adapters only,
through a parameter group selected by a regular expression, was added.
The changes affect crates/burn-core/src/module/lora.rs,
and the test is in crates/burn-core/tests/reparameterization.rs.
The corresponding paragraph in the project’s book, stating that a
snapshot folds adapters and discards checkpointing strategies, was
replaced.
#5799,
squash commit 2e5d2152be,
changed how restoration of optimiser state is checked. Instead of a
byte-for-byte comparison of the record, functional equivalence is
checked, that is, that the step after restoration matches the step
without it. The check was applied to Adagrad, Adam, AdamW, Adan, Lion,
Muon, RMSProp and SGD.
#5647,
squash commit 6212b88fe0,
is marked as breaking. The fields of an autodiff tensor became
crate-internal, a node guard object was introduced in crates/burn-autodiff/src/ops/base.rs,
and preparation of a backward operation takes such guard objects and
releases them only after the child has been registered. The effect is
that an abandoned graph correctly gives up its buffers. Users of custom
backends need to move to accessor methods instead of touching fields
directly, and that is recorded in the migration guide. The test
test_mm_reclaims_abandoned_graph_buffers_after_unrelated_backward
in crates/burn-backend-tests/tests/autodiff/memory_management.rs
is marked as a best-effort check: it allows up to sixty-four attempts
and is excluded for NdArray. It follows that the release may be
deferred, and the test gives no strict guarantee about when memory is
returned.
#5645,
squash commit 68a92fdfc9,
closed a silent bug. Before it, a repeated backward pass over an already
consumed tape could silently produce an incorrect parameter update. The
name of the new test says exactly that:
consumed_graph_is_rejected_before_an_incorrect_parameter_update
in crates/burn-backend-tests/tests/autodiff/graph_reuse.rs.
Such a call is now rejected, and the provenance check happens before
additional steps are consumed, so fresh branches remain usable. A leaf
flag appeared on the parent node, thanks to which parameters remain
valid points at which to obtain a gradient. Documentation was added in
crates/burn-tensor/src/tensor/api/autodiff.rs
stating outright that retaining the graph for a repeated backward pass
is not supported and that losses should be combined or the forward pass
repeated instead.
#5692,
squash commit 98e48ddbdd,
split out a separate operation for raising to an integer scalar power
and registers an explicit zero gradient for the zeroth power, adding
fast paths for powers one, two, minus one and minus two. The tests are
in crates/burn-backend-tests/tests/autodiff/pow.rs
and cover non-finite inputs among other cases.
#5837,
squash commit 28c871ce18,
takes one line but fixes a genuine bug. Insertion of a dimension and the
subsequent summation used an index one greater than required. Because
repetition along a dimension lays copies out in blocks, with a
non-uniform incoming gradient the result was not merely permuted but
outright incorrect. The test
should_diff_repeat_non_uniform_grad in crates/burn-backend-tests/tests/autodiff/repeat_dim.rs
supplies a non-uniform gradient and expects values the former code could
not produce.
#5547,
squash commit 708573a449,
fixed the backward pass of the product, where global reductions give
rank one while binary operations require matching ranks. There is no
test of its own in this PR. Coverage of empty axes came separately in #5598, squash
commit 0d26e3f83a,
and #5641,
squash commit ee94f264ca,
with tests in crates/burn-backend-tests/tests/autodiff/aggregation.rs.
#5624,
squash commit 7ecb912cab,
added transfer of a tensor between backends while preserving the graph.
Moving to a device preserves the graph, branching off a copy does not,
and moving to another device does not apply that device’s autodiff
settings.
#5889,
squash commit 697ae01af7,
deserves separate mention because it shows the state of the former
coverage. A test that had lost its expected-panic attribute, and
therefore checked nothing, was removed. Three tests sharing an expected
panic were tied to specific messages. The full-precision check
previously only made sure gradients were present; it now verifies the
type and the values. Autodiff tests of the attention mechanism, which
did not exist at all, were added, along with tests of index retrieval in
minimum, maximum and max pooling. This sharpens August’s caveat that
tests check specific scenarios: some of them did not check even
that.
#5775,
squash commit f9faec9a66,
added Adafactor in crates/burn-optim/src/optim/adafactor.rs.
Second moments are stored in factored form by rows and columns for
matrices and for tensors above rank two, so optimiser state does not
grow as the number of parameters. The learning rate is passed to the
step rather than to the configuration, by default the relative step is
bounded by the inverse square root of the step number, and scaling by
the root-mean-square value of the parameter is enabled. Tests in the
same directory check the vector and matrix paths against a scalar
reference implementation over two steps, check a round trip through
serialisation, check type preservation at reduced precision with state
in single precision, and check rejection of invalid epsilons and of the
clipping threshold.
#5741,
squash commit c7c72164e3,
added Lion, which stores one moment instead of two. It matters here not
to retell the tests as claiming more than they do. Both tests are marked
as smoke tests: in the first, which trains a tiny linear regression, the
comment says outright that this is not a reproduction of the paper’s
accuracy. The second compares not memory at run time but the length of
serialised state on a single linear layer, and the comment calls it a
stable, backend-independent indirect indicator.
#5843,
squash commit 81b5a3b315,
fixed gradient clipping by norm at half precision. The sum of squares
overflowed the half-precision range, the norm became infinite, the
clipping coefficient went to zero and the gradient was zeroed out
entirely. The norm, the coefficient and the scaling are now computed in
single precision with the result cast back. The tests
test_clip_by_norm_f16_does_not_overflow and a variant for a
norm above the half-precision limit are in crates/burn-optim/src/grad_clipping/base.rs.
The most coherent performance line of September is not about individual kernels but about the fact that Burn used to force data into contiguous form instead of working with its actual layout. A convolution may return a tensor physically laid out so that the channel varies fastest, even though the API presents it as a tensor with channels in the second dimension. Any following operation read such a tensor with a large stride and did not vectorise.
#5625,
squash commit 5533260720,
introduced launching in memory order for binary, scalar and unary
element-wise operations. Operands are fed in the order in which they are
laid out, the order is chosen by a byte-weighted vote, and the result is
permuted back. The helper functions that determine dimension order moved
from burn-cubecl-fusion into crates/burn-std/src/tensor/layout.rs.
The PR author’s measurement: CUDA, one L4 card, median of twenty runs.
The training step fell from 418.7 to 364.2 ms, of which the backward
pass went from 225.7 to 181.5 ms. The forward pass did not change, 166.7
against 166.0 ms.
#5620,
squash commit 265bb944d9,
lifted a restriction that could make this whole mechanism not work at
all. A tensor with alignment gaps in its allocation is now admitted to
the vote on a block’s layout, with a gap allowed but overlap not. Before
that, on a backend with an aligning allocator the layout propagation
never triggered, which from the outside looked like an unexplained
difference between backends. The author’s measurement: L4, median of ten
runs. On element-wise normalisation with an activation over a
channels-last representation, shapes 4 by 48 by 384 by 384, the gap
behind the contiguous layout narrowed from 6.28× to 1.10×. The
convolutional encoder sped up by 1.58×, the full forward pass by 1.6×,
the training step by 1.11×.
#5796,
squash commit 0e38b76bf9,
extended the same approach to type casting. The author’s measurement on
wgpu: casting from half to single precision for a channels-last tensor
sped up from 8.6 to 2.2 ms, against 2.1 ms on a contiguous layout.
#5760,
squash commit 2e626a85d7,
replaced the transposed two-dimensional convolution kernel with a
channels-last variant. The author’s measurement on an Apple M4 under
wgpu showed that the ratio of the data-gradient time to the forward pass
for one-dimensional convolution fell from 6.9 to 0.3. A substantive bug
was fixed along the way: the end of the window was derived from the
start plus a fixed width, so with a stride greater than the dilated
kernel size the output collapsed to a single offset.
An important caveat applies to all four numbers. These are PR authors’ measurements, taken on different configurations and against different baselines, so the percentages cannot be added up. On the forward pass, which fuses into a single operation from end to end, there is no measurable effect, and that is visible in #5625’s own data.
#5622,
squash commit 00215d4092,
changed the moment at which a block is committed: only the part of it
that cannot be fused is now committed. The description contains the
picture before the change, namely 3,588 single-element segments per step
and 3,238 unfused element-wise operations, of which 2,504 came in runs
of two or more. The author’s measurement on an L4 at batch size four,
medians of twenty runs: from 366.1 to 320.0 ms, and again from 370.8 to
324.8 ms.
#5850,
squash commit 219077efa1,
removed quadratic complexity in the router: the fusion estimate was
recomputed over the whole accumulated graph after every operation. The
count is now kept incrementally, and the full recomputation was kept as
a test reference in crates/burn-router/src/fusion.rs.
There are no numbers in the PR, so the size of the gain cannot be
stated.
#5901,
squash commit 7bc1dd6204,
eliminated the case where a tensor released on another thread before its
producer had run forced the fusion server to flush the entire operation
stream. The motivation is taken straight from the commit body and
relates to remote execution: the training event stream releases each
step’s outputs, and on top of a remote backend every step was cut in a
new place. A release now waits in a deferred queue.
#5621,
squash commit 1fd8d24bee,
sped up scatter-add by allocating one worker per value and using atomic
addition. The limits are stated explicitly: the addition operation only,
and only types with atomic addition, whereas assignment, multiplication
and logical OR remained sequential. There is a cost here that must be
named next to the gain: the order of additions is not fixed, so the
rounding of a floating-point sum may differ between runs. The gate is in
crates/burn-cubecl/src/kernel/index/scatter.rs.
The author’s measurement on an L4, a bank of one-dimensional
convolutions, 1,600,000 values into 8,192 slots: the training step from
95.9 to 81.2 ms.
Fusion fixes over the month produced two cases with visible
consequences. #5871, squash
commit d9b8377fae,
eliminated the choice of vector size along an unsuitable axis for a
transposed operand, which the trailing part of a fused kernel also
reads. The consequence was observable: an expression with a transposed
operand returned garbage, which made the Newton–Schulz iteration in the
Muon optimiser diverge into NaN. #5847, squash
commit 96425f5959,
fixed a mismatch in which the shape was taken from the output arguments
and the strides from the input ones, which gave either foreign strides
or an out-of-bounds index. #5725, squash
commit 18b3132c68,
removed autotuning for empty reductions, and #5902, squash
commit f328596411,
a double remapping of the input offsets of a fused matrix
multiplication.
#5822,
squash commit 80d3a024f3,
factored out a shared policy for bounding autotuning by a roofline model
and removed the duplicate table in burn-cubecl. The fused
matrix-multiplication and reduction tables got bounds for the first
time. There are no numbers in the PR.
#5673,
squash commit d40f99ea69,
added generation of fusion implementations for backend extensions in the
new crates/burn-backend-extension/src/fusion.rs.
This is a direct consequence of the extension mechanism from section 1: operations that have been moved out must
be able to fuse.
Tiled storage went through three steps in a row: #5631, squash
commit df92ec0750,
carried such a tensor through all the transformations, #5839, squash
commit 9ef294317c,
added the operation that converts to tiled form, and #5894, squash
commit a2edbd2663,
renamed the operations in the wake of the external dependency. That is,
the line moves at upstream’s pace and changed names once in a single
month.
#5545,
squash commit 7eaab04251,
is often described as speeding up convolution in general, and that is
wrong. A single file, crates/burn-cubecl/src/kernel/conv/direct.rs,
was changed, and the optimisation is enabled by a condition requiring
the hardware’s maximum plane size to equal one. That is, it works only
on the CubeCL CPU runtime, while on a GPU the former kernel is compiled.
The author’s measurements on a 5700X processor under the CPU runtime,
medians of four runs: ResNet50 at batch size one from 165.1 to 145.1 ms,
at batch size eight from 818.7 to 604.5 ms, MobileNet unchanged. The
reason for the gate is also stated in the commit body: on three NVIDIA
cards under CUDA the new shape was nine and twelve per cent slower in
single precision and up to twice as slow in half precision.
Eight days later this kernel left the repository. #5646, squash
commit da9e4224e4,
replaced it with a call to the direct convolution of an external
dependency, shrinking the same file from 422 lines to 87, and the commit
body says the kernel moved verbatim.
#5811,
squash commit d85aa0ab39,
moved the folding of four-dimensional blocks onto grouped convolution,
thanks to which the unrolling weight shrank along the input channel.
There are no measurements in the PR, so the conclusion about a gain
follows from the shape of the weight and not from a measurement.
Twenty PRs landed in Flex over the month, of which five concern
performance, fourteen fixes and one documentation. The performance list:
#5617, squash
commit 78159b741e,
on views, copies, in-place operations and parallelism; #5636, squash
commit 6e8932ca2a,
on parallelising attention over batch-and-head pairs in crates/burn-flex/src/ops/attention.rs;
#5603, squash
commit 89bb954882,
on in-place slice assignment in crates/burn-flex/src/ops/slice.rs;
#5761, squash
commit 1dafe9f4a2,
on the time and memory of transposed convolution in crates/burn-flex/src/ops/conv_transpose.rs;
#5865, squash
commit 943b90eb20,
on updating the vectorisation library and hoisting the SIMD path choice
out of hot loops, with changes in crates/burn-flex/src/simd/kernels.rs
among other places.
A direct caveat is needed here, otherwise the conclusion will outrun the data. This line contains no published measurements. #5617, #5636, #5761 and #5603 contain no numbers at all, only added benchmarks. The single number for Flex over the whole month appears in #5865, where a matrix multiplication of shape one million by eight sped up from roughly 650 to 290 microseconds, with the hardware not named. So it cannot be claimed on September’s data that Flex became faster by any particular fraction. This differs from August, where a measurement was given with a specific configuration.
It is worth noting separately that #5603 is less a speed-up than the removal of a hidden cost: conversion to contiguous form returned a clone, a second owning pointer kept the reference count at two, and an in-place change copied the whole destination on every slice assignment. A consuming variant of the operation was added.
Correctness fixes in Flex matter more than the above for a user who
was recommended this backend for CPU in August. #5634, squash
commit 3d72c99c84,
closed a summation fast path that accepted a zero-stride view, so that
an expanded tensor gave an inflated sum. #5650, squash
commit b7946a625c,
restored accounting for trailing padding in three-dimensional
convolution. #5663, squash
commit c0a8b0d13d,
replaced round-half-away-from-zero with round-half-to-even in grid
sampling, which changes the selected pixel relative to PyTorch. The
remaining Flex fixes are described in the section on
numeric semantics.
#5259,
squash commit 5f74b3938a,
closed August’s WIP on batched SVD, but not in August’s form. The power
method was replaced with the one-sided Jacobi method with a tournament
scheme for traversing pairs, which gives a linear rather than quadratic
number of kernel launches per pass. The right singular vectors are
obtained by an inverse transformation, and wide matrices are handled
through transposition. The public entry point is in crates/burn-linalg/src/functions/svd.rs,
that is, already in the new crate from section
1.
#5583,
squash commit 10b74fbb3d,
added einsum with two interfaces, a runtime call on an equation string
and a macro. #5582, squash
commit 766436c4ef,
deserves separate attention in connection with August’s deprecation
theme: before it every backend failed on the minimum and maximum
variants of scatter and select-assign, and the implementation was added
at once in NdArray, Flex, CubeCL and LibTorch together with the backward
pass. That is, new functionality is still being written for backends
declared as deprecating too. The matches in routing became exhaustive,
so a missed combination is now a compilation error.
#5372,
squash commit 4e8380f15e,
added adaptive three-dimensional average pooling in CubeCL, and #5616, squash
commit d38cb74d6b,
integer exponentiation.
#5688,
squash commit ac79057b48,
introduced device identity and enumeration of physical cards, under
which the same card reachable through CUDA and through Vulkan gives a
single entry. Grouping is by bus address, or by device identifier on
Windows. #5739, squash
commit bbea651c02,
arranged for a lazily initialised parameter to be created on the device
it moved to, rather than on the original one followed by a copy. #5746, squash
commit b9175d2225,
made it possible to count a model’s parameters from shapes without
allocating anything in memory. The purpose of the latter is stated
outright in the commit body: planning how to split a model across
devices. These are the only two primitives that reached
main from the theme discussed in the WIP
section.
Of the window’s 246 PRs, 36 change only the revision of the external CubeCL and Cubek dependencies in the root manifest together with the lock file. A further 18 PRs move the same pin but also carry their own code changes, so they cannot be called dependency maintenance. In all, 54 PRs touch the revision pin, that is 22 per cent of the month. Some fixes come precisely from there, for example the rejection of invalid kernel shapes in #5890 and the reading of an empty tensor in #5704. The limit of inference matters here: the public history shows the frequency of integration but not the reason for each revision, so it cannot be concluded from this data that Burn’s development is constrained by the external dependency. The only direct fact in that direction for September is the move of the convolution kernel from #5545 into Cubek eight days later.
Twelve PRs with scope remote landed over the month, and
eleven of them were merged between 24 and 30 September. The twelfth, #5724, squash
commit 07575391ff,
left installation of a process-wide log subscriber to the program rather
than the library. In effect all work on this theme is concentrated in
the last week of the period.
Before examining the content, two different notions that are easily
conflated need separating. The file crates/burn-tensor/src/tensor/distributed.rs
was not changed once in September. That is, distributed training with
data replication and collective operations did not advance in this
window. Everything discussed below concerns the transport and the
session lifecycle of burn-remote.
Three changes close one class of failure and work only together.
#5825,
squash commit e0319c6de7,
made a failure detectable at all. Before it neither side sent probe
messages and nothing timed out, so a client could wait on a read
indefinitely. A TCP-level liveness check on the connection was added:
the first probe after ten seconds, then every five, with a reset after
four unanswered. As a result a vanished peer shows up as a read error in
about thirty seconds, that is, the same way the Iroh transport does it
through its own idle timeout.
#5828,
squash commit 905c1954f6,
closed a leak on the server. The session service loop exited through the
error-propagation operator on any read error, bypassing session
teardown. A client that disconnected without an explicit close — and
that is every client, including an orderly endpoint close and an idle
timeout — left the session registered forever. The worker thread did not
exit and tensors stayed on the device. The consequence is named in the
description outright: on a card with six gigabytes, several sessions in
a row filled the memory, after which the next run failed for want of it.
The read loop was moved into a named function, the handshake response is
encoded before the session is bound, and the thread binding is performed
under a single lock. The changes run through crates/burn-remote/src/server/pump.rs.
#5878,
squash commit d2412e60df,
closed the same path for a panic. A panic on a session’s worker thread
unwound past teardown, objects exposed for transfer within the host
stayed in the server’s registry, and the memory was not returned to the
device. Unwind interception and response tasks with abort handles were
added in crates/burn-remote/src/server/worker.rs.
There is a citation trap here worth knowing about. The body of squash
commit d2412e60df
mentions #5903
and #5904, but
their changes are not in its own diff: both commits are already its
ancestors. The diff of #5878 is limited
to the worker, startup and local-exchange files. Citing this PR by the
text of its message credits it with someone else’s work.
#5829,
squash commit d373a5f90a,
eliminated a stall of blocking reads inside an asynchronous task after
roughly 128 operations with the processor fully loaded. The cause is
that every poll spent the task’s cooperative budget, which is only
replenished when control is yielded. The remedy is to lift the limit
explicitly for such a region.
#5903,
squash commit 5dac07a9ed,
raised the WebSocket frame size limit, which defaulted to sixteen
megabytes and therefore made any upload of a larger tensor drop the
session. The new limit is aligned with what the Iroh transport uses.
#5898,
squash commit 52bb097cbf,
eliminated an overlap of device addresses. Remote devices occupied type
identifier zero, so a remote device with a given number received the
same identifier as a local CUDA card with that number and shared its
runtime worker thread. The remote device type was moved to the end of
the range next to the graph-capture type in crates/burn-backend/src/backend/router_device_type.rs.
#5851,
squash commit 9b5402e9f1,
closed a leak during data loading: an initialisation operation captured
by the fusion scheduler in the router waited on the execution of a
stream that does not exist on a loader thread. The description quotes
consumption growing over an epoch from 74 to 1,865 megabytes, but that
figure was obtained on a remote-training example that is not in
main, so it cannot be reproduced from the repository.
#5678,
squash commit dec3046c39,
does not concern burn-remote at all, although it is
thematically close. A transfer between two devices of the same wgpu,
ROCm or Metal runtime panicked on the device threads, the panics were
swallowed, and the destination tensor was left unwritten. The path
through the collective-operations library for CUDA was kept.
#5910,
squash commit d97e804890,
is the theme’s largest PR of the month and the only one that adds a
capability rather than removing a failure. Channel and peer builders for
the Iroh transport appeared, along with token authorisation with
constant-time digest comparison, loading or creation of a secret in a
file readable only by its owner — written through a temporary file and a
rename that does not replace an existing one — a choice of relay mode
from three variants, a configurable port, and connection to a peer by
its identifier. Offloading of segmentation was moved into the
configuration file and is off by default after #5819, squash
commit 017464b63d.
Debug output of the peer and session-initialisation structs does not
print the secret. The update of the transport library itself to version
1.2.0 landed separately in #5830, squash
commit 757438aa98,
and consists of one line in the manifest.
What is described above is the release of resources on failure, not
resilience to it. Reconnection of an already established session is not
in main at the 30 September snapshot: the ensure-connection
function in crates/burn-remote/src/client/service.rs
returns immediately if the writing half is present, and it remains
present after the connection has died. The retries from #5824, squash
commit 740cc4f810,
with delays from a quarter of a second to eight seconds, apply only to
the first opening of the channels. A client’s tensors are not restored
after a disconnect.
The client’s behaviour on a disconnect is not uniform. Synchronisation and asynchronous reading of a tensor return an execution error, writes that do not wait for a response are logged and discarded, but a query of the data types in use and the handshake panic.
In that light #5904, squash
commit e501c844bf,
moves the panic rather than removing it. The server stopped crashing on
a request for a non-existent device index and now logs and closes the
stream instead, but the client in that same scenario crashes while
awaiting a successful handshake. There is no typed session failure in
main: the corresponding enumeration does not exist, and the
connection error type that does exist covers only a missing address, a
bind error and an abort.
Session teardown is also not bounded in time. An ordinary disconnect
does not abort the response tasks forcibly — that is done only after a
panic — the executor synchronisation runs without a timeout, and the
wait in the local exchange in crates/burn-remote/src/server/local_comm.rs
is performed in an unbounded loop. A hung read can delay the closing of
a session.
This section was reconstructed retrospectively from the Git history and current pull-request metadata. PR state was taken on 1 October 2026 rather than on 30 September, so no particular PR can be credited with draft status or with having an approval as of the period’s closing date: the available data contains no state-change history. A subsequent merge does not count towards September’s results. The size of a branch and the number of commits in it are not an estimate of readiness, and reasons for a delay cannot be inferred from them either. The dates of branches’ last commits are given in UTC−4, since in their original zone some of them look like 1 October although in the series’ zone they belong to 30 September.
Splitting a model by layers across devices. #5702 adds a
pipeline trait, a layout by stages and a matching example. As of 30
September the branch holds eleven commits outside main, the
last of them from 23 September. None of the types this PR introduces are
present in snapshot 91d892a4. Only two primitives from this
theme reached main, described in the section on memory layout: creation of a
lazily initialised parameter on the target device in #5739 and
counting the number of parameters from shapes without allocating memory
in #5746. The
purpose of the second is stated outright in its commit body, namely
planning how to split a model across devices. So the repository contains
groundwork for this capability, but the capability itself is not in
main.
Returning execution errors to the training loop. #5896 replaces
panics in the training loop with structured error types. It is the
period’s largest open initiative, twenty-four commits outside
main, the last from 30 September. The theme directly
continues August’s change in which a compute failure stopped breaking
independent work in the queue: there the error was tied to the tensors
affected, here it is meant to reach the calling training code. The
profiling that is sometimes associated with this branch is present in
main, but came from a different and merged #5726, squash
commit d0365962df.
A typed failure for connecting to a remote device.
#5922 makes
connecting return a result instead of panicking and introduces an
enumeration of session failure causes. Alongside it are #5723 on serving
streams a peer opens on a connection we dialed, #5920 on builds
with a partial feature set, #5919 with a
regression test for tensors dropped on another thread, and #5906 with an
example of training on another machine’s GPU. None of these changes are
in main, and the corresponding enumeration of failure
causes is not there either. This matters for reading section 8: typed handling of session failure remains
unfinished as of the end of September.
Retaining the graph for a repeated backward pass. #5043 moves the
step and backward-operation traits onto shared ownership so as to allow
several backward passes. Ten commits outside main, the last
from 22 September. The contrast with what was merged in the window
matters here: #5645 added
documentation to main stating outright that retaining the
graph is not supported and that losses should be combined or the forward
pass repeated instead. That is, the absence of the capability is pinned
down in the repository as of the end of the period, while the work to
add it proceeds in an open branch.
The learning rate as a device tensor for graph-captured
steps. #5872 moves the
learning rate onto an enumeration that admits a tensor on the device and
replaces a process-wide counter of active graph captures with a
per-device query to the backend. Ten commits outside main,
the last from 28 September. This is September’s only substantive contact
with the computation-graph capture theme, and it is not merged. The
graph-capture crate itself was touched by only two PRs tangentially in
the window.
The remaining substantive initiatives, grouped by
area. Reading a checkpoint from memory without a file continues
in #5755,
which belongs to section 2.
Three-dimensional pooling primitives in #5801 and
differentiable trilinear grid sampling in #5803 complement
the breaking change to the pooling interface from section 1. The data gradient for convolutions in #5813 and the
folding path with autotuning for three-dimensional transposed
convolution in #5793 form a
symmetric continuation of August’s weight-gradient work, and neither of
them landed in main. An arbitrary-size Fourier transform in
#5757 and a
complex inverse transform in #5840 extend the
new signal-processing crate, while a linear solve in #5856 and a
Cholesky decomposition in #5858 extend the
new linear-algebra crate. That is, both extracted crates immediately
acquired a work queue of their own. The Matthews correlation metric in
#5818 belongs
to section 5 and is absent from
main. Alongside are #5862 on an
unused gate in the coupled LSTM implementation, #5908 on
connected components in the CubeCL backend, #5911 on dropping
an incomplete batch in the data loader, and #5674 on exposing
tolerance fields.
Long-lived open PRs and a correction to the August material. #5504 with the GroupNorm primitive has been open since August, its branch’s last commit dating from 29 August. The internal August reference report described it, incorrectly, as merged. As of 30 September it is still open, as are six further PRs attributed to August’s work in that document: #3608, #5190, #5309, #5367, #5387 and #5541. Not one of the seven was integrated during the month.
| PR | Purpose | Branch’s last commit (UTC−4) | Commits outside main |
|---|---|---|---|
| #5896 | refactor(train): Re-surface execution errors in training loop | 2026-09-30 | 24 |
| #5702 | feat(core): pipeline trait to split a model by layers across devices | 2026-09-23 | 11 |
| #5906 | feat(examples): train and run MNIST on another machine’s GPU with remote-mnist | 2026-09-30 | 11 |
| #5043 | Retain graph feature for partial derivatives | 2026-09-22 | 10 |
| #5872 | feat(optim): accept a device learning rate for graph-captured steps | 2026-09-28 | 10 |
| #5504 | feat: add GroupNorm backend primitive and CubeCL kernel | 2026-08-29 | 8 |
| #5920 | fix(remote): gate what partial feature builds leave unused | 2026-09-30 | 8 |
| #5723 | fix(remote): serve the streams a peer opens on a connection we dialed | 2026-09-28 | 7 |
| #5801 | feat: add 3D spatial pooling primitives (avg_pool3d,
max_pool3d) (#5785) |
2026-09-30 | 5 |
| #5755 | add from bytes | 2026-09-26 | 3 |
| #5757 | feat(signal): support arbitrary-size rfft/irfft via Bluestein’s algorithm | 2026-09-30 | 3 |
| #5922 | feat(remote)!: connecting to a remote device returns a Result | 2026-09-30 | 3 |
| #5674 | Expose Tolerance fields | 2026-09-14 | 2 |
| #5803 | feat: add differentiable trilinear grid_sample_3d |
2026-09-23 | 2 |
| #5919 | test(remote): feed a queued reader from tensors dropped on another thread | 2026-09-30 | 2 |
| #5921 | fix(metal): require native MSL for explicit Metal devices | 2026-09-30 | 2 |
| #5541 | Validate CubeCL complex cast lowering | 2026-08-31 | 1 |
| #5793 | feat(cubecl): add a col2im path and autotune for conv_transpose3d | 2026-09-23 | 1 |
| #5813 | perf(conv): add im2col data-gradient path for dense convolutions | 2026-09-24 | 1 |
| #5818 | feat(train): add Matthews correlation coefficient metric | 2026-09-24 | 1 |
| #5840 | feat(signal): add complex inverse FFT (ifft) | 2026-09-30 | 1 |
| #5856 | feat(linalg): add batched linear solve | 2026-09-26 | 1 |
| #5858 | Cholesky implementation | 2026-09-26 | 1 |
| #5862 | fix(nn): omit unused forget gate in coupled LSTM | 2026-09-26 | 1 |
| #5908 | fix(vision): call the CPU connected components directly in the cube backend | 2026-09-29 | 1 |
| #5911 | Add drop_last to the data loader |
2026-09-29 | 1 |
Below are all 246 PRs merged into main in the window,
each exactly once. The grouping is by area of change, not by importance:
the themes examined above appear here on equal terms with the rest, so
that the list stays complete and checkable. For each PR the merge date
in UTC−4 and the squash commit in main are given, by which
the statement can be verified directly.
There is not a single direct commit in the window, so no separate section for them is needed. This differs from August, where the window contained one direct commit.
| PR | Merged | Squash commit | Content |
|---|---|---|---|
| #5548 | 2026-09-02 | 435b354974 |
update cube |
| #5550 | 2026-09-02 | ab0439eef6 |
Update cubek |
| #5560 | 2026-09-03 | 7f5736d248 |
update cube |
| #5561 | 2026-09-03 | 6fe858268a |
chore(deps): update cubecl and cubek for fallible throughput probes |
| #5556 | 2026-09-04 | 6e3ada50d0 |
Chore/update cubecl runtime erasure |
| #5565 | 2026-09-04 | 530f68153e |
update cube |
| #5566 | 2026-09-04 | 1f96ddaacc |
Update/cube3 |
| #5586 | 2026-09-06 | 05ac3f79c1 |
chore: update cubecl and cubek revs |
| #5587 | 2026-09-06 | 3073dba638 |
Update cubek |
| #5589 | 2026-09-07 | abb50388bc |
update cubek |
| #5591 | 2026-09-07 | b9a1563d9f |
Update CubeCL & CubeK |
| #5611 | 2026-09-08 | 5ecfb38965 |
update cubek |
| #5627 | 2026-09-08 | 8b35c753cc |
update cubek |
| #5632 | 2026-09-09 | d4cd33666b |
update cube |
| #5644 | 2026-09-10 | eb03c79803 |
update cube |
| #5649 | 2026-09-10 | 09f9b3f78b |
update cube |
| #5668 | 2026-09-14 | e523ad3c3f |
update cube |
| #5670 | 2026-09-14 | cfe65e1cfc |
Update/cubecl crate layout |
| #5691 | 2026-09-15 | c5aa775981 |
update cube |
| #5697 | 2026-09-16 | f621d6edb6 |
update cube |
| #5701 | 2026-09-16 | d01bb5499a |
chore: update cubek |
| #5719 | 2026-09-17 | 3b8fb7387d |
update cube |
| #5703 | 2026-09-17 | 577fece293 |
chore: update cubek |
| #5732 | 2026-09-18 | 1d6f50330d |
update cube |
| #5763 | 2026-09-21 | e13601d3c7 |
Chore/cubecl device capacity |
| #5770 | 2026-09-21 | 28550733d5 |
update cube |
| #5771 | 2026-09-22 | c6f9fd4421 |
Chore/cubecl environment records |
| #5797 | 2026-09-23 | 0fa63b3e45 |
Chore/cubecl adaptive pool |
| #5810 | 2026-09-23 | fafc9d7a92 |
Update cubecl and cubek to the tune plan record |
| #5841 | 2026-09-25 | 86781f0887 |
chore(deps): bump cubek to a066bdd |
| #5848 | 2026-09-25 | c240f0011c |
chore(deps): bump cubek to b40b522 |
| #5854 | 2026-09-25 | 474f46b73a |
chore(deps): bump cubecl to a1bb768 and cubek to 9db95ba |
| #5870 | 2026-09-28 | 49dcd5ba9e |
chore(deps): bump cubek to 7d60e30 |
| #5887 | 2026-09-28 | 3d605097e2 |
chore(deps): bump cubek to 41a4ab0 and cubecl to 33b6dfb |
| #5895 | 2026-09-29 | 53ebdce405 |
chore: bump cubecl to f0cf8383 and cubek to 9f35e2ab |
| #5905 | 2026-09-29 | 5f3cc76826 |
chore: bump cubek to 6927503 |
| #5907 | 2026-09-29 | 98ca0bdac8 |
update cube |
| #5913 | 2026-09-30 | 9fa49d9cdd |
chore: bump cubek to bb461fb |
| #5915 | 2026-09-30 | f5d57e66b9 |
chore: bump cubek to 5753e85 |
| #5917 | 2026-09-30 | a6de085337 |
chore: bump cubecl to 48301b5 and cubek to 2e07d52 |
| #5918 | 2026-09-30 | 903f6897e5 |
chore: bump cubek to e84f83d and cubecl to 48301b5 |
| PR | Merged | Squash commit | Content |
|---|---|---|---|
| #5593 | 2026-09-10 | 33df298f84 |
fix(store): harden and restructure the PyTorch reader |
| #5664 | 2026-09-15 | 317fb474fc |
fix(store): collect PyTorch tensors nested in lists/tuples with indexed names |
| #5656 | 2026-09-16 | 0fdac31a3d |
refactor(store): extract the PyTorch reader into a burn-free pytorch-reader crate |
| #5728 | 2026-09-18 | e68583e42f |
fix(pytorch-reader): report values load_config cannot represent instead of defaulting |
| #5714 | 2026-09-18 | b2faf4f336 |
perf(pytorch-reader): read stored ZIP entries outside the archive lock |
| #5700 | 2026-09-21 | 96e2ec1e12 |
fix(pytorch-reader): remove unsafe visitor cloning and deserialization panics |
| #5737 | 2026-09-21 | d79dab2d12 |
fix(pytorch-reader): open a checkpoint once and never reopen it by path |
| #5747 | 2026-09-21 | b3984918b5 |
docs(pytorch-reader): sync README with lib.rs |
| #5764 | 2026-09-21 | 03e2da00c5 |
fix(pytorch-reader): handle unsupported values in deserialize_any |
| #5749 | 2026-09-22 | 4d26deac29 |
fix(pytorch-reader): thiserror source chains + reject non-checkpoint files |
| #5756 | 2026-09-22 | a36db79bc0 |
fix(pytorch-reader): accept checkpoints saved with compute_crc32=False |
| #5766 | 2026-09-23 | 855fe4e812 |
feat(pytorch-reader): load torch.save(model) full-model pickles |
| PR | Merged | Squash commit | Content |
|---|---|---|---|
| #5489 | 2026-09-01 | 137b73a5ae |
fix(store): write safetensors files atomically |
| #5490 | 2026-09-02 | c1d6874cb9 |
fix(store): reject a loaded tensor whose dtype is the wrong kind |
| #5558 | 2026-09-04 | ef521c0896 |
fix(store): reject invalid path filter regex |
| #5576 | 2026-09-08 | a7be13739c |
fix(pack): validate tensor byte lengths when reading |
| #5666 | 2026-09-14 | 52c0535842 |
test(pack): check streaming memory with a drop hook, not a global allocator |
| #5750 | 2026-09-21 | bd9a529bbc |
fix(burn-store): scope contiguous index mapping per prefix |
| #5781 | 2026-09-23 | d17a1ff8b8 |
fix(store): enforce overwrite(false) at publish time, not just before the save |
| #5832 | 2026-09-28 | 18ac73b70b |
feat(pack): make atomic writes the default |
| PR | Merged | Squash commit | Content |
|---|---|---|---|
| #5724 | 2026-09-18 | 07575391ff |
fix(remote): leave the process-wide log subscriber to the program |
| #5819 | 2026-09-24 | 017464b63d |
fix(remote): turn off iroh GSO until iroh#4555 is fixed |
| #5830 | 2026-09-25 | 757438aa98 |
chore(deps): bump iroh to 1.2.0 |
| #5824 | 2026-09-25 | 740cc4f810 |
fix(remote): try again while a server is not reachable yet |
| #5825 | 2026-09-25 | e0319c6de7 |
fix(remote): notice a WebSocket peer that vanished without closing |
| #5829 | 2026-09-28 | d373a5f90a |
fix(remote): blocking reads inside a tokio task stall after 128 |
| #5828 | 2026-09-28 | 905c1954f6 |
fix(remote): end a server session when its client disconnects |
| #5873 | 2026-09-28 | ccdfede767 |
test(remote): serve the tests on ports the OS picks, with burn-remote linked once |
| #5903 | 2026-09-29 | 5dac07a9ed |
fix(remote): let a WebSocket server read what its clients send |
| #5904 | 2026-09-29 | e501c844bf |
fix(remote): refuse a session for a device the server does not host |
| #5898 | 2026-09-29 | 52bb097cbf |
fix(remote): keep remote devices off local GPUs’ runner threads |
| #5878 | 2026-09-30 | d2412e60df |
fix(remote): clean up a session whose task panicked |
| #5910 | 2026-09-30 | d97e804890 |
feat(remote): configure an Iroh server’s relays, port and authorizer, and dial one by its id |
| PR | Merged | Squash commit | Content |
|---|---|---|---|
| #5547 | 2026-09-02 | 708573a449 |
fix(autodiff): prod backward broadcasting |
| #5623 | 2026-09-09 | 04e0f2ded6 |
perf(autodiff): reduce a broadcast gradient over all dims at once |
| #5624 | 2026-09-09 | 7ecb912cab |
feat(autodiff): support graph-preserving cross-backend transfers |
| #5647 | 2026-09-11 | 6212b88fe0 |
fix(autodiff)!: retain input nodes until child registration |
| #5645 | 2026-09-11 | 68a92fdfc9 |
fix(autodiff): explicitly reject consumed graph reuse and preserve reusable leaves |
| #5679 | 2026-09-15 | 7d9e2be68c |
perf(autodiff): avoid redundant traversal in checkpoint topological sort |
| #5692 | 2026-09-17 | 98e48ddbdd |
fix(autodiff): preserve gradients for zero scalar exponents |
| #5837 | 2026-09-25 | 28c871ce18 |
fix(autodiff): correct repeat_dim gradient ordering |
| #5889 | 2026-09-29 | 697ae01af7 |
test(autodiff): tighten weak tests and cover untested backward paths |
| PR | Merged | Squash commit | Content |
|---|---|---|---|
| #5618 | 2026-09-08 | 2440750e19 |
fix(optim): keep a parameter’s checkpointing strategy across an update |
| #5574 | 2026-09-08 | b6a4f30fc0 |
fix(rl): preserve deterministic mode in async batches |
| #5578 | 2026-09-08 | 2646751f3b |
fix(train): exclude padded samples from accuracy aggregation |
| #5684 | 2026-09-15 | fb90882e3e |
fix(train): handle tied scores in average precision |
| #5729 | 2026-09-21 | 288b697e89 |
fix(train): flush partial gradient accumulation and step LR per optimizer update |
| #5805 | 2026-09-24 | d26efc54a4 |
feat(train): add labels to training progress loggers |
| #5866 | 2026-09-28 | c357872392 |
fix(tui): handle already-joined thread in manual close |
| #5900 | 2026-09-29 | 58cc57d767 |
fix(train): compute AUROC with sorted score groups |
| PR | Merged | Squash commit | Content |
|---|---|---|---|
| #5775 | 2026-09-22 | f9faec9a66 |
feat(optim): add Adafactor optimizer |
| #5741 | 2026-09-22 | c7c72164e3 |
feat(optim): add Lion optimizer |
| #5799 | 2026-09-23 | 2e5d2152be |
test(optim): strengthen optimizer save-load round-trip coverage |
| #5843 | 2026-09-25 | 81b5a3b315 |
fix(optim): compute gradient norm clipping in F32 for half precision |
| PR | Merged | Squash commit | Content |
|---|---|---|---|
| #5569 | 2026-09-04 | 1bceb880ec |
fix(nn): exclude pad tokens from cross-entropy normalization |
| #5592 | 2026-09-09 | 986f0c1da1 |
fix(nn): honor count_include_pad for asymmetric average pooling |
| #5676 | 2026-09-16 | 55ee6c82f8 |
perf(nn): use a dedicated BatchNorm training op with closed-form backward |
| #5860 | 2026-09-28 | c0878a0ab0 |
fix(nn): validate scalar module configurations |
| #5892 | 2026-09-29 | 9e2eefff24 |
fix(nn): clamp probabilities in smoothed cross-entropy |
| #5875 | 2026-09-29 | 440a9e2299 |
fix(nn): apply attention dropout after softmax |
| PR | Merged | Squash commit | Content |
|---|---|---|---|
| #5619 | 2026-09-09 | dd37d19950 |
docs(flex): fix stale statements in burn-flex docs and comments |
| #5603 | 2026-09-09 | 89bb954882 |
perf(flex): write slice_assign in place when the destination is uniquely owned |
| #5617 | 2026-09-09 | 78159b741e |
perf(flex): optimize views, copies, in-place ops, and Rayon parallelism (#5613) |
| #5629 | 2026-09-09 | 4a0d726530 |
fix(flex): avoid overflow in boolean select updates |
| #5630 | 2026-09-10 | 9cc7f63ac7 |
fix(flex): support unsigned dtypes in int_argmax/int_argmin |
| #5634 | 2026-09-10 | 3d72c99c84 |
fix(flex): guard sum fast path with layout_covers_storage_once |
| #5640 | 2026-09-14 | 18af6a8c8f |
fix(flex): validate gather_nd and scatter_nd coordinates |
| #5663 | 2026-09-14 | c0a8b0d13d |
fix(flex): round nearest grid_sample ties to even |
| #5661 | 2026-09-14 | 7ef8198a48 |
fix(flex): make u64 right shift logical |
| #5650 | 2026-09-15 | b7946a625c |
fix(flex): honor asymmetric padding in conv3d |
| #5636 | 2026-09-15 | 6e8932ca2a |
perf(flex): parallelize attention across (batch, head) pairs (#5612) |
| #5652 | 2026-09-15 | e205024136 |
fix(flex): handle remainder edge cases |
| #5662 | 2026-09-15 | 11f8c072e1 |
fix(flex): propagate NaN through max pooling |
| #5658 | 2026-09-16 | 52c101f455 |
fix(flex): propagate NaN through relu and clamp_min/clamp_max |
| #5717 | 2026-09-18 | dae47a34b8 |
fix(flex): handle NaN clamp bounds |
| #5780 | 2026-09-23 | d918eff531 |
fix(flex): compute f32 layer_norm variance in two passes |
| #5761 | 2026-09-23 | 1dafe9f4a2 |
perf(flex): reduce ConvTranspose time and memory use |
| #5808 | 2026-09-24 | ebad4f9d6d |
fix(flex): mask shift amounts to each int dtype’s own width |
| #5865 | 2026-09-28 | 943b90eb20 |
perf(flex): upgrade macerator to 0.5.0 and hoist SIMD dispatch out of hot loops |
| #5888 | 2026-09-28 | 3bf6e03948 |
fix(flex): enforce quantization block alignment |
| PR | Merged | Squash commit | Content |
|---|---|---|---|
| #5543 | 2026-09-01 | 13637e9558 |
fix(no-std): use shared sync primitives across crates |
| #5553 | 2026-09-03 | 37f87ab37f |
fix(ndarray): preserve SIMD unary layout and recip precision |
| #5639 | 2026-09-10 | d26b1efbb7 |
test(burn-std): import vec! macro in layout.rs tests for no-std |
| #5665 | 2026-09-15 | 66a8a5ff8a |
fix(burn-ndarray): sign(NaN) must be 0, not its hidden sign bit |
| #5733 | 2026-09-21 | 6ade1047f8 |
fix(ndarray): preserve NaN through float clamps |
| #5738 | 2026-09-23 | 9b08e4aba5 |
fix(std): ignore unit axes when checking contiguity |
| PR | Merged | Squash commit | Content |
|---|---|---|---|
| #5620 | 2026-09-08 | 265bb944d9 |
perf(fusion): let a padded tensor vote for the block’s layout |
| #5622 | 2026-09-08 | 00215d4092 |
perf(fusion): settle only the unfusable head of a block |
| #5633 | 2026-09-09 | ba1886d2cf |
test(fusion): change f32 expected value to oracle |
| #5725 | 2026-09-21 | 18b3132c68 |
fix(fusion): avoid autotuning empty reductions |
| #5850 | 2026-09-28 | 219077efa1 |
perf(router): avoid rescanning the graph when scoring fusion |
| #5871 | 2026-09-28 | d9b8377fae |
fix(fusion): keep the default vectorization axis for matmul operands the epilogue reads |
| #5847 | 2026-09-28 | 96425f5959 |
fix(fusion): resolve reduce reference strides against the reference’s argument list |
| #5851 | 2026-09-28 | 9b5402e9f1 |
fix(router): close the fuser on an upload so the upload can be freed |
| #5877 | 2026-09-28 | 64e563fd9e |
fix(router): serve unsigned int tensors in the interpreter |
| #5893 | 2026-09-29 | 11a0757036 |
test(fusion): relax half-precision absolute tolerance for matmul epilogue view |
| #5890 | 2026-09-29 | bd0f1dbc40 |
fix(fusion): update cubek to reject invalid VecMat kernel shapes |
| #5902 | 2026-09-29 | f328596411 |
fix(fusion): avoid remapping fused matmul input offsets twice |
| #5901 | 2026-09-30 | 7bc1dd6204 |
perf(fusion): defer a cross-thread drop until its producer runs instead of draining the stream |
| PR | Merged | Squash commit | Content |
|---|---|---|---|
| #5372 | 2026-09-08 | 4e8380f15e |
feat(cubecl): support AdaptiveAvgPool3d |
| #5616 | 2026-09-08 | d38cb74d6b |
feat(cubecl): implement integer powi operations |
| #5621 | 2026-09-09 | 1fd8d24bee |
perf(cubecl): scatter-add runs one unit per value with atomic adds |
| #5635 | 2026-09-09 | 20d53f5cfc |
fix(cubecl): gate storage_tiled tests behind runtime features |
| #5625 | 2026-09-09 | 5533260720 |
perf(cubecl): walk elementwise kernels in operands’ memory order |
| #5646 | 2026-09-11 | da9e4224e4 |
refactor(cubecl): call cubek’s direct convolution routine |
| #5654 | 2026-09-11 | 1414c8a14e |
fix(cubecl): update cubek to fix sums over overlapping views |
| #5678 | 2026-09-16 | dec3046c39 |
fix(cubecl): same-runtime device moves without a peer transport, and quantized moves |
| #5704 | 2026-09-16 | 7bc513b711 |
fix(cubecl): update dependencies to fix empty tensor readback |
| #5718 | 2026-09-18 | 274cb46cb9 |
fix(cubecl): propagate NaN through two-sided clamp |
| #5767 | 2026-09-21 | 18fa3bccae |
fix(cubecl): handle broadcasting in mask_where and mask_fill |
| #5760 | 2026-09-23 | 2e626a85d7 |
perf(cubecl): NHWC direct kernel for conv_transpose2d |
| #5796 | 2026-09-23 | 0e38b76bf9 |
perf(cubecl): walk cast in its input’s memory order |
| #5815 | 2026-09-24 | caa39cf874 |
Metal/wgpu msl tests |
| #5867 | 2026-09-28 | af3100e5af |
fix(deps): restore gpu-allocator’s windows version to match wgpu-hal |
| #5844 | 2026-09-28 | bea4cdef84 |
fix(cubecl): reject reshape that splits a quantization block across rows |
| #5881 | 2026-09-29 | baf6943408 |
fix(cubecl): skip the reduce launch when another axis leaves the output empty |
| #5909 | 2026-09-30 | 56584ee48b |
fix(cube): propagate NaNs in clamp and stabilize GPU tests |
| PR | Merged | Squash commit | Content |
|---|---|---|---|
| #5520 | 2026-09-01 | 0e2dc6b631 |
refactor(dispatch): unify routing and strengthen autodiff contract |
| #5551 | 2026-09-02 | cf1b6d2a60 |
fix(dispatch): restore autodiff context promotion |
| #5537 | 2026-09-03 | af55cf4432 |
feat(module): separate gradient control from module freezing |
| #5521 | 2026-09-03 | 7b815b01bd |
feat(backend)!: support asymmetric padding in conv1d and conv2d |
| #5570 | 2026-09-04 | 54d36fa936 |
fix(dispatch): decouple CubeCL runtime facade crates |
| #5598 | 2026-09-08 | 0d26e3f83a |
test(backend): cover empty-axis autodiff reductions |
| #5641 | 2026-09-10 | ee94f264ca |
test(backend): cover empty-axis autodiff product reductions |
| #5673 | 2026-09-16 | d40f99ea69 |
feat(extension): generate Fusion implementations for backend extensions |
| #5705 | 2026-09-17 | b9088436b6 |
fix(core): let an init_mapper parameter train on an autodiff device |
| #5721 | 2026-09-17 | 8f08fdad05 |
refactor(module)!: merge AutodiffModule into
Module |
| #5722 | 2026-09-18 | 596d9acfbd |
fix(dispatch): require explicit backend selection and support backend-free builds |
| #5746 | 2026-09-21 | b9175d2225 |
fix(core): count a module’s parameters without initializing them |
| #5739 | 2026-09-21 | bbea651c02 |
feat(core): a lazy parameter initializes on the device it moves to |
| #5653 | 2026-09-22 | d580fbca80 |
fix(backends): correct average-pooling gradients with ceil mode |
| #5778 | 2026-09-22 | 1726b8d47e |
fix(module): preserve lazy initialization when collecting devices |
| #5794 | 2026-09-23 | ed0be6f22d |
test(backend): skip ue4m3 scale accuracy cases where no 8-bit type exists |
| #5798 | 2026-09-24 | 9eadada8ef |
fix(backend): match PyTorch pooling output size in ceil_mode |
| #5821 | 2026-09-25 | b9dec8235c |
fix(module): preserve reparameterizations during validation |
| #5869 | 2026-09-28 | 1a5a4a35ee |
fix(backend)!: separate device settings queries from initialization |
| #5916 | 2026-09-30 | 8453f26d8e |
test(backend): run the empty select and roll tests |
| PR | Merged | Squash commit | Content |
|---|---|---|---|
| #5259 | 2026-09-02 | 5f74b3938a |
feat(tensor): add batched SVD decomposition to linalg |
| #5555 | 2026-09-03 | 38f3d8a7d1 |
fix(tensor): validate matmul batch broadcast in TensorCheck |
| #5557 | 2026-09-03 | f7f08fd57c |
feat(tensor): replace no_grad with explicit autodiff
conversions |
| #5559 | 2026-09-03 | a30a9214a7 |
docs(tensor): fix unfold window formula |
| #5539 | 2026-09-04 | 163b7f442d |
feat(tensor)!: add multi-axis vector norm variants and update empty
max_abs_dims semantics |
| #5554 | 2026-09-04 | 5bd0170384 |
feat(tensor): add assert_shape! and debug_assert_shape! macros |
| #5564 | 2026-09-04 | a181e614aa |
feat(tensor): apply TensorCheck to remainder, powi, powf, hypot, atan2 |
| #5571 | 2026-09-04 | d16f7ba2ed |
feat(tensor): add is_autodiff and
is_tracked state inspection |
| #5580 | 2026-09-08 | 64a69aca89 |
feat(tensor): validate matmul rank in TensorCheck |
| #5572 | 2026-09-08 | 47ceefde21 |
feat(linalg)!: extract tensor linalg into burn-linalg
extension crate |
| #5585 | 2026-09-09 | 5871c72ef6 |
fix(tensor): avoid cosine similarity denominator underflow |
| #5648 | 2026-09-11 | 545682e718 |
fix(linalg): add autotune feature propagation |
| #5582 | 2026-09-14 | 766436c4ef |
feat(tensor): implement Min and Max scatter/select_assign across backends |
| #5693 | 2026-09-16 | d6883eb4fa |
fix(linalg): support negative even lp norm orders |
| #5583 | 2026-09-16 | 10b74fbb3d |
feat(tensor): add einsum with runtime and macro APIs |
| #5720 | 2026-09-17 | 1c337e0d80 |
feat(signal)!: extract tensor signal into burn-signal
extension crate |
| #5688 | 2026-09-18 | ac79057b48 |
feat(tensor): device identity and one entry per physical GPU |
| #5743 | 2026-09-21 | 752e82b252 |
fix(tensor): fix quiet softmax for negative infinity slices |
| #5745 | 2026-09-21 | b683dc45cc |
fix(tensor): preserve f64 precision in degree/radian conversions |
| #5754 | 2026-09-21 | ab84b8e69f |
fix(tensor): return NaN for infinite dividends in fmod_scalar |
| #5790 | 2026-09-23 | 1471422107 |
fix(tensor): keep input dtype on empty slice/repeat/cat results and one_hot_fill |
| #5804 | 2026-09-24 | b930e39f47 |
fix(linalg): import Vec in fusion svd for no_std builds |
| #5827 | 2026-09-25 | e814443c76 |
fix(tensor): reject quantized inputs in *_like ops and one_hot |
| #5846 | 2026-09-25 | af8b3306db |
refactor(tensor): remove into_tiled from the public API |
| #5838 | 2026-09-25 | e201487165 |
fix(tensor): make TensorData fields private and validate construction |
| #5849 | 2026-09-28 | 84b9669a59 |
feat(tensor)!: accept output_size or scale_factor in InterpolateOptions |
| #5868 | 2026-09-29 | cdeccbcf6c |
fix(tensor): display quantized tensors as metadata without panicking |
| #5853 | 2026-09-30 | 9adff02644 |
refactor(tensor)!: move pooling args into options structs with asymmetric padding |
| #5912 | 2026-09-30 | 91d892a4ad |
fix(tensor): roll shift direction to match torch.roll |
| PR | Merged | Squash commit | Content |
|---|---|---|---|
| #5544 | 2026-09-01 | ef7c981590 |
chore(deps): reduce unnecessary dependencies |
| #5669 | 2026-09-14 | 8f7a1c435c |
ci: reduce redundant compilation in macOS tests |
| #5672 | 2026-09-14 | 7832f67a00 |
fix(deps): update rustls for cargo audit |
| #5671 | 2026-09-14 | 45734c0170 |
ci: reduce redundant compilation across test suites |
| #5675 | 2026-09-14 | faa20a1f4a |
chore(deps)!: make backend tracing opt-in |
| #5680 | 2026-09-15 | 11f9ff07d0 |
chore(deps): reduce dataframe dataset dependencies with polars-core |
| #5681 | 2026-09-15 | f6e8251512 |
ci: reduce examples overhead by skipping coverage setup and enabling caching |
| #5694 | 2026-09-16 | 8116dff453 |
fix(deps): avoid enabling CubeCL through linalg defaults and respect vision defaults |
| #5715 | 2026-09-17 | f8534af9e8 |
chore(deps): trim zip default features and fix burn-dataset nlp feature |
| #5727 | 2026-09-18 | da2209a9de |
fix(ci): publish burn-einsum and include it in no-std checks |
| #5735 | 2026-09-18 | 3bd1a6e91f |
fix(deps): restore the windows crate versions the cube update changed |
| #5736 | 2026-09-18 | 3a93fbfc0f |
chore: remove recursion limit workarounds and update CubeCL configs |
| #5776 | 2026-09-22 | d3c8d7c615 |
chore: bump version to 0.22.0-pre.4 |
| #5777 | 2026-09-22 | 9147c11a19 |
fix(ci): use cargo info to check published crate versions |
| #5786 | 2026-09-23 | ee211424ac |
chore(deps): bump actions/checkout from 6 to 7 |
| #5831 | 2026-09-25 | 28235bc19e |
chore(deps): drop unused bincode workspace dependency |
| PR | Merged | Squash commit | Content |
|---|---|---|---|
| #5731 | 2026-09-18 | 28bfe74d5f |
docs(book): update backend extension guides and include maintained example sources |
| #5748 | 2026-09-21 | 70326cfc4b |
docs: fix typos and grammar in books and API comments |
| #5765 | 2026-09-21 | f726002c08 |
docs: clarify guidelines for minor documentation fixes |
| #5762 | 2026-09-21 | 9f1bb5112f |
docs: update guides and examples for Burn 0.22 |
| #5768 | 2026-09-21 | faec398324 |
docs: cover remaining 0.22 migration points and refresh stale pages |
| #5773 | 2026-09-22 | bc2b9832f8 |
docs: select the package when running examples from the repo root |
| #5800 | 2026-09-23 | f56efabd89 |
clarify 0.22 migration guidance and update API examples |
| #5845 | 2026-09-25 | f5bd4c075c |
docs(book): document burn.toml runtime configuration |
| #5863 | 2026-09-28 | 6352b35ef4 |
docs: fix dead links in README and the Burn Book |
| PR | Merged | Squash commit | Content |
|---|---|---|---|
| #5540 | 2026-09-02 | 046a21c436 |
Refactor/fusion write scope |
| #5545 | 2026-09-03 | 7eaab04251 |
perf(conv): accumulate direct convolution channels in vector registers on CPU |
| #5631 | 2026-09-09 | df92ec0750 |
Feat/storage tiled carrier |
| #5546 | 2026-09-09 | e80e7981b9 |
refactor(dataset): back SqliteDataset with Turso instead of rusqlite |
| #5667 | 2026-09-14 | 5685f782e9 |
fix(burn-dataset): cap MNIST item counts at the split size |
| #5726 | 2026-09-18 | d0365962df |
Feat/profiling |
| #5814 | 2026-09-24 | a9de0f0487 |
fix(test): align dtype-support expectations with the runtimes |
| #5816 | 2026-09-24 | ffc4fe0223 |
fix(examples): correctly format web inference probability labels |
| #5806 | 2026-09-24 | 0b6f7cbf04 |
fix: remove rl from default features |
| #5823 | 2026-09-24 | b90efe979d |
fix(examples): correct WebGPU inference and keep live predictions responsive |
| #5811 | 2026-09-24 | d85aa0ab39 |
perf: use grouped convolution for fold4d |
| #5839 | 2026-09-25 | 9ef294317c |
Qa/tiled storage |
| #5822 | 2026-09-28 | 80d3a024f3 |
Share the autotune roofline policy and bound fused matmul and reduce tuning |
| #5894 | 2026-09-29 | a2edbd2663 |
Refactor/storage tile naming |
Burn and Tracel AI are trademarks of their respective owners. PRA is not affiliated with Tracel AI.