← PRA / report example

Burn — technical evidence review, August 2026

Companion to the management memo. Prepared 29 September 2026 from the local repository history and saved PR metadata.

Scope and evidence

Repository: tracel-ai/burn, branch main. Window: [2026-08-01T00:00:00-04:00, 2026-09-01T00:00:00-04:00). The main line of history (first-parent) contains 137 PR-associated commits and one direct commit in this window. Changes are selected by their committer timestamps on main, not their author timestamps on feature branches. The period ends at this main revision.

This review covers all changes merged in the window. Each of the 137 pull requests and the direct commit is assigned to one topic section. Work still in progress at the cutoff and subsequent September merges are described separately and excluded from the August totals.

All changes are covered, at two levels of detail:

The analysis is based on code and tests at the linked revisions. Tests were read, not run for this review: the text describes what they check, not the results of an independent test run. Production workloads, adoption and financial outcomes were not evaluated. Where a PR description differs from the code, the merged implementation takes precedence. Benchmarks from PRs are attributed to their authors, with measurement conditions where the source provides them.

Period summary

Measure August 2026
Pull requests merged into main 137
Direct commits to main 1
Pull requests marked breaking 6
Days with at least one merge 24
Pre-releases 2 (0.22.0-pre.2, 0.22.0-pre.3)

The six changes marked as breaking are #5343 (Candle removal), #5262 (quantization scale model), #5290 (NaN propagation in extrema), #5316 (TensorData access APIs), #5400 (fusion and autodiff edge cases) and #5411 (store transport type). Each is described and linked in its own section below. The count comes from the breaking marker in the squash commit’s subject on main.

Contributor rankings, diff sizes and co-author statistics are omitted: they do not explain what changed for users of the project.

1. Backend scope, deprecation and dispatch

What Flex and CubeCL do

A backend implements tensor computation on a device. Flex is a Rust CPU backend with multithreading, SIMD and optimized matrix multiplication. CubeCL supplies the common execution infrastructure behind Burn’s CUDA, ROCm, WGPU and CubeCL CPU backends. The distinction is not simply CPU versus GPU: both Flex and CubeCL offer CPU execution, while CubeCL also targets GPU platforms.

These roles are documented in the August-end Flex README and backend/device guide. They describe existing capabilities, not features all introduced in August.

What changed

#5343, merged 11 August, removed burn-candle, its facade features and workspace dependencies, and its publishing job. Existing configurations requesting the removed Candle feature need to change. This is an actual removal, following earlier deprecation.

#5344, also merged 11 August, marked NdArray deprecated in crate metadata, device constructors and backend documentation. Its build and test support remained. In particular, NdArray continued to serve as a reference backend in CubeCL tests.

#5525, merged 31 August, applied the same deprecation approach to LibTorch. Its constructors and documentation now direct users toward CubeCL or Flex, while existing builds and publication continue during the deprecation window.

Evidence and interpretation

The NdArray change retains its reference-test role; the LibTorch change updates migration guidance and annotations. Neither is an immediate removal. The removal of Candle is visible in its merged diff.

Candle configurations must switch backends; NdArray and LibTorch users can continue using them during deprecation. The documentation recommends Flex and CubeCL, but these PRs do not show how many users have migrated or how maintenance costs have changed.

Other routing changes

Dispatch is the layer that routes an operation from the Burn API to a concrete backend. Four changes affected routing.

2. Computation graph capture and QLoRA fixes

Graph capture

#5377, merged 20 August, added burn-capture and the shared graph representation burn_ir::GraphIr. Instead of performing calculations, the backend records operations and their connections. After a recording session (capture scope), public APIs expose the graph and initial tensor values. Another tool can use these to export the model; capture itself does not convert the model to an export format.

The capture implementation and tests include backend_router_operations_are_captured and captures_operations_values_and_explicit_boundaries. These check operation recording and scope boundaries, rather than an external export toolchain.

Two follow-ups on 21 August clarify how tensors can move between capture scopes. #5408 retains initialized values after scope completion, allowing materialized tensors to initialize another capture scope. #5409 updates dispatch tests to distinguish initialized tensors, which can move, from tensors defined by recorded operations. The latter have not been computed and have no materialized data, so their transfer is rejected.

These PRs test graph recording, not an end-to-end export workflow. The reference to burn-onnx in #5408 indicates an intended use; the integration’s readiness needs to be checked against burn-onnx code and tests.

QLoRA correctness and parameter handling

QLoRA trains small adapter matrices while keeping the model’s quantized base weights frozen. Burn already had APIs for this. On 6 August, #5317 filled missing quantized autodiff operations. Quantization disconnects the result from the original graph; dequantization creates a floating-point tensor without gradient tracking. Gradients flow to the trainable adapter, not to the frozen base or through the rounding operation.

The quantization tests compare gradients for an adapter composed over a dequantized base, check that the base receives none, and exercise mixed quantized/float matrix multiplication.

#5311, merged 7 August, generalized the mechanism into Reparameterization and Reparameterizer, preserving the LoRA/QLoRA convenience APIs. A WeightNorm integration test exercises another use of the shared abstraction.

#5362, merged 13 August, creates adapter matrices in persistent memory so they survive temporary memory-pool resizing or rebuilding. #5364, merged 14 August, makes the numeric type (dtype) of adapter matrices explicit over a packed quantized base. #5365, merged later that day, corrects how the base and adapter update are added. A dense base is cast to the adapter dtype when needed; a packed base is dequantized during mixed addition, using the other operand’s dtype. This replaces the separate dequantization step in the earlier version.

The final behavior is covered by packed-base composition and dense-base dtype tests. These tests check parameter-handling correctness, not training speed. Graph capture and QLoRA are independent capabilities, discussed together here for convenience.

Other quantization and module-state changes

Quantization gained a new scale model, and several fixes addressed paths that previously panicked or silently produced wrong values.

Module-state changes affect checkpoint contents and how training mode is passed to nested layers.

3. Runtime failures and model storage

Failure isolation in fusion

#5535, merged 31 August, addresses a panic leaving queue bookkeeping inconsistent and causing subsequent work to fail. The merged implementation associates failure with affected tensor handles using TensorError and Handle::Errored.

Operations whose inputs carry an error are skipped and propagate its root cause. Independent operations may continue. A fused kernel is handled as one unit: a failure affects every tensor it writes to. A subsequent successful write can clear the error flag. This describes failures handled by these execution paths, not recovery from every device fault or process abort.

Evidence: execution ordering and handle implementation/tests. Tests cover errored handles, clearing errors on write, preserving root causes, and releasing failures with their tensors. The later WriteScope redesign is listed under WIP, not credited to August.

On-demand tensor data and atomic BurnpackStore saves

#5349, merged 20 August, allows Tensor::deferred to provide bytes on demand. The writer requests a tensor’s data once, writes it and releases it before requesting the next. The temporary buffer for this data only needs to accommodate the largest tensor, rather than the entire model. This is not a bound on total application RAM or device memory.

BurnpackStore uses Writer::write_to_file_atomic: it writes a temporary file beside the destination and renames it when complete. A data-provider failure does not truncate the existing destination. Ordinary write_to_file retains in-place behavior, so the guarantee must not be generalized to all saving APIs or formats.

The lazy writer tests check provider order, length mismatches, failed writes and rename failures. They include both a_failed_write_leaves_an_existing_file_intact and a_failing_provider_does_truncate_an_existing_file_in_place, documenting the API distinction. The memory test counts live payload-sized allocations to check that data is not all materialized up front. Its measurement excludes ordinary bookkeeping allocations.

Explicit errors when reading tensor data

#5316, merged 18 August by committer date, introduces TensorReadError, structured DataError variants and conversion and vector-extraction methods that return errors. The conversion implementation contains try_to_vec, try_into_vec and dtype-converting variants. This gives callers ways to handle failures explicitly; it does not make every caller panic-free. The same pull request removed implicit Deref access in favor of shape() and dtype() and renamed conversions around try_cast_as and *_data_dtype across the repository, which is why it is marked breaking.

Other model-storage and loading changes

Tensor data now uses a shared representation across the store, record and pack layers. Other changes address error handling, file reading and loading diagnostics.

4. Convolution and compute-kernel performance

#5428, merged 25 August, adds wgrad_im2col as an autotuning candidate for dense convolution weight gradients. The old fallback expressed the gradient as a convolution with a kernel as large as the input feature map. The new path constructs input columns and uses matrix multiplication, at the cost of materializing repeated input data. For a 3×3 kernel the column representation can hold each pixel nine times. It is therefore offered to autotuning rather than imposed on every workload.

Evidence: candidate registration, implementation and convolution tests changed by the PR.

The PR author reports these measurements on Radeon 8060S (gfx1151), Vulkan, f32, batch 4, EfficientNet V2-S backbone at 384×384:

Measurement Before After
3×3 weight gradient, 64→256 channels, 128×128 map 22.10 ms 3.90 ms
Whole backbone backward 871 ms 715 ms
Training step 1.235 s 1.085 s

The approximately 12% training-step reduction in the memo is calculated from the last row. The author’s measurements were not reproduced for this review; the result applies to the stated configuration.

#5361, merged 18 August, rewrites Flex pooling loops. Valid index ranges are computed before the inner loop, removing repeated boundary checks; average pooling also gains SIMD dispatch. The merged implementation supports that mechanism. The merged diff touches only the pooling implementation and carries no benchmark, so no timing is reported here and no whole-model CPU speedup is inferred.

Other forward- and backward-pass convolution changes

Three further changes affect convolution gradients and the forward pass. They use different algorithms for different configurations.

Memory layout and fusion correctness

Fusion combines adjacent operations into one kernel. In the implementation discussed here, convolution output is stored in NHWC order but has a logical NCHW shape: the channel dimension occupies different positions in the two representations. Choosing a compatible layout for fused operations can avoid transposes. The following changes affect that choice, output correctness and read ordering across threads.

Autotuning, CPU kernels and device characteristics

5. PyTorch import and input validation

#5392, merged 27 August, parses and validates serialized tensor strides, which describe the steps through memory between elements. A tensor view can describe logical values whose storage is offset, permuted or expanded; treating its bytes as a contiguous array produces incorrect weights. The updated reader materializes non-contiguous views in logical row-major order while retaining the contiguous fast path. Missing or truncated storage now reports an error instead of creating zero-filled data. Storage validation is lazy, so metadata can still be inspected when an individual tensor cannot be loaded.

Evidence: PyTorch reader and tests with non-contiguous fixtures. This is a specific import-correctness improvement, not proof of compatibility with every PyTorch checkpoint or model architecture.

#5360, merged 24 August, validates TensorData byte length against shape and dtype during serde deserialization, with checked multiplication to reject overflow. Quantized QFloat takes a different storage layout and is excluded from that simple length formula. NdArray construction switches from unchecked to checked array creation, covering callers that bypass serde.

Evidence: validation and regression tests and NdArray construction. Tests include oversized and overflowing shapes. These checks address malformed tensor data; they are not a security audit of the complete loading stack.

6. Training metrics, modules and optimizers

#5320, merged 10 August, adds a flush barrier that waits for queued events to be processed before metric-driven checkpoint selection and early stopping. Without it, asynchronous epoch-end events could remain queued when those decisions queried metrics. Single-device, multi-device and distributed data-parallel (DDP) strategies now wait for pending events when these features are used. The processor implementation and test include flush_waits_for_queued_training_events, checking that previously submitted training and validation events have been processed before return.

#5515, merged 28 August, changes Dice aggregation to accumulate raw numerator and denominator across batches. Averaging individual ratios with batch-size weights can produce a different score, especially when the proportion of foreground pixels varies. The implementation and tests cover uneven batch sizes, empty masks, background exclusion and state reset.

#5517, also merged 28 August, replaces the ROUGE-L final-value todo!() with the accumulated metric value. Its regression test checks final aggregation across unequal batch sizes and state clearing. These fixes concern metric calculation, not the model’s prediction accuracy.

Other module, optimizer and image-processing changes

7. Tensor APIs, indexing and numerical correctness

This section covers API additions and computation fixes not featured in the management memo. They affect specific operations, tensor shapes and numeric types, including incorrect results, crashes and memory-safety violations.

A shared interface for indexed updates

Indexed updates write values at positions specified by an index tensor. A shared interface now replaces specialized backend methods for these operations.

Selection and padding also moved to the backend layer, allowing backends to provide their own implementations:

Empty tensors and special-value fixes

These changes define behavior for empty tensors, zeros, NaNs and other special input values.

Memory safety and aliasing

Test and diagnostic fixes

8. Dependencies, releases, CI and the direct commit

This section covers dependency updates, pre-releases, build changes and documentation. PRs that only update pinned dependency revisions are listed as a group; accompanying code changes are described separately.

CubeCL and CubeK synchronization

Twenty-one PRs updated Burn’s pinned CubeCL and CubeK revisions. Some also required changes to Burn’s integration code:

At the end of August both CubeCL and CubeK were pinned to tracel-ai Git revisions rather than published packages, as the root manifest at the final August revision shows. 0.22.0-pre.2 had temporarily moved both to pre-release packages (#5338); later August updates returned them to Git revisions, and #5476 advanced Burn to 0.22.0-pre.3 in the middle of that sequence. These pre-releases should not be treated as the stable 0.22.0 release. What a version contains is determined by its revision, not by the end of the reporting month.

Build, CI and contributor workflow

Documentation

The Burn Book was revised: backend generics were removed from its examples, the device, learner and module chapters were updated, an optimizer section was added, and the distributed-computing chapter was expanded. Optimizer crate documentation was updated in the same change. #5276 (merged diff)

The direct commit

One change reached main on 10 August without a pull request. It removed a circular publishing dependency by making burn-router depend on a local burn-flex development dependency rather than the workspace dependency. Commit on main.

9. Work in progress at the August cutoff

This section describes branch work that had not reached main by the end of August. These PRs are excluded from the August totals. Later integration dates are shown separately to distinguish the cutoff state from subsequent events.

The state is reconstructed from Git history and saved PR data, not an August-end GitHub snapshot. Commit timestamps do not establish when a branch was first published; review and draft statuses at the cutoff were not verified. Branch size and PR age do not establish readiness or explain delays.

Continuations of August themes that merged in September

The following branches contain work dated before the cutoff, while their integrating commits entered main afterwards.

Initiative Relevance and evidence of pre-cutoff branch work Integration into main, outside August
#5489 — atomic SafeTensors saving Extends protection against partial checkpoint overwrite to another format; a 31 August branch change deals with the shared atomic-file path. 1 September
#5520 — dispatch routing Shared routing machinery and an explicit autodiff context; 31 August branch work addresses dispatch contracts. 1 September
#5540 — Fusion WriteScope Continues failure attribution with a write scope and structured execution results; 31 August branch work predates integration. 2 September
#5259 — batched SVD A separate linear-algebra capability; a 31 August branch change handles convergence errors. 2 September

Six further PRs merged into main in the first weeks of September and continue themes covered here. Their purpose and merge dates are listed below; a detailed review of the September changes is outside this report’s scope:

Branches not integrated as of the data snapshot

The following are not present in main in the data used for this review, whose history ends 22 September 2026 and whose PR metadata snapshot is dated 23 September 2026. These statuses refer to those dates, not the report’s preparation date of 29 September. Review progress was not assessed.

How to use this review

The management memo highlights the main changes and explains their relevance. This review connects those conclusions to PRs, code and tests, and includes work omitted from the shorter memo. All 137 PRs are accounted for, but brief entries describe a change’s main purpose rather than every edit it contains.

The review text does not replace primary evidence. The links allow readers to inspect the implementation; tests and benchmarks need separate reproduction. Assessing costs, adoption and production readiness requires information beyond the public development history.

Privacy

Burn and Tracel AI are trademarks of their respective owners. PRA is not affiliated with Tracel AI.