Companion to the management memo. Prepared 29 September 2026 from the local repository history and saved PR metadata.
Repository: tracel-ai/burn, branch main.
Window:
[2026-08-01T00:00:00-04:00, 2026-09-01T00:00:00-04:00). The
main line of history (first-parent) contains 137
PR-associated commits and one direct commit in this window. Changes are
selected by their committer timestamps on main, not their
author timestamps on feature branches. The period ends at this
main revision.
This review covers all changes merged in the window. Each of the 137 pull requests and the direct commit is assigned to one topic section. Work still in progress at the cutoff and subsequent September merges are described separately and excluded from the August totals.
All changes are covered, at two levels of detail:
main; specific
mechanisms have been checked against the code changes. Not every entry
includes links to individual files.The analysis is based on code and tests at the linked revisions. Tests were read, not run for this review: the text describes what they check, not the results of an independent test run. Production workloads, adoption and financial outcomes were not evaluated. Where a PR description differs from the code, the merged implementation takes precedence. Benchmarks from PRs are attributed to their authors, with measurement conditions where the source provides them.
| Measure | August 2026 |
|---|---|
Pull requests merged into main |
137 |
Direct commits to main |
1 |
| Pull requests marked breaking | 6 |
| Days with at least one merge | 24 |
| Pre-releases | 2 (0.22.0-pre.2,
0.22.0-pre.3) |
The six changes marked as breaking are #5343 (Candle removal), #5262
(quantization scale model), #5290 (NaN propagation in extrema), #5316
(TensorData access APIs), #5400 (fusion and autodiff edge
cases) and #5411 (store transport type). Each is described and linked in
its own section below. The count comes from the breaking marker in the
squash commit’s subject on main.
Contributor rankings, diff sizes and co-author statistics are omitted: they do not explain what changed for users of the project.
A backend implements tensor computation on a device. Flex is a Rust CPU backend with multithreading, SIMD and optimized matrix multiplication. CubeCL supplies the common execution infrastructure behind Burn’s CUDA, ROCm, WGPU and CubeCL CPU backends. The distinction is not simply CPU versus GPU: both Flex and CubeCL offer CPU execution, while CubeCL also targets GPU platforms.
These roles are documented in the August-end Flex README and backend/device guide. They describe existing capabilities, not features all introduced in August.
#5343,
merged 11 August, removed burn-candle, its facade features
and workspace dependencies, and its publishing job. Existing
configurations requesting the removed Candle feature need to change.
This is an actual removal, following earlier deprecation.
#5344, also merged 11 August, marked NdArray deprecated in crate metadata, device constructors and backend documentation. Its build and test support remained. In particular, NdArray continued to serve as a reference backend in CubeCL tests.
#5525, merged 31 August, applied the same deprecation approach to LibTorch. Its constructors and documentation now direct users toward CubeCL or Flex, while existing builds and publication continue during the deprecation window.
The NdArray change retains its reference-test role; the LibTorch change updates migration guidance and annotations. Neither is an immediate removal. The removal of Candle is visible in its merged diff.
Candle configurations must switch backends; NdArray and LibTorch users can continue using them during deprecation. The documentation recommends Flex and CubeCL, but these PRs do not show how many users have migrated or how maintenance costs have changed.
Dispatch is the layer that routes an operation from the Burn API to a concrete backend. Four changes affected routing.
Autodiff wrapper. #5326From<WgpuDevice> resolves through a
feature-priority chain, preserving the conversion when Cargo enables the
Metal, Vulkan and WebGPU features together instead of failing on the
ambiguity. #5323rfft and irfft, so signal operations are
available through it rather than unimplemented. #5513#5377,
merged 20 August, added burn-capture and the shared graph
representation burn_ir::GraphIr. Instead of performing
calculations, the backend records operations and their connections.
After a recording session (capture scope), public APIs expose the graph
and initial tensor values. Another tool can use these to export the
model; capture itself does not convert the model to an export
format.
The capture
implementation and tests include
backend_router_operations_are_captured and
captures_operations_values_and_explicit_boundaries. These
check operation recording and scope boundaries, rather than an external
export toolchain.
Two follow-ups on 21 August clarify how tensors can move between capture scopes. #5408 retains initialized values after scope completion, allowing materialized tensors to initialize another capture scope. #5409 updates dispatch tests to distinguish initialized tensors, which can move, from tensors defined by recorded operations. The latter have not been computed and have no materialized data, so their transfer is rejected.
These PRs test graph recording, not an end-to-end export workflow.
The reference to burn-onnx in #5408 indicates an intended
use; the integration’s readiness needs to be checked against
burn-onnx code and tests.
QLoRA trains small adapter matrices while keeping the model’s quantized base weights frozen. Burn already had APIs for this. On 6 August, #5317 filled missing quantized autodiff operations. Quantization disconnects the result from the original graph; dequantization creates a floating-point tensor without gradient tracking. Gradients flow to the trainable adapter, not to the frozen base or through the rounding operation.
The quantization tests compare gradients for an adapter composed over a dequantized base, check that the base receives none, and exercise mixed quantized/float matrix multiplication.
#5311,
merged 7 August, generalized the mechanism into
Reparameterization and Reparameterizer,
preserving the LoRA/QLoRA convenience APIs. A WeightNorm
integration test exercises another use of the shared
abstraction.
#5362, merged 13 August, creates adapter matrices in persistent memory so they survive temporary memory-pool resizing or rebuilding. #5364, merged 14 August, makes the numeric type (dtype) of adapter matrices explicit over a packed quantized base. #5365, merged later that day, corrects how the base and adapter update are added. A dense base is cast to the adapter dtype when needed; a packed base is dequantized during mixed addition, using the other operand’s dtype. This replaces the separate dequantization step in the earlier version.
The final behavior is covered by packed-base composition and dense-base dtype tests. These tests check parameter-handling correctness, not training speed. Graph capture and QLoRA are independent capabilities, discussed together here for convenience.
Quantization gained a new scale model, and several fixes addressed paths that previously panicked or silently produced wrong values.
ScaleDtype. This
replaced the earlier QuantLevel/QuantParam
model and is one of the month’s breaking changes. #5262linear() with a rank-2 input and a
quantized weight, where a shape-preserving no-op had been rejected. The
added tests dequantize the same tensor in two transformation orders and
check exact agreement. They detect transformation errors, not precision
loss from quantization itself. #5321Module-state changes affect checkpoint contents and how training mode is passed to nested layers.
Param<Flag> gives parameterless but
training-sensitive layers a ParamId, so group traversal
reaches dropout, Gaussian noise and RReLU. Frozen batch normalization
uses stored statistics, container-held modules retain their state,
inference round-trips preserve it, and module formatting exposes it. #5498Param::valid() and Param::from_inner()
lost the parameter mapper, which converts between in-memory and
checkpoint layouts. Because Learner returns its trained
model through valid(), saving that model wrote mapped
parameters as transposed tensors. The file remained valid in format, so
the problem appeared on loading rather than saving. Both methods now
preserve the mapper; regression tests were added. #5509 (merged
diff)ParamGroup. On loading, callers can explicitly allow unused
parameters. #5336GradientCheckpointingStrategy offers Balanced
and Disabled, devices carry a default, individual tensors
may override it, and dispatch retains the strategy but excludes it from
device equality checks. The same PR placed MetricsRenderer
in a Box, restricted cached fusion-plan shape IDs to those
assigned by the plan’s own operations, and fixed NHWC relayout for
read-only tensors. It is marked breaking. #5400#5535,
merged 31 August, addresses a panic leaving queue bookkeeping
inconsistent and causing subsequent work to fail. The merged
implementation associates failure with affected tensor handles using
TensorError and Handle::Errored.
Operations whose inputs carry an error are skipped and propagate its root cause. Independent operations may continue. A fused kernel is handled as one unit: a failure affects every tensor it writes to. A subsequent successful write can clear the error flag. This describes failures handled by these execution paths, not recovery from every device fault or process abort.
Evidence: execution
ordering and handle
implementation/tests. Tests cover errored handles, clearing errors
on write, preserving root causes, and releasing failures with their
tensors. The later WriteScope redesign is listed under WIP,
not credited to August.
#5349,
merged 20 August, allows Tensor::deferred to provide bytes
on demand. The writer requests a tensor’s data once, writes it and
releases it before requesting the next. The temporary buffer for this
data only needs to accommodate the largest tensor, rather than the
entire model. This is not a bound on total application RAM or device
memory.
BurnpackStore uses
Writer::write_to_file_atomic: it writes a temporary file
beside the destination and renames it when complete. A data-provider
failure does not truncate the existing destination. Ordinary
write_to_file retains in-place behavior, so the guarantee
must not be generalized to all saving APIs or formats.
The lazy
writer tests check provider order, length mismatches, failed writes
and rename failures. They include both
a_failed_write_leaves_an_existing_file_intact and
a_failing_provider_does_truncate_an_existing_file_in_place,
documenting the API distinction. The memory
test counts live payload-sized allocations to check that data is not
all materialized up front. Its measurement excludes ordinary bookkeeping
allocations.
#5316,
merged 18 August by committer date, introduces
TensorReadError, structured DataError variants
and conversion and vector-extraction methods that return errors. The conversion
implementation contains try_to_vec,
try_into_vec and dtype-converting variants. This gives
callers ways to handle failures explicitly; it does not make every
caller panic-free. The same pull request removed implicit
Deref access in favor of shape() and
dtype() and renamed conversions around
try_cast_as and *_data_dtype across the
repository, which is why it is marked breaking.
Tensor data now uses a shared representation across the store, record and pack layers. Other changes address error handling, file reading and loading diagnostics.
TensorSnapshot was removed and
burn_pack::Tensor became the single transport object for
metadata and deferred byte sources, with a bridge module performing lazy
conversion at the burn-core/burn-pack seam.
Module container metadata that burn-pack cannot own moves
to a borrowed ModuleContext, and the panic guard moves to
provider construction so a backend panic becomes a typed error. Routing
every materialization through one declared-length check surfaced a
latent PyTorch-reader bug in which a storage offset past the end of the
archive returned f32 zeros whatever the dtype. Marked breaking. #5411 (merged
diff)rejects_data_truncated_into_alignment_padding. #5530 (merged
diff)ModuleSnapshot::apply moves values out of a module
while applying a mapper. A panic in caller code — an adapter
implementation or a backend’s from_data — could leave the
owner attempting to free the moved value a second time. This was
reachable through a safe method on a public trait. An
AbortOnUnwind guard now terminates the process: continuing
would be unsafe because the mapper has consumed the original module and
it cannot be restored. #5488 (merged
diff).bpk, so the path the caller gave is the path that is read.
#5494TensorData directly instead of
reconstructing it from raw bytes on every load. #5407Send/Sync requirements and single-use
behavior of deferred tensor sources are documented at their API
boundary, which is the contract the streaming writer depends on. #5500#5428,
merged 25 August, adds wgrad_im2col as an autotuning
candidate for dense convolution weight gradients. The old fallback
expressed the gradient as a convolution with a kernel as large as the
input feature map. The new path constructs input columns and uses matrix
multiplication, at the cost of materializing repeated input data. For a
3×3 kernel the column representation can hold each pixel nine times. It
is therefore offered to autotuning rather than imposed on every
workload.
Evidence: candidate registration, implementation and convolution tests changed by the PR.
The PR author reports these measurements on Radeon 8060S (gfx1151), Vulkan, f32, batch 4, EfficientNet V2-S backbone at 384×384:
| Measurement | Before | After |
|---|---|---|
| 3×3 weight gradient, 64→256 channels, 128×128 map | 22.10 ms | 3.90 ms |
| Whole backbone backward | 871 ms | 715 ms |
| Training step | 1.235 s | 1.085 s |
The approximately 12% training-step reduction in the memo is calculated from the last row. The author’s measurements were not reproduced for this review; the result applies to the stated configuration.
#5361, merged 18 August, rewrites Flex pooling loops. Valid index ranges are computed before the inner loop, removing repeated boundary checks; average pooling also gains SIMD dispatch. The merged implementation supports that mechanism. The merged diff touches only the pooling implementation and carries no benchmark, so no timing is reported here and no whole-model CPU speedup is inferred.
Three further changes affect convolution gradients and the forward pass. They use different algorithms for different configurations.
Fusion combines adjacent operations into one kernel. In the implementation discussed here, convolution output is stored in NHWC order but has a logical NCHW shape: the channel dimension occupies different positions in the two representations. Choosing a compatible layout for fused operations can avoid transposes. The following changes affect that choice, output correctness and read ordering across threads.
RefLayoutSetting::Any is restricted to dense layouts. A
fused kernel treats its position as an offset into the reference buffer,
which only agrees with the tensor when that buffer is densely packed; on
a padded one the kernel left the tail of the tensor unwritten and could
overwrite real elements. CUDA and ROCm pad rows through CubeCL’s pitched
allocator, which is why the fault reproduced there and not on WGPU,
Metal or CPU. #5395 (merged
diff)ReadPlan, so
fusion composition and autotune keys no longer depend on thread timing.
#5282MatVec and VecMat was intended to avoid a
worker-stack failure on tall matrices (#5423). Because
Gemm was registered in that group without being one of
those kinds, it lost its CPU priority too, and since strided pointwise
convolutions now route through matmul, this also affected them. The
author reported a ResNet-50 CPU performance regression of 32% at batch 1
and 42% at batch 8; the processor is not identified. The subsequent fix
keeps GEMV restricted to vector cases and registers Gemm
separately at high priority for MatmulKind::General. #5538 (merged
diff)conv_plane_accumulate and bias addition run through
macerator::with_simd rather than relying on incidental
autovectorization. #5267int_cast
code size. #5357autotune_bounds interface instead of
duplicating per-kernel logic. Without such a bound a tuning run has no
time limit and cannot short-circuit. #5188in_channels twice,
collapsing distinct configurations onto one key. #5289MemoryPoolLayout, with integration tests at the
burn crate level. #5425#5392, merged 27 August, parses and validates serialized tensor strides, which describe the steps through memory between elements. A tensor view can describe logical values whose storage is offset, permuted or expanded; treating its bytes as a contiguous array produces incorrect weights. The updated reader materializes non-contiguous views in logical row-major order while retaining the contiguous fast path. Missing or truncated storage now reports an error instead of creating zero-filled data. Storage validation is lazy, so metadata can still be inspected when an individual tensor cannot be loaded.
Evidence: PyTorch reader and tests with non-contiguous fixtures. This is a specific import-correctness improvement, not proof of compatibility with every PyTorch checkpoint or model architecture.
#5360, merged 24 August, validates TensorData byte length against shape and dtype during serde deserialization, with checked multiplication to reject overflow. Quantized QFloat takes a different storage layout and is excluded from that simple length formula. NdArray construction switches from unchecked to checked array creation, covering callers that bypass serde.
Evidence: validation and regression tests and NdArray construction. Tests include oversized and overflowing shapes. These checks address malformed tensor data; they are not a security audit of the complete loading stack.
#5320,
merged 10 August, adds a flush barrier that waits for queued events to
be processed before metric-driven checkpoint selection and early
stopping. Without it, asynchronous epoch-end events could remain queued
when those decisions queried metrics. Single-device, multi-device and
distributed data-parallel (DDP) strategies now wait for pending events
when these features are used. The processor
implementation and test include
flush_waits_for_queued_training_events, checking that
previously submitted training and validation events have been processed
before return.
#5515, merged 28 August, changes Dice aggregation to accumulate raw numerator and denominator across batches. Averaging individual ratios with batch-size weights can produce a different score, especially when the proportion of foreground pixels varies. The implementation and tests cover uneven batch sizes, empty masks, background exclusion and state reset.
#5517, also
merged 28 August, replaces the ROUGE-L final-value todo!()
with the accumulated metric value. Its regression
test checks final aggregation across unequal batch sizes and state
clearing. These fixes concern metric calculation, not the model’s
prediction accuracy.
Tensor::no_grad() disables gradient tracking without
moving data. #5319softplus computed exp before
log; the exponential overflows above roughly 88.7 in f32,
or around x = 18 with beta 5 because beta scales the input.
This produced a NaN gradient that propagated into shared parameter
gradients and corrupted the update step. SoftMarginLoss
inherited the fault for strongly incorrect predictions. The threshold
for switching to a linear approximation defaults to 20.0, as in PyTorch,
and is now also capped at the safe exponential range of the input dtype
— about 11.09 for f16. Using log1p instead of
log(1 + …) avoids rounding small results to zero at large
negative inputs. softplus(tensor, beta) keeps its
signature; softplus_with_threshold is the explicit form, so
this is not a breaking change. Forward and autodiff tests were added;
the function previously had no autodiff tests. #5374mean_dim, avoiding half-precision
underflow for large groups. #5410AdaptiveAvgPool3d was added as a module, and CubeCL’s
incomplete host fallback was replaced with explicit stubs pending a
compute kernel. The CubeCL implementation itself is not August work; see
the WIP section. #5115build methods became public so grouped
modules can construct optimizers through
ModuleOptimizer::with_group. #5314filter2d, and configurable
padding through Transform2D::with_padding_mode. #5339filter2d kernels retain the input height and
width by using asymmetric padding where necessary. #5414This section covers API additions and computation fixes not featured in the management memo. They affect specific operations, tensor shapes and numeric types, including incorrect results, crashes and memory-safety violations.
Indexed updates write values at positions specified by an index tensor. A shared interface now replaces specialized backend methods for these operations.
IndexingUpdateOp, and
Assign joined Add as a supported operation.
Generic float_scatter and float_select_assign
backend methods route Add to the existing add
implementation and refuse anything a backend has not implemented, so
support could be added per backend. NdArray implements
Assign for floats and integers, autodiff gains its backward
pass — the value tensor receives the gradient at overwritten positions
while the base’s gradient is zeroed there — and the IR, dispatch, fusion
and router layers carry the operation. The public documentation states
that duplicate indices are undefined for Assign in both the
forward result and the gradients. #5245 (merged
diff)Assign was then implemented for CubeCL, Flex and
LibTorch, and the per-backend skips were removed from the shared test
suite, so one set of tests covers the operation everywhere. #5532 (merged
diff)scatter_add and select_add backend methods
were removed in favor of scatter and
select_assign, which take the update operation as an
argument. This removes duplicated code from the backend traits and seven
implementations. #5523Selection and padding also moved to the backend layer, allowing backends to provide their own implementations:
mask_select became a backend operation with a default
that composes argwhere and select, so LibTorch
can use its native selection and the Float implementation preserves
QFloat quantization instead of dequantizing. The public API
and its behavior are unchanged. #5283float_pad and int_pad
with Constant, Reflect and Edge modes, which also makes it available to
fusion (#5505), and
padding preserves the input dtype (#5386).These changes define behavior for empty tensors, zeros, NaNs and other special input values.
false for any, 1 for product,
true for all, and NaN for a floating-point
mean, matching NumPy and PyTorch. Extrema and integer mean return an
error. The rule depends on the reduced axis length, not whether the
output tensor is empty, so max_dim behaves the same for
shapes [0, 0] and [3, 0]. Two pre-existing CPU
inconsistencies were fixed in the same change: Flex returned the
accumulator’s initial infinity value for empty f16/bf16 extrema, and
NdArray accepted an empty axis for one shape while panicking on another.
Empty-axis coverage was added to the shared suite for eight reductions
over both output shapes. #5333 (merged
diff)[0] against [1] gave
[1]. That overstated the element count, took LibTorch’s
in-place fast path against a one-element buffer and aborted, which any
backward() over a zero-sized tensor hit. The shape is now
zero where either operand’s dimension is zero, and the in-place path is
gated on shape equality rather than element count. #5305 (merged
diff)mask_select
backward over an all-false mask. Broadcasting now computes the shared
non-one size, and each broadcast launcher returns an empty output
directly rather than entering a kernel that assumes a non-empty one. #5335 (merged
diff)topk rejects a k larger than the selected
dimension before reaching a backend. #5307topk gained a backward pass that scatters selected
gradients to their source positions. #5531grad * prod(x) / x_i, producing NaN at
zero-valued input positions. The replacement counts zeros per reduced
group and selects among three cases without a host-side branch: no zero
gives the ordinary quotient, exactly one zero gives the product of the
other elements at that position and zero elsewhere, and two or more
zeros give an all-zero gradient. Tests cover a broadcast upstream
gradient and groups holding zero, one and two zeros, reduced along both
an interior and the last dimension. #5534 (merged
diff)float_sign uses libm::copysign. The
previous three-way branch compiled into a constant-pool lookup that some
targets cannot lower, which broke the whole crate’s build for
xtensa-esp32s3-none-elf whether or not sign
was called. The zero branch is retained because
copysign(1.0, -0.0) is -1.0. The author
reports verifying equivalence over all 2³² f32 bit patterns and 20
million f64 patterns plus edge cases, and adds a negative-zero
regression test that the previous suite lacked. #5331 (merged
diff)UnsafeSharedRef::get copied a
stored &mut T. Under the aliasing model checked by
Miri, this invalidated previously issued views while Rayon workers
continued writing through them. The fault affected convolution, pooling,
interpolation, matrix multiplication and grid sampling. The cell now
stores RawArrayViewMut, and each call creates an
independent view without an intermediate &mut reference
to the entire array. The author reports Miri flagging the invalid write
on the previous implementation and being clean on the new one, with a
handles_stay_valid_while_another_is_alive test under Miri.
The safety contract was also widened from disjoint writes to disjoint
accesses, since a concurrent read and write are equally unsound. #5346 (merged
diff)permute and swap_dims return
views over the parent buffer but built the child with a fresh storage
handle, so the reference count reported the buffer as exclusively owned.
A binary operation between a tensor and a view of itself then took
LibTorch’s in-place path over overlapping memory and hit LibTorch’s own
overlap assertion. The child is now built from the existing storage,
which shares the handle when the buffer is aliased. flip is
deliberately unchanged, because LibTorch’s flip allocates
rather than returning a view. #5376 (merged
diff)debug_assert_eq! in grad_replace
required the replacement gradient to carry the same checkpointing
strategy as the tensor it replaced, which only holds for a gradient
obtained from grad(). The public API takes an inner-backend
tensor, so other ways of constructing a gradient could trigger the
assertion incorrectly. CI used --release, which disables
this check. The author reports that removing it resolved two NdArray
debug-test failures. #5294 (merged
diff)This section covers dependency updates, pre-releases, build changes and documentation. PRs that only update pinned dependency revisions are listed as a group; accompanying code changes are described separately.
Twenty-one PRs updated Burn’s pinned CubeCL and CubeK revisions. Some also required changes to Burn’s integration code:
MemorySpec representation for memory-throughput
measurement modes, replacing four separate variants.BlockLevel::BlockTensor,
allowing Burn to compile before support was implemented.At the end of August both CubeCL and CubeK were pinned to
tracel-ai Git revisions rather than published packages, as
the root
manifest at the final August revision shows.
0.22.0-pre.2 had temporarily moved both to pre-release
packages (#5338); later
August updates returned them to Git revisions, and #5476 advanced
Burn to 0.22.0-pre.3 in the middle of that sequence. These
pre-releases should not be treated as the stable 0.22.0
release. What a version contains is determined by its revision, not by
the end of the reporting month.
tracel-ai/github-actions workflows moved
from v9 to v11 for the main test, GPU, Valgrind and vulnerability jobs
(#5390), for
publishing (#5391) and for
the remaining reference (#5398).tracel-xtask moved to v5.0.0, replacing its procedural
dispatch macro with explicit command variants and handlers. #5404cargo run-checks now runs formatting, typo
checks, audit, lint, a quick host no_std check and Flex
tests by default, with --backend to select another; the
broader documentation and target matrix stays in CI. #5356topk example (#5302). Chained
reduction examples (#5301) and the
equivalent ordering and sorting examples (#5298) were
corrected so a demonstration does not reuse an already-reduced
variable.h2 version was updated (#5381).The Burn Book was revised: backend generics were removed from its examples, the device, learner and module chapters were updated, an optimizer section was added, and the distributed-computing chapter was expanded. Optimizer crate documentation was updated in the same change. #5276 (merged diff)
One change reached main on 10 August without a pull
request. It removed a circular publishing dependency by making
burn-router depend on a local burn-flex
development dependency rather than the workspace dependency. Commit
on main.
This section describes branch work that had not reached
main by the end of August. These PRs are excluded from the
August totals. Later integration dates are shown separately to
distinguish the cutoff state from subsequent events.
The state is reconstructed from Git history and saved PR data, not an August-end GitHub snapshot. Commit timestamps do not establish when a branch was first published; review and draft statuses at the cutoff were not verified. Branch size and PR age do not establish readiness or explain delays.
The following branches contain work dated before the cutoff, while
their integrating commits entered main afterwards.
| Initiative | Relevance and evidence of pre-cutoff branch work | Integration into main, outside August |
|---|---|---|
| #5489 — atomic SafeTensors saving | Extends protection against partial checkpoint overwrite to another format; a 31 August branch change deals with the shared atomic-file path. | 1 September |
| #5520 — dispatch routing | Shared routing machinery and an explicit autodiff context; 31 August branch work addresses dispatch contracts. | 1 September |
| #5540 — Fusion WriteScope | Continues failure attribution with a write scope and structured execution results; 31 August branch work predates integration. | 2 September |
| #5259 — batched SVD | A separate linear-algebra capability; a 31 August branch change handles convergence errors. | 2 September |
Six further PRs merged into main in the first weeks of
September and continue themes covered here. Their purpose and merge
dates are listed below; a detailed review of the September changes is
outside this report’s scope:
conv1d and conv2d (3
September).AdaptiveAvgPool3d kernel whose host fallback
#5115 removed (8 September).burn-signal
extension crate (17 September).The following are not present in main in the data used
for this review, whose history ends 22 September 2026 and whose PR
metadata snapshot is dated 23 September 2026. These statuses refer to
those dates, not the report’s preparation date of 29 September. Review
progress was not assessed.
argwhere, nonzero and
mask_select, continuing the backend-operation work in
#5283.TensorData constructors, continuing the API migration
in #5316.The management memo highlights the main changes and explains their relevance. This review connects those conclusions to PRs, code and tests, and includes work omitted from the shorter memo. All 137 PRs are accounted for, but brief entries describe a change’s main purpose rather than every edit it contains.
The review text does not replace primary evidence. The links allow readers to inspect the implementation; tests and benchmarks need separate reproduction. Assessing costs, adoption and production readiness requires information beyond the public development history.
Burn and Tracel AI are trademarks of their respective owners. PRA is not affiliated with Tracel AI.