On September 14, 2026, Lorenzo Stoakes posted a v2 series of 21 patches to the mailing list that speed up Linux kernel builds. The best result is a 36% cut in full-build time on a dual-EPYC server, from 188.0 seconds to 121.1 seconds. But the compiler is not what got faster. Even with hundreds of threads, a "serial tail" remains at the end of a build: dependency parsing, huge intermediate files, work divided into pieces too small to distribute efficiently, and compression. The series removes these one by one. An LLM helped with the search, but what this series shows is less the speed of AI than the human work of turning generated assistance into changes that can go upstream.

AD

What the 21 v2 patches target: the wait after compilation

The changes listed in the v2 cover letter span 47 files. They cover Kbuild, kallsyms, and modpost, and extend to objtool, the Rust build, and gzip compression. The total is 2,982 lines added and 865 lines removed. They look like unrelated small tweaks, but they share one aim: after a large number of C files have been compiled in parallel, one process or a handful of jobs remains, and every CPU waits for it to finish. The series shortens that wait.

In a full build, compilation still accounts for most of the time. In an incremental build that changes a single file, or a build that only checks whether anything changed (a no-op build), the amount of compilation drops sharply. Fixed costs such as reading dependencies, relinking, generating symbols, and compression barely shrink, so their share jumps. In v2, full builds with all modules enabled got 19–36% shorter, while incremental builds got 61–66% shorter and no-op builds 89–95% shorter. That pattern reflects this structure.

So reading the 21 patches as a single speed-up algorithm misses what is going on. They eliminate unnecessary sorting, avoid generating huge text files, stop repeating the same checks, batch jobs that take only milliseconds, and parallelize formerly serial parsing and compression where it is safe. Each gain is small, but cutting the remaining wait from several places means the effect is larger in environments where the post-compilation wait makes up a bigger share of the total.

36% is an upper bound, not a promise for everyone

The 36% figure comes from a 256-core, 512-thread machine with two EPYC 9754 processors, running a full GCC build of an all-modules-enabled configuration (allmodconfig). The time fell from 188.0 seconds to 121.1 seconds, a 66.9-second reduction, which divided by 188.0 seconds and rounded gives 36%. With Clang, the time fell from 259.5 seconds to 184.6 seconds, a 29% reduction. On a 64-core, 128-thread Threadripper 9980X, GCC improved 19% and Clang 22%.

By contrast, a GCC build of the default configuration (defconfig) on an 8-core 2022 M2 MacBook Pro went from 519.3 seconds to 512.4 seconds, an improvement of only 1%. With Clang the improvement was 10%. The v2 full-build improvement was 36% for EPYC with allmodconfig and GCC, but only 1% for the M2 with defconfig and GCC.

These two extremes cannot be treated as a comparison that varies only core count. The machine, architecture, configuration, and toolchain all differ at once. The EPYC side builds a vast set of x86 modules, so generated artifacts and post-processing also grow. The M2's default configuration involves less work, and some of the x86-oriented optimizations do not apply. The gap does not show a causal effect; it shows that even within the same series, the amount of waiting that can be cut varies greatly by environment.

The measurement conditions also deserve attention. The author added a separate series improving srcversion generation and a series from Josh Poimboeuf related to noreturn, and measured with KBUILD_RUST_THREADS=8, -j $(nproc), and pigz installed. The figures are best values, not medians or averages over multiple runs. There is also no per-patch ablation table for the 21 patches. The 36% is an upper bound awaiting reproduction, not a speed guaranteed on a typical development machine.

AD

Eliminating huge generated files and fine-grained jobs

The .cmd files read by GNU make are heavy even when nothing has changed. For large targets such as amdgpu, each object file has over a thousand dependencies, and make parses them serially as make syntax. The depcheck patch uses a small helper program written in C to check the dependency files once each and return to make only the targets that need rebuilding, as small fragments. If the helper fails, the build falls back to the conventional path. In individual v1 measurements, make parsing for amdgpu dropped from 380 ms to 10 ms, and for the whole directory from 2.2 seconds to 0.4 seconds.

Splitting work too finely is also costly. In an x86-64 all-modules-enabled configuration, module finalization can turn into about 22,000 jobs of a few milliseconds each. The finalization patch batches them in groups of 128. If the time spent launching shells, passing environments, and collecting jobs dominates the time the CPU spends on real work, bundling the load is faster than increasing the degree of parallelism further.

The 11,189 *.mod.c files that modpost generates had the same problem. Compiling small C files one by one also brings along header dependency generation and objtool checks. In the v1 measurements this used 6,300 CPU-seconds in total and still took 64 seconds with 128 threads. The .cmd files alone, which are re-read later, reach about 1.3 GiB. The proposed patch generates the required descriptors directly in assembly, removing the detour through the C compiler.

kallsyms was writing about 158,000 symbols out as 37 MiB of assembly. The binary-ization patch outputs the raw data to a 2.6 MiB file and pulls it in with .incbin from a small 9.8 MiB assembly file. The resulting object is reportedly identical. Here too, the patch does not make the computation smarter; it removes the very reason for making the compiler read a huge string representation.

Final compression is also a serial section. In v1 measurements for pigz support, the time to compress the 36 MiB vmlinux.bin with gzip fell from 1.6 seconds to 0.09 seconds. However, although gzip and pigz output can both be decompressed to the same content, the compressed byte sequences are not identical. For reproducible builds, different builders must be pinned to the same compressor. Speed is not free.

v2 kept the speed-up while moving toward safety

v2 went from 23 patches to 21: two were already merged upstream, and two others were removed because a separate series covers equivalent work. The merged ones were mksysmap-related fixes, which became commits 281b61d408d4 and 59351365ac27. Meanwhile, two srcversion-related patches were dropped from the series to avoid duplicate work. Because new patches were added and others split or reordered, the change is not a simple 23 minus 4. Still, it is clear that v2 narrowed its scope through review.

The changes toward safety matter more than the performance numbers. Because the Rust compiler's parallel frontend does not produce reproducible output, it is no longer enabled automatically and is used only when KBUILD_RUST_THREADS is set. In the parallelization of objtool's instruction decoding, part of the destination-resolution work was returned to serial execution, and the shared pv_ops state is now locked. Sashiko's automated checks produced many false positives, but also some valid findings. ThreadSanitizer additionally found a race that Sashiko had not reported.

The author built a wide range of configurations for arm64, arm, RISC-V, powerpc64, s390, and LoongArch, and booted nine architectures under QEMU. They also checked that System.map matches across several architectures, and report having verified the kallsyms self-test, module loading and unloading, and out-of-tree modules. This is strong verification, but it does not guarantee correctness for every configuration or toolchain. That is exactly why each change needs to be split out to the responsible subsystem and pass through the maintainer's judgment.

AD

Third-party reruns confirmed double-digit gains

Measurements by people other than the author were more modest than the 36% maximum. In follow-up tests of v1 gathered on LKML, an EPYC 9454P saw a reduction of about 13%. On an 80-core Ampere Altra, the time went from 6 hours 21 minutes 42 seconds to 5 hours 31 minutes 15 seconds, a difference of 3,027 seconds, or 13.22%. On a 32-core AMD machine it fell from 3 hours 38 minutes 55 seconds to 3 hours 13 minutes 18 seconds, a 11.7% improvement.

The three reruns suggest the series is not a fluke that works only on the author's two machines. On the other hand, each machine had a different configuration and build matrix, and the target was v1, not v2. The set of additional series was also not aligned. What third parties confirmed, therefore, is only the direction: double-digit reductions appear in different environments. It is not an independent reproduction of the 36% on a 512-thread machine.

The "up to 70% incremental" used in the v1 headline also needs updating. In v2's remeasurement, incremental builds with all modules enabled improved by up to 66%, with EPYC/GCC going from 70.9 seconds to 24.3 seconds. The maximum for no-op builds, under the same conditions, is 30.6 seconds down to 1.5 seconds, a 95% reduction. When discussing v2, the old maximum should not simply be carried over.

AI found candidates; humans and checking tools made them adoptable

Stoakes explains that he used the LLM to find bottlenecks, draft improvements, run builds and tests, debug, and analyze results. He has not disclosed the model names, input volume, cost, or number of attempts. All commits carry Assisted-by: LLM, but because much of the generated code was "ugly," he audited and rewrote it himself and verified performance and correctness by hand.

Linus Torvalds's reply neatly expresses this division of responsibility. He said he initially braced himself on seeing the disclosure of LLM use, but that the cleaned-up patches were not as bad as he had feared and looked like small, safe changes. He was also positive about merging them through the appropriate trees. However, he did not apply and measure them himself. He also held back on some points: Rust is not ready yet, and objtool requires approval from its maintainer.

It is confirmed that Stoakes used the LLM to search for candidates. But the patch series alone cannot measure how many hours shorter development was than it would have been without AI, or which changes would not have been found without the LLM. The effect on search speed or development time cannot be evaluated. What can be confirmed is something else: only because humans read the changes, sorted the false positives from automated checks, made up for oversights with dynamic checks, and maintainers judged whether to accept them, did the candidates take a form that can go into Linux.

The Linux 7.4 adoption process and what remains to evaluate

As of September 14, 2026, v2 is under proposal and review, and its adoption into Linux 7.4 has not been decided. What Phoronix reported was an expectation that it "could" go in. The 21 patches may not all go in together; like the already-upstreamed mksysmap fix, they may proceed separately through the trees responsible for Kbuild, objtool, Rust, and so on.

The procedure that can be confirmed for upstream adoption is the judgment of each responsible maintainer, including for objtool. For readers to judge how general the performance gains are, however, a different kind of verification is needed: separating each patch's contribution from measurements that include the additional series, and showing medians and variance rather than just best values. v2 also needs to be rerun on typical development machines and on distribution configurations, and the effect of pigz and Rust parallelization on reproducible builds needs to be checked. These are unresolved evaluation issues, not formal adoption conditions set by the maintainers.

The test is not only whether 36% can be reproduced on a machine with hundreds of threads. The series must also avoid regressions on small development machines and build correctly on multiple architectures. Only when it also reproduces artifacts under identical conditions and consists of changes that maintainers can read over the long term can this speed-up become part of the standard Linux workflow.