Developer StayLameBro has published an experiment in which an iPhone 17 Pro Max, connected to an M4 Pro MacBook Pro with 24GB of memory, speeds up local AI prompt processing by up to 44%. Using custom software called "Backburner" to run Qwen3.8-27B, the setup hands part of the computation the Mac would normally do for input processing over to the iPhone. In the developer's measurements, the time to read about 2,000 tokens of additional material, with 16K tokens of context already held, fell from 18.8 seconds to 13.1 seconds.
The speed at which answers are generated did not improve by 44%, however. Adding a spare iPhone is most likely to help when you want to cut the wait for reading large amounts of material, or when memory runs short for holding a long conversation.
What does a 44% speedup mean in seconds saved?
The measurements on October 1, 2026 used Qwen3.8-27B quantized in IQ4_XS format, an M4 Pro MacBook Pro with 24GB of memory, and an iPhone 17 Pro Max. The two were connected with a USB-C cable supporting 10Gb/s. The comparison is between running the same Backburner on the Mac alone and running it with the iPhone added, so it should be kept separate from any effect of switching from a general-purpose inference program to Backburner.
The measurement covers input processing speed when new files or tool outputs are added to an already saved conversation. An LLM first reads and computes over the input, then generates its answer step by step. The "up to 44%" figure refers to the former. In the public logs, the added input is 2,048 tokens, and the measurement script is set to perform only input processing, without generating an answer.
| Context already held | Mac-only input processing | Input processing with iPhone | Change in reading time |
|---|---|---|---|
| 16K | 109 tokens/sec | 157 tokens/sec | 18.8 sec → 13.1 sec |
| 32K | 101 tokens/sec | 130 tokens/sec | 20.3 sec → 15.8 sec |
| 48K | 87 tokens/sec | 113 tokens/sec | 23.5 sec → 18.1 sec |
The figures are rounded from the results published by the developer. K denotes the number of tokens of context already held, and input processing was run twice under each condition. The speeds and times shown compare input processing only; they do not include the time to restore a saved conversation or to generate an answer.
Under the 16K condition measured on October 1, input processing speed rose about 44%, which shortened reading time by about 30%. The speed gain is 157 ÷ 109 − 1, or about 44%, and the time reduction is (18.8 − 13.1) ÷ 18.8, or about 30%. These are the same improvement expressed against different baselines, so "44% faster" cannot be read as "44% less waiting."
The original post also includes a correction: the speed shown on the iPhone screen was a figure for only the layers that device handles. The comparison above uses the numbers the developer gave as the input processing speed of the whole system. The raw logs and measurement procedure are public, but this is not performance that has been confirmed by independent replication.
The iPhone handles the second half of the model
For context up to 64K, the Mac handles layers 1–40 of the model and the iPhone handles layers 41–64. The Mac splits the input into small chunks of 256 tokens, processes them, and sends the intermediate data to the iPhone. While the iPhone's GPU computes the second half of the model, the Mac starts on the next chunk of input. Running the two devices in parallel shortens the time needed to finish processing a long input.
The whole model's data is not re-sent over the cable each time. The iPhone holds the model data for the layers it is responsible for and receives the intermediate data needed for newly processed input. However, the design has the Mac process the small leftover portion of input at the end of each batch and keep the final output on the Mac side.
The matrix-computation capability in the A19 Pro's GPU is also used. According to the developer, with this feature enabled the iPhone's share of the processing ran 2.4 times faster than with it disabled. This compares speeds inside the iPhone only; it does not mean the whole system combined with the Mac became 2.4 times faster.
Apple describes the A19 Pro as having a Neural Accelerator in each GPU core, plus a separate 16-core Neural Engine. What Backburner uses for input processing up to 64K is the GPU; the Neural Engine is used in a separate process for handling long context, described below.
Faster answer generation comes from Mac-side optimization, not the iPhone
With 27K–33K of context, the developer measured answer generation at 11.3 tokens/sec for standard llama.cpp, 25.0 tokens/sec for Backburner on the Mac alone, and 25.1 tokens/sec with the iPhone added. The large speedup comes from changes in Backburner itself, while adding the iPhone to the same Backburner made almost no difference.
Backburner is software derived from llama.cpp. It uses SME2, a matrix-computation feature in the Mac's CPU, and adds its own optimizations to GPU processing. It also combines speculative decoding, in which a small model proposes candidate output first and the large model verifies it. The developer explains that the main reason answer generation is faster with context under 64K is these Mac-side optimizations.
The iPhone also sees little use when inputs are short. In the current implementation, work sharing begins only for input processing beyond roughly 512 tokens. In a session the developer actually used, the iPhone took part in 7 of 36 requests, but those 7 accounted for about 83% of all tokens read. The configuration is better suited to AI agents that read large code files and tool outputs than to repeatedly asking short questions.
In other words, even if answer generation feels faster than with a general-purpose inference program, that improvement cannot be attributed to adding the iPhone. To check the effect of the iPhone, you need to use the same Backburner and compare with and without it under identical input conditions.
Beyond 64K, the iPhone holds past context
When context exceeds 64K, the iPhone's role changes. Rather than computing the second half of the model, it holds the KV cache for older context and takes on the computation that refers to that part. The Mac processes all 64 layers of the model and moves part of the KV cache corresponding to past context to the iPhone.
The KV cache stores computation results that are reused when referring to past input, and the memory it needs grows as the conversation gets longer. The iPhone not only stores the old KV cache but also performs the attention computation for that portion and merges the result with the Mac-side results.
For input processing it uses the iPhone's GPU, and during answer generation the Neural Engine also handles some of the computation. The design treats the already fixed past Keys and Values as fixed weights of a model for the Neural Engine, thereby sharing the work.
In the developer's configuration, the longest 8-bit context confirmed to work on the 24GB Mac alone was 64K. With the iPhone added, operation up to 128K at 8-bit was confirmed. The figure of 196K–229K, calculated from the iPhone's free memory at startup, should be distinguished from context lengths actually verified to work. A separate test run at 140K quantized the context to 4-bit.
IQ4_XS, which compresses the model weights, is a separate setting from the 8-bit and 4-bit context quantization discussed here. In the speed comparison beyond 64K, the Mac alone used 4-bit while the iPhone-assisted run used 8-bit, so it is not a completely like-for-like speedup comparison.
The benefit of adding an iPhone is not only speed: it also lets a higher-precision KV cache be spread outside the Mac's memory. That alone, however, has not been shown to improve answer quality.
In the 128K test, the developer says all three pieces of information embedded at distant positions in the context were retrieved. At 140K, a 32-token output matching the one produced by the Mac alone was also confirmed. Both are limited tests by the developer and do not guarantee that information at any position in a long document can be referenced accurately.
The Mac-side Neural Engine is not used in this implementation. In the developer's tests, enabling it slowed answer generation by 26%, which the developer attributes to memory bandwidth shared with the GPU. This shows that adding more compute units does not always speed things up; the design has to account for memory bandwidth and the division of work.
The iPhone has to be dedicated as a compute resource
Backburner is a pre-release version with code and setup instructions published, and App Store distribution is among the plans. On the Mac side, about 24GB of downloads, including the model, is needed, and the second-half model data stored on the iPhone is about 5.1GB. Simply connecting by USB-C is not enough; a dedicated app must be installed and the model data prepared.
The installation methods provided are building from Xcode and using AltStore. The developer has confirmed that an app installed via Xcode can use about 6GB of memory even with a free Apple account. For AltStore, however, a version that retains the entitlement to raise the memory usage limit is needed, and it is stated that the developer has not yet confirmed installation that way. That the steps are published is not the same as every installation method having been verified to work.
While in use, the iPhone app must be kept in the foreground. In particular, beyond 64K, while the iPhone holds part of the KV cache, locking the screen or unplugging the cable so that it stops responding for 15 seconds causes the server to stop, and you must reopen the app and restart the server.
If a problem occurs on the iPhone side during short-context input processing, the Mac can redo the work. But once part of the KV cache has been moved to the iPhone in a long context, processing cannot be returned to the Mac alone in the same way. Only one request can be handled at a time.
For people who already own a Mac and a compatible iPhone and use local AI with long code or documents, this is an interesting experiment in shortening input-processing waits with equipment already at hand. On the other hand, it is hard to reconcile time spent using the iPhone as an ordinary phone with time it is tied up as an AI compute resource and KV cache store.
If long-running stable operation is confirmed across multiple devices and installation methods, and the benefits of shorter input processing and longer context outweigh the inconvenience of tying up the iPhone, this could become an option for broadening what local AI can do on a 24GB Mac.
