Flagship smartphone chips from Apple, MediaTek and Qualcomm all moved to the 2nm generation in September 2026. Apple's A20 Pro, MediaTek's Dimensity 9600 Pro and Qualcomm's Snapdragon 8 Elite Gen 6 family each highlight stronger on-device AI.

However, the figures in each company's announcements, such as "2x," "51% improvement" and "55% improvement," measure different kinds of processing.

To understand the race to run large AI models on smartphones, it is not enough to look at the raw performance of the compute units. You also need to look at which parts of a model are actually run, where the data is kept, and how the heat that is generated is dissipated.

The claim of "support for 30 billion parameters" also does not mean that all 30 billion are computed every time.

Comparing Qualcomm's technical explanation with Apple's model research reveals techniques for running only the necessary parts and for handling large models within limited memory.

Beyond the move to 2nm, the competition is really about design: how a smartphone's limited power and memory are allocated across different AI workloads.

AD

Separating the 2nm-generation headline numbers by what they measure

Apple announced the iPhone 18 Pro, featuring the A20 Pro, on September 9. MediaTek announced the Dimensity 9600 Pro on September 15, and on September 22 Qualcomm launched two products: the higher-end Snapdragon 8 Elite Extreme Gen 6 and the Snapdragon 8 Elite Gen 6.

Both of Qualcomm's products use a 2nm-generation process, but their AI performance is not the same. Qualcomm announcement

Organizing the figures each company presented by what they measure makes the differences clearer.

Product / announcement date Main AI-related claims What is measured / basis of comparison
Snapdragon 8 Elite Extreme Gen 6 (Sept. 22) 35% higher NPU performance; up to 33% better performance per watt NPU performance and power efficiency vs. previous generation
Snapdragon 8 Elite Gen 6 (Sept. 22) 14% higher NPU performance; up to 20% better AI Engine power efficiency NPU performance and overall AI Engine efficiency, evaluated separately
Dimensity 9600 Pro (Sept. 15) 51% faster LLM input processing; 55% better power efficiency in generation; 40% lower power for always-on AI vs. previous generation. Input processing, generation, and always-on operation via a low-power NPU, evaluated separately
A20 Pro (Sept. 9) 2x AI processing capability; 50% more memory bandwidth vs. A19 Pro. Not a direct comparison of actual model response speed

The table is compiled from Qualcomm's higher-end model and standard model pages, MediaTek's announcement and Apple's announcement.

MediaTek states explicitly that its performance figures were measured on its own lab demo devices.

Because the three companies did not compare the same model, input length and conditions, the improvement rates in this table cannot be used on their own to rank performance.

A language model first processes the input, such as a question or conversation history, all at once, and then generates tokens, the text units, one after another.

In other words, the time from handing over a long document to the first reply and the speed at which text is generated once the reply starts are separate measures of performance.

Apple's Core ML technical write-up likewise evaluates the time to first output separately from the number of tokens generated per second afterward.

That explanation concerns a 2024 Mac and is not a measured value for the A20 Pro itself. Still, it offers a clue for reading the numbers that appear in smartphone chip announcements.

The manufacturing timeline also needs to be considered separately.

TSMC's N2 entered mass production in the fourth quarter of 2025 and is the company's first process to use nanosheet transistors. The improved N2P is slated to begin mass production in the second half of 2026.

The factory-side production start described in TSMC's technology overview is a different matter from the announcement timing of these SoC products.

Even under the same "2nm" generation name, circuit design and memory configuration are not identical from product to product.

Even with 30 billion parameters, only about 3 billion are run each time

In a September 10 technical explanation of its next high-end platform, Qualcomm gave an example of a 30-billion-parameter MoE model that uses about 3 billion selected parameters each time it generates a token.

MoE (mixture of experts) is a model structure that selects and uses only the necessary parts from among multiple specialized components, depending on the input.

Compared with running the whole model every time, it keeps the amount of computation per step lower. Hexagon NPU explanation

However, even if fewer parts are used for computation, the rest of the weights do not become unnecessary.

Parameters hold the relationships the model has learned as numerical values, and parts not used at a given moment may be selected for the next input, so they must be kept in storage.

The capacity needed to store the whole model, the memory capacity needed during computation, and the speed of moving data are each separate constraints.

The figure of about 3 billion refers to the size of the selected expert portions; it does not mean the memory required across the smartphone drops to one-tenth.

Qualcomm combines a management mechanism that reads expert portions from flash memory with a cache that keeps frequently used data near the NPU.

The next Hexagon increases shared memory by 50%, widening the room to hold model state and intermediate computation data on the chip.

If the number of trips to external DRAM can be reduced, the compute units spend less time waiting for data, and the power used for transfers is easier to hold down.

That said, "50% more shared memory" does not mean that the RAM capacity installed in the smartphone increases by 50%.

The compute units themselves have also been improved.

Qualcomm explains that the new Element Accelerator handles Transformer processing, the foundation of language models, and speeds up input processing for INT4 models by up to 50%.

INT4 is a scheme that represents numbers in 4 bits.

Smaller numeric representations reduce the amount of data handled, but how well answer quality is preserved varies with the model and the quantization method.

The 50% increase in shared memory and 50% faster input processing are values for the next high-end design explained ahead of its announcement, and they cannot be applied uniformly to all products, including the standard version.

These techniques matter not only when answering a single short question.

In AI that reads documents during a conversation, receives results from apps, and decides on the next step based on them, input processing and generation are repeated many times.

The KV cache, which reuses past computation results, also consumes memory, so speeding up only the compute units cannot eliminate all the waiting time in long conversations.

Both how the model is used and where the data is kept matter.

AD

"55% better efficiency" is not the same as "55% less power"

In the Dimensity 9600 Pro, MediaTek combined the NPU 1090, which handles heavy inference, with a second-generation Super Efficient NPU for always-on operation.

It says power efficiency in generation improved 55% with the former, and power consumption for always-on AI was cut 40% with the latter.

These are numbers for different kinds of processing.

The design separates quickly finishing heavy tasks when needed from sustaining light tasks that run for long periods on little power. MediaTek's September 15 announcement

Assuming the same model and generation conditions, a 55% improvement in power efficiency during token generation works out to roughly a 35.5% reduction in the energy needed to generate one token.

If generation speed per unit of power becomes 1.55 times higher, the energy needed to generate the same amount is the reciprocal, 1 ÷ 1.55.

The reduction is therefore "1 − 1 ÷ 1.55," or about 35.5%.

This is a calculated conversion of the efficiency ratio MediaTek published, not a measured result that smartphone battery life improves by 35.5%.

The announcement does not specify which model was used, what the generation conditions were, or which components were included in the power measurement.

How much real devices improve would need to be measured separately.

For always-on AI, how long processing continues is also important.

Running a large model for a short time and running a small recognition task continuously for a long time load the battery in different ways.

The 40% reduction from the low-power NPU is a figure aimed at improving the latter.

This 40% cannot be treated as a reduction in whole-device standby power that includes the display and connectivity.

Still, if AI is expanding from a feature that waits for user input to one that continuously tracks its surroundings, power consumption during always-on operation becomes an important metric to look at on its own.

The Dimensity 9600 Pro also supports LPDDR6 memory and UFS 5.0 storage, combined with CPU-NPU coordination and 34.5MB of cache.

Both the path that delivers in-process data from memory and the path that loads the model itself from storage have been updated.

However, support for more standards does not determine the capacity or speed actually installed in a given smartphone.

Measuring model startup time separately from generation speed after startup makes it easier to tell which improvements actually affect real-world usability.

Apple changed not only compute performance but also heat and how models are loaded

Apple's A20 Pro has a 32-core Neural Engine in total and doubles AI processing capability over the A19 Pro. Memory bandwidth is also widened by 50%.

In addition, the silicon die and memory are placed side by side, taking the memory out of the path that carries heat from the die to the cooling components.

The A20 Pro is bonded directly to the vapor chamber, whose surface area is tripled from the previous generation. Apple says this improves sustained performance by up to 40%. Apple announcement

A vapor chamber is a component in which an internal liquid repeatedly evaporates and condenses, moving heat across a wide area.

Even if a chip can briefly deliver high performance, it is hard to maintain that speed for long if the heat it generates cannot be dissipated sufficiently.

While raising compute performance and data transfer capability, Apple also changed the path that carries heat away.

However, the tripled surface area is a figure for the cooling component; it does not mean AI inference becomes three times faster.

On the model side, too, the way limited memory is used has changed.

According to a 2026 research write-up on Apple Foundation Models, the on-device model "AFM 3 Core Advanced" has 20 billion parameters in total and activates 1 billion to 4 billion of them depending on the request.

Apple explains that it optimized this model for systems with its most capable Apple silicon.

That said, this research announcement alone does not settle which A20 Pro devices will be able to use it or to what extent it will be offered as features.

The whole model is stored in flash memory, and when an input is received, only the necessary expert portions are selected and loaded into DRAM.

During generation, the parts to use are reselected at regular intervals, combining an always-used shared portion with portions that vary by input.

Apple explains that the reason for this approach is that transferring from flash to DRAM is too slow to swap weights for every token.

So it selects the parts to use at the initial input-processing stage, reducing how often weights are reloaded.

The aim is to keep the overall model large while adjusting how much is loaded into DRAM during actual processing.

However, this explanation alone does not reveal how often Qualcomm's MoE implementation loads weights or how fast it can transfer them.

Qualcomm works on managing expert portions and on shared memory near the NPU.

Apple designs the model side, deciding the unit at which it is loaded.

MediaTek strengthens the paths that carry data, such as LPDDR6 and UFS 5.0.

The approaches differ, but all three tackle the same problem: how to run large AI models within a smartphone's limited memory.

Simply reducing the three companies to one trait each, "large models," "heat dissipation" and "always-on AI," misses this commonality.

Even if model-side techniques reduce DRAM usage, the initial load and transfers when switching the parts in use still take time.

Conversely, even if memory bandwidth is widened, a chip may throttle its performance when heat builds up during long workloads.

Only when the places where the model is stored, where it is computed, and where the heat is released are designed as one flow can on-device AI be used comfortably for long periods.

AD

On real devices, waiting time and power consumption need to be measured separately

To compare devices that can run the same language model, you would first match input length and numeric precision, then measure the time from loading the model to the first reply.

Next, you would measure generation speed after the reply starts, and the energy used during that time.

When comparing products that use different models, you also need to give them the same task and check answer quality, not just waiting time and power consumption.

AI quality cannot be ranked by parameter count alone.

Always-on AI requires yet another kind of evaluation.

You would check how much standby power rises when recognition features are enabled, and whether response speed holds up even when running alongside gaming or photography.

For long AI workloads, you also need to look separately at the speed right after starting and the speed that can be maintained once the device warms up.

If whole-device power consumption, including the display and connectivity, were shown, it would also be easier to judge how much NPU-level improvements affect actual battery life.

2nm-generation smartphone chips are moving toward allocating compute power for heavy AI processing and power for always-on operation in finer detail.

If device makers and app developers implement models and features that support these designs, and show on real devices that long conversations and continuous processing can keep waiting time and battery drain in check, smartphones could move closer to being devices that support multiple tasks continuously, rather than AI devices that merely answer short questions.