Contents
Tags

Two days ago we published a piece arguing that the spec deciding on-device AI is memory bandwidth, and that you would never hear it from Apple. We called it “the other half of the equation, and the one nobody publishes for phones”.
Apple opened with it. From the newsroom release, verbatim: the A20 Pro features “50 percent more memory bandwidth than A19 Pro”, enabled by “the widest memory interface ever shipped in an iPhone”.
The prediction was wrong. The thesis it rested on has just been endorsed by Apple's own marketing department, which is a better outcome than being right.
Straight from the release, so you can check every one:
Generating a token means reading every active weight in the model once. That is a great deal of data movement and comparatively little arithmetic, so the speed at which a phone produces text tracks memory bandwidth rather than compute.
Which means the two headline numbers do very different work. The 50% bandwidth increase translates almost directly into tokens per second on the same model. The doubled neural-engine throughput is a compute claim, and compute was not the constraint on single-stream generation — it helps prompt processing, computational photography and image work, and it does not make text appear twice as fast.
If you take one thing from the keynote for on-device AI, take the 50% and not the 2x. Apple put them in the same paragraph; only one of them governs how fast Siri answers you.
Apple published no RAM figure. It never does, and this release is no exception — there is no unified memory number anywhere in it.
That matters because bandwidth and capacity do different jobs. Bandwidth decides how fast a model runs; capacity decides how large a model can be resident at all. Apple has told us the first improved by half and said nothing about the second, so the ceiling on model size is exactly as unknown tonight as it was this morning. Teardowns usually settle it within a few days.
One more correction while we are at it: several outlets have framed the new neural accelerators as being aimed at running third-party LLMs on device. That framing is the press, not Apple — the release makes no such claim, and we are not going to attribute it to them.
Very little directly, and quite a lot indirectly. A phone is not going to run the models you care about; the memory budget puts it in the 3B class while a 24 GB desktop card reaches 27B.
What has changed is the argument. The largest consumer hardware company in the world just built its chip announcement around memory bandwidth as the thing that makes AI feel fast. That is the same case we make when we tell people to rank GPUs on bandwidth rather than TFLOPs, and it is a much easier case to make tomorrow than it was yesterday.
Rank any device on the number Apple just led with.
Estimate tokens per second →Our pre-event piece, including the prediction that did not survive.
Read what we got wrong →Tools
Find AI models that fit your exact hardware. Enter your specs and get a ranked list instantly.
Newsletter