← Blog/Apple Just Put the Number We Said They'd Hide On Stage
deep-dive
Runyard Team
@runyard_dev
8 min read

Tags

#apple#a20-pro#on-device#bandwidth#iphone-18
Runyard.dev — Find AI Models That Run on Your Hardware

Apple Just Put the Number We Said They'd Hide On Stage

A20 Pro specifications: 50% more memory bandwidth, dual 16-core Neural Engine, 2x AI processing, 7-core GPU, 2-nanometre process
Apple's own numbers. The one that governs on-device generation is the big one.

Two days ago we published a piece arguing that the spec deciding on-device AI is memory bandwidth, and that you would never hear it from Apple. We called it “the other half of the equation, and the one nobody publishes for phones”.

Apple opened with it. From the newsroom release, verbatim: the A20 Pro features “50 percent more memory bandwidth than A19 Pro”, enabled by “the widest memory interface ever shipped in an iPhone”.

The prediction was wrong. The thesis it rested on has just been endorsed by Apple's own marketing department, which is a better outcome than being right.

What Apple actually announced

Straight from the release, so you can check every one:

  • “50 percent more memory bandwidth than A19 Pro”, from “the widest memory interface ever shipped in an iPhone”.
  • “A new Dual 16-core Neural Engine”, with “32 total cores for double the AI processing power of A19 Pro”.
  • “The new 6-core CPU features integrated Neural Accelerators”.
  • “The new 7-core GPU design is up to 40 percent faster than A19 Pro”.
  • Built on 2-nanometre process technology — the first 2nm smartphone chip.
  • “Siri AI is an entirely new version of Siri powered by Apple Intelligence”.

Why the bandwidth line is the important one

Generating a token means reading every active weight in the model once. That is a great deal of data movement and comparatively little arithmetic, so the speed at which a phone produces text tracks memory bandwidth rather than compute.

Which means the two headline numbers do very different work. The 50% bandwidth increase translates almost directly into tokens per second on the same model. The doubled neural-engine throughput is a compute claim, and compute was not the constraint on single-stream generation — it helps prompt processing, computational photography and image work, and it does not make text appear twice as fast.

If you take one thing from the keynote for on-device AI, take the 50% and not the 2x. Apple put them in the same paragraph; only one of them governs how fast Siri answers you.

What we still do not know

Apple published no RAM figure. It never does, and this release is no exception — there is no unified memory number anywhere in it.

That matters because bandwidth and capacity do different jobs. Bandwidth decides how fast a model runs; capacity decides how large a model can be resident at all. Apple has told us the first improved by half and said nothing about the second, so the ceiling on model size is exactly as unknown tonight as it was this morning. Teardowns usually settle it within a few days.

One more correction while we are at it: several outlets have framed the new neural accelerators as being aimed at running third-party LLMs on device. That framing is the press, not Apple — the release makes no such claim, and we are not going to attribute it to them.

What this means if you run models yourself

Very little directly, and quite a lot indirectly. A phone is not going to run the models you care about; the memory budget puts it in the 3B class while a 24 GB desktop card reaches 27B.

What has changed is the argument. The largest consumer hardware company in the world just built its chip announcement around memory bandwidth as the thing that makes AI feel fast. That is the same case we make when we tell people to rank GPUs on bandwidth rather than TFLOPs, and it is a much easier case to make tomorrow than it was yesterday.

Rank any device on the number Apple just led with.

Estimate tokens per second

Our pre-event piece, including the prediction that did not survive.

Read what we got wrong

RUNYARD.DEV

Hardware-aware AI model discovery. Know exactly what runs on your machine — before you download.

© 2026 RUNYARD.DEV — All rights reserved.

Built for local AI.

Tools

Try Runyard

Find AI models that fit your exact hardware. Enter your specs and get a ranked list instantly.

Newsletter