Contents
Tags

Apple's keynote is on 9 September. You will hear about the A20 Pro, the 2nm process, the neural engine and how much faster everything is. You will not hear a RAM figure, because Apple has never put one on a keynote slide.
RAM is the number that matters, and this year it is not a subtle point. Reporting ahead of the event says Apple is gating Apple Intelligence features on memory outright.
That last pair is the story. A cloud fallback exists for most of Apple Intelligence, and for these features it explicitly does not. Apple is saying, through its product boundaries, that some things only work if the model is resident in memory on the device.
A phone does not hand its RAM to one process. iOS, the compositor and every background app take their share, so a resident model realistically gets somewhere between a quarter and a third of the total. That fraction is an assumption; the arithmetic on top of it is the same we use for every local model — parameters times bits-per-weight over eight, plus a cache term, plus runtime overhead.
Put another way: three extra gigabytes of RAM buys well under a gigabyte of extra model budget, and that is the entire difference between qualifying for the top tier and not. Memory on a phone is expensive in a way that is easy to miss when a desktop card carries 24 GB.
Expect the keynote to spend real time on the A20 Pro's neural engine. It will be faster, and for on-device generation that matters less than it sounds.
Producing one token means reading every active weight once. That makes generation memory-bound rather than compute-bound — the processor spends most of its time waiting on memory, not calculating. A faster NPU shortens the part that was not the bottleneck. It helps prompt processing, image work and anything compute-heavy; it does not let a bigger model fit.
Which is why the feature gate is expressed in gigabytes rather than in TOPS. Apple knows exactly which number is binding.
The same relationship, on hardware you can actually configure.
See how bandwidth sets speed →Apple will not say it, so it will come from the teardowns and the developer documentation in the days after. When it does, this is how to read it:
Apple has been forced into the trade every local-AI user already makes. A model that fits, doing what a small model does well, with the hard work sent elsewhere. The difference is scale: a 12 GB phone runs something in the 3B to 5B class, while a 24 GB desktop card runs a 27B needing about 18 GB, and frontier models live in datacentres.
The gap between an on-device assistant and a frontier model is not a tuning problem or a software problem. It is roughly two orders of magnitude of parameters, and it is set by how much memory the device has.
One caveat worth stating plainly: every RAM figure above is pre-event reporting rather than confirmed specification. The memory arithmetic is ours and you can check it; the specs are journalism, and Wednesday will settle them.
This post went up before the keynote. The event has happened, and the answer arrived by a route it did not predict — so here is the scorecard, including the part that did not hold.
The distinction Apple is drawing is a fine one and worth naming. It will not tell you how much memory a phone has. It will tell you how much memory a feature needs. Those two facts together give you the answer, and Apple has now supplied the second half.
A company that publishes a memory requirement has conceded the argument that capacity is what decides on-device AI. That was the point of this post, and it is now Apple's position as much as ours.
For balance, our companion piece argued Apple would talk about neural-engine figures and stay quiet about memory bandwidth, which we called the number nobody publishes for phones. Apple opened with it: 50 percent more memory bandwidth than the A19 Pro, from the widest memory interface ever shipped in an iPhone.
So the record for the week is one prediction half-right by an unexpected route, and one flatly wrong. Both landed the same way in the end — with Apple making the argument we had been making, in its own words, about both halves of the memory question.
The full tier breakdown, and why the cliff sits at 12GB.
Read the 12GB piece →And the bandwidth prediction we got wrong.
Read what we got wrong →Tools
Find AI models that fit your exact hardware. Enter your specs and get a ranked list instantly.
Newsletter