Contents
Tags

Apple says the A20 Pro has 50 percent more memory bandwidth than the A19 Pro. It does not say what either figure is in gigabytes per second, and it never has for a phone. So the honest way to use the claim is as a ratio, which turns out to be enough.
Tokens per second is roughly memory bandwidth divided by the bytes read per token, which for a dense model is the size of its weights. Nothing else in that expression changed tonight.
So a 50 percent increase in bandwidth is about a 50 percent increase in generation speed, on the same model at the same quantisation. If a 3B model produced 28 tokens per second on the previous chip, expect something near 42 on this one.
Those absolute figures are illustrative rather than measured — Apple has never published a phone's bandwidth in GB/s, so nobody can give you a real starting number. The ratio is what is confirmed, and the ratio is what carries through to tokens per second.
Capacity. How large a model can be resident is set by memory size, not memory speed, and Apple published no RAM figure. A faster memory interface moves the same 3B-class model more quickly; it does not let a 7B model fit where it did not before.
This is the distinction that makes most phone-AI coverage confusing. Two numbers govern local inference and they are frequently conflated:
Phone memory interfaces do not usually move like this. Generational gains of ten or fifteen percent are ordinary; half again in one step is the kind of change that comes from widening the interface rather than clocking the same one higher — which is precisely what Apple says it did.
For on-device AI it is the most consequential thing in the announcement. It is also, quietly, an admission of where the bottleneck was: you do not widen the memory interface to fix a compute problem.
Everything above uses Apple's stated ratio and our own arithmetic. Where we have used absolute numbers they are labelled illustrative, because no real ones exist yet.
The same relationship, on hardware you can actually measure.
Estimate tokens per second →And why capacity, not speed, is what limits a phone.
Read the on-device memory piece →Tools
Find AI models that fit your exact hardware. Enter your specs and get a ranked list instantly.
Newsletter