All tools
Hardware · free, no sign-up

How much unified memory do I need on a Mac?

Apple Silicon shares one memory pool between CPU, GPU and the system, which makes capacity cheap and bandwidth the thing that separates the tiers.

3 inputs5 questions answeredUpdated for 2026 hardware
Mac Unified Memory Calculator — How much unified memory do I need on a Mac?
Answer first

The short answer

Computed from this tool’s default settings — model size and the rest as most people start. Change them below for your own case.

Smallest Mac that fits16 GB

A 8B model at Q4_K_M needs 6.89 GB, so 16 GB is the smallest configuration that holds it. Step up a chip tier for bandwidth rather than capacity if speed matters more than size.

The calculator

Mac Unified Memory Calculator

What each memory configuration actually runs, and how fast.

Your setup
8,192
Smallest Mac that fits16 GB

A 8B model at Q4_K_M needs 6.89 GB, so 16 GB is the smallest configuration that holds it. Step up a chip tier for bandwidth rather than capacity if speed matters more than size.

Model needs6.89 GB
QuantisationQ4_K_M
Expected speed3 tok/s
MachineUsableBandwidthSpeed
Raspberry Pi 5 · 8 GB4.80 GB17 GB/sWill not fit
Jetson Orin Nano Super · 8 GB5.60 GB102 GB/sWill not fit
iPhone · A19 · 8 GB3.60 GB68 GB/sWill not fit
iPhone · A18 Pro · 8 GB3.60 GB60 GB/sWill not fit
iPhone · A19 Pro · 12 GB5.40 GB77 GB/sWill not fit
Snapdragon 8 Elite · 12 GB5.40 GB77 GB/sWill not fit
Raspberry Pi 5 · 16 GB9.60 GB17 GB/s3 tok/s
Snapdragon 8 Elite · 16 GB7.20 GB77 GB/s14 tok/s
MacBook Air M4 · 16 GB12.0 GB120 GB/s21 tok/s
MacBook Air M4 · 24 GB18.0 GB120 GB/s21 tok/s
MacBook Pro M4 Pro · 48 GB36.0 GB273 GB/s49 tok/s
MacBook Pro M4 Pro · 64 GB48.0 GB273 GB/s49 tok/s
MacBook Pro M4 Max · 128 GB96.0 GB546 GB/s97 tok/s
Inputs

What each setting changes

Every input moves the result for a reason. This is what each one does and where to find the value for your own machine.

SettingDefaultWhat it changes
Model size8B6 options, from 3B to 405B.
Context length8192 tokensAnywhere from 1,024 to 131,072 tokens.
QuantisationQ4_K_M — 4.5 bits/weight5 options, from Q8_0 — 8.5 bits/weight to Q3_K_M — 3.5 bits/weight.
Worked examples

Real answers across model size

The same calculation run at a range of settings, with everything else left at its default. These are computed by the tool itself, not written by hand.

Model sizeSmallest Mac that fitsModel needsQuantisationExpected speed
3B8 GB3.54 GBQ4_K_M8 tok/s
8B16 GB6.89 GBQ4_K_M3 tok/s
14B16 GB10.7 GBQ4_K_M12 tok/s
32B48 GB21.8 GBQ4_K_M12 tok/s
70B64 GB44.5 GBQ4_K_M6 tok/s

A 3B model at Q4_K_M needs 3.54 GB, so 8 GB is the smallest configuration that holds it. Step up a chip tier for bandwidth rather than capacity if speed matters more than size.

Method

How this is calculated

No lookup tables and no invented constants. Here is the arithmetic, so you can check it against your own numbers.

Apple Silicon shares one memory pool between CPU, GPU and the operating system, so the usable fraction is lower than on a discrete card — macOS caps what a single GPU process may allocate, and the desktop needs its share.

Bandwidth, not capacity, is what separates the chip tiers. A Max part has roughly three to four times the bandwidth of a base part at the same capacity, and that ratio is close to the ratio of tokens per second you see.

These are well-founded engineering estimates, not benchmark results. Your quantisation, runtime and context length all move the real number, and usable memory is an assumption rather than a specification. See the full methodology for every assumption behind these figures.

Use cases

Who this is for

The three situations that bring people to this calculation.

Use case 01

Configuring a Mac

Choose a memory tier from the models you want to run.

Use case 02

Comparing chips

See what a Pro or Max buys in speed.

Use case 03

Existing Mac

Find what your machine already handles.

Walkthrough

How to use this calculator

Four steps, no account, nothing leaves your browser.

  1. Set model size

    Start at the top of the panel. Every figure recalculates as you change it — there is no submit button, because watching the number move is the point.

  2. Adjust the rest to match your setup

    2 further settings: context length, quantisation. Defaults are the common case, so change only what differs for you.

  3. Read the headline, then the table

    The large figure answers the question. The table underneath shows how the answer changes across nearby settings, which is usually where the decision actually gets made.

  4. Check it against the method

    The arithmetic is written out above. If a number looks wrong for your hardware, the assumptions are the first place to look — usable memory and quantisation are the two that vary most.

Questions

How much unified memory do I need on a Mac: common questions

The questions people ask about this, answered without hedging.

Is 16GB enough to run local LLMs on a Mac?

For 7–8B models at 4-bit, yes — about 5 GB of weights plus cache leaves room to work in. For 32B you want 32GB or more, and 70B needs 64GB before it is comfortable.

How much unified memory do I need for local AI?

16GB runs 7–8B models comfortably at 4-bit. 32GB opens 32B. 64GB and above is where 70B becomes practical. Memory is not upgradeable later, so buy ahead of what you need.

Why can I not use all my unified memory?

macOS reserves a share for the system and caps what a single GPU process may allocate. The usable fraction is meaningfully lower than the number on the spec sheet.

Is a Mac good for running LLMs?

Very good on capacity per dollar and excellent on power draw. A Max chip has bandwidth competitive with mid-range discrete cards; a base chip is much slower.

Base, Pro or Max?

For capacity alone, memory size is what matters. For speed, the chip tier is what matters — a Max has roughly three to four times a base chip’s bandwidth, and tokens per second follow that ratio closely.

The rest of the set

All 50 run on the same arithmetic, so answers across them agree.

Coverage

Searches this page answers

Different ways of asking the same question, all resolved above.

mac unified memory llmhow much ram for llm on macm4 max llm performancemacbook local ai memoryapple silicon llm speedis 16gb enough for local llm macmac vs gpu for llmhow much unified memory do i need on a macmac unified memory calculatormac unified memory calculator onlinefree mac unified memory calculator
Next step

Now find the models that fit

Sizing is only half the problem. Model Radar takes your hardware and shows which models actually run on it, ranked by what they are good at — the same arithmetic as this page, applied to every model worth running.