Guide · Local AI
How Much Memory Does an AI Model Need?
October 2026
Before asking whether a computer is fast enough to run an AI model, ask whether it has enough memory. If the model doesn't fit, nothing else matters. The good news is that a rough answer takes one multiplication. This guide shows the rule of thumb, what it leaves out, and how it applies to a very large model that was announced this week.
Step 1: The weights
A model is mostly a long list of numbers called parameters. The memory it needs is the number of parameters times the bytes each one takes. At full 16-bit precision that is 2 bytes, so a 70-billion-parameter model takes about 140GB (LLM Stats).
Quantization stores each number with fewer bits to shrink the model:
- 16-bit: about 2 bytes per parameter.
- 8-bit: about 1 byte per parameter.
- 4-bit: half a byte in theory, but real 4-bit files come out nearer 0.6 bytes per parameter (Spheron).
Going much lower saves more memory but costs quality. Very low-bit versions are often too lossy to trust (DEV Community).
Step 2: Add the context
While a model works, it keeps a record of the conversation called the KV cache. It grows with the length of the context, not with the size of the model. One guide puts a Llama 3.1 8B model at about 4GB of cache at 32K tokens of context and about 16GB at 128K, and notes that some techniques can shrink that by 4 to 8 times (LLM Stats). Add one to two gigabytes for the software around it (Spheron).
Step 3: Remember "mixture of experts"
Many large models are mixtures of experts. They have a total parameter count and a smaller active count, the part used for each token. All of the parameters still have to be stored, so total size sets how much memory you need. The active count affects how fast the model can respond (LLM Stats). As an example, Meta's Llama 4 Scout has 109 billion parameters in total and 17 billion active, and the same guides put it at roughly 60 to 70GB at 4-bit.
A worked example: Reflection's Beam
Reflection AI announced Beam on October 5. It's a mixture-of-experts model with 501 billion total parameters and 23 billion active, with a context of up to 1 million tokens. The company says it will release the weights under an Apache 2.0 licence later this month. Until then, access is by early-access sign-up, and its benchmark results are the company's own (Reflection).
Applying the rule, these are rough estimates for the weights alone, before any context:
- 16-bit: about 1TB.
- 8-bit: about 500GB.
- 4-bit: about 250 to 300GB.
Set against the machines in our overview of local AI hardware: an RTX Spark system tops out at 128GB, so it can't hold Beam at any of these sizes. A Mac Studio M5 Ultra configuration with 512GB, due in late October, could in principle hold a 4-bit version with room left for context. Two 128GB machines would add up to 256GB, which looks too tight at 4-bit once context and overhead are added. These are calculations, not tests. Nobody can run Beam locally until the weights are out.
What to watch
- The actual file sizes of quantized Beam releases once the weights are published.
- How much memory the 1 million token context adds in practice.
- Speed: the active parameter count helps, but memory bandwidth also decides how quickly a model answers.
- How much quality is lost at 4-bit and below on real tasks.
What we don't know yet
Beam's weights aren't public, so the sizes above come from a general rule, not from a measurement. The rule is a rough guide: real memory use depends on the file format, the runtime and the context you ask for. We haven't run Beam, and Reflection's performance claims haven't been independently verified.
More from Offgrid Studio
Local AI Hardware in Autumn 2026: What You Can Actually Run at Home
Why Cloud AI Bills Are So Hard to Budget