Computing

Why 128GB of shared memory matters more than your GPU for local AI

(today) · 3 min read · By Future Technology · Edited by Nath Connell

Key takeaways

  • Unified memory means the CPU and GPU draw from one shared pool, so a model is not limited by a separate graphics card's VRAM
  • Rough rule: a 4-bit model needs about half a gigabyte per billion parameters, plus room for context
  • Very roughly, 8GB suits small models, 32GB suits mid-size ones, and 128GB opens up the biggest open models

A 70 billion parameter model stored at 4 bits per weight needs roughly 35GB just for its weights. Most graphics cards cannot hold that. A machine with 128GB of unified memory can, with room to spare. That one number decides almost everything about what you can run at home.

What unified memory is

On a typical desktop PC, the CPU has its own system memory and the graphics card has separate VRAM. A model has to fit in the VRAM, or it spills into slower system memory and speed collapses. Unified memory removes the split. The CPU and GPU read from one pool, so the limit is the total, not the size of the graphics card.

That design is why Apple's Macs became popular for local models, and it is the idea behind Nvidia's new RTX Spark platform, which offers up to 128GB shared between its Grace CPU and Blackwell GPU. We cover the launch in our RTX Spark news piece.

How much memory a model needs

Very roughly, a model quantised to 4 bits uses about half a gigabyte per billion parameters, and 8 bit uses about one gigabyte. Add a few gigabytes for the context window and the operating system. These are rules of thumb, not guarantees, since exact use varies by software and settings.

At 8GB you are in small model territory, around 7 billion parameters at 4 bit, with little room left over. At 32GB, models in the 30 billion range fit comfortably at 4 bit, which covers most of what people use day to day. At 128GB, a 70 billion model fits at 8 bit with space for long context, and larger open models become possible at lower precision.

What to buy, and what to skip

If you want an affordable way in, the Apple Mac mini M4 with 24GB of unified memory is a known option, and you can check current pricing on Amazon. Our guide to running models locally with vLLM, Ollama, LocalAI and Exo covers the software side.

Unified memory is not magic. Shared memory is usually slower than the dedicated memory on a top graphics card, so a big model will answer more slowly than the same one on fast VRAM. The advantage is that it runs at all. Memory capacity decides what you can try; bandwidth decides how patient you need to be.

Watch the RTX Spark prices when they land. A 128GB machine is only interesting if the price lets ordinary buyers reach it.

Some links in this article are affiliate links. We may earn a small commission at no extra cost to you.

More from Future Technology