There is no single VRAM requirement for “local AI.” For an LLM, estimate weight memory from parameter count and precision, add context/KV-cache and runtime overhead, then choose whether you need full GPU placement or can accept CPU offload. Check the exact downloadable file before buying.
1. Estimate the weights
For unquantized weights, a useful approximation is:
- FP32: about 4 bytes per parameter.
- FP16/BF16: about 2 bytes per parameter.
- 8-bit: roughly 1 byte per parameter, plus format/runtime details.
- 4-bit: roughly half a byte per parameter, plus format/runtime details.
So an 8-billion-parameter model is roughly 16 GB at FP16/BF16 before overhead. Quantized files can be much smaller, but “4-bit” is a family of formats rather than one guaranteed file size. The exact model file is the better planning number.
2. Add what the weight file does not show
Inference also uses memory for the context/KV cache, runtime buffers, and sometimes vision or other model components. Longer context generally increases memory use. Batch size and concurrent users can increase it further.
Puget Systems found in its tested llama.cpp configurations that modest headroom above the model’s weight memory could be sufficient for usable context. Treat that as a testing result, not a universal percentage: architecture, runtime, cache precision, context, and workload can change the requirement.
| If your target is… | Plan around… | Then verify… |
|---|---|---|
| A small everyday assistant | A current small/mid-size quantized model | Exact file size plus your normal context |
| A coding assistant | The model plus longer working context | First-token delay and memory at realistic repository context |
| A larger model via offload | Combined VRAM and system RAM capacity | Whether the resulting speed is acceptable |
| Multiple simultaneous users | Model plus per-session cache and concurrency | Throughput under expected load |
3. Decide whether offload is acceptable
If all model layers fit in fast accelerator memory, generation is usually more responsive. llama.cpp and other runtimes can split work between GPU and CPU memory, allowing a larger model to run—but system RAM has different bandwidth and latency. Offload expands capacity; it does not promise the same performance as a fully GPU-resident model.
Apple unified memory does not divide capacity into conventional system RAM and discrete VRAM. That can make larger models practical in one pool, but the usable amount is still shared with the operating system and applications.
4. Choose quantization deliberately
Quantization trades some numerical precision for lower memory use and often faster local inference. A recent Hugging Face llama.cpp example shows one 4B model shrinking from 8.42 GB in BF16 to 2.74 GB at Q4_K_M, 3.14 GB at Q5_K_M, and 3.53 GB at Q6_K. Those are model-specific values, but they illustrate why the exact variant matters.
A practical approach is to begin with a well-supported Q4 variant, test quality on your actual tasks, and move to Q5/Q6 only if the improvement is useful and memory allows. Do not assume the largest file is automatically the best everyday choice.
5. Use this pre-purchase check
- Open the model’s official or publisher-linked download page.
- Choose the exact quantization/file you expect to run.
- Add room for your intended context and runtime.
- Confirm backend support for the exact GPU, OS, and driver.
- Use My rig for an estimate, then verify with the runtime’s own memory report.
Once you have a target, read how to balance the rest of the rig or move directly to running a first model.
Sources and review notes
- Hugging Face LLM optimization guide — parameter precision and baseline memory math.
- Hugging Face bitsandbytes guide — 8-bit and 4-bit quantization concepts.
- Hugging Face llama.cpp quantization example — model-specific Q4/Q5/Q6 file sizes.
- Puget Systems hardware primer — measured context/headroom discussion.
Memory figures are planning estimates. Architecture, runtime, context, batch size, and model format can change actual use.