Choose what you want the machine to do, identify a realistic model class, and size memory for that model plus context and runtime overhead. Spend on accelerator memory first when fast inference matters; keep enough system RAM and SSD space to avoid turning the rest of the build into the bottleneck.
1. Define the job before the parts
“Run AI locally” is too broad to guide a purchase. A private chat assistant, a coding model with a large repository in context, image generation, and a shared office server stress hardware differently. Write down one primary job and one secondary job.
| Workload | What usually matters most | Easy mistake |
|---|---|---|
| Chat and summarization | Model quality, memory capacity, responsive token generation | Buying for a model far larger than you will use |
| Coding | Enough memory for the model and useful context | Ignoring context/KV-cache growth |
| Image or video | Supported GPU backend, VRAM, workflow-specific benchmarks | Applying LLM sizing rules to diffusion/video workloads |
| Multi-user server | Concurrency, reliability, cooling, power, network access controls | Sizing only for one interactive user |
2. Set a memory target, not just a model name
Model weights must fit somewhere. Hugging Face’s optimization guide gives a useful baseline: a model with X billion parameters needs roughly 2X GB for FP16/BF16 weights alone. Quantization can reduce that substantially, but the model file is not the whole memory budget. Context, KV cache, runtime buffers, and other software need room too.
That is why “an 8B model needs exactly 8 GB” is not a dependable buying rule. Quantization format, context length, architecture, runtime, and offloading change the result. Use the VRAM guide for the sizing method, then compare candidate models using My rig.
3. Choose a platform you can actually support
NVIDIA generally has broad CUDA ecosystem support, but exact runtime and driver requirements still matter. AMD support varies more by GPU, operating system, and ROCm version, so check the current compatibility matrix before buying. Apple silicon uses unified memory and Metal: it can make large capacities available in one pool, but capacity alone does not guarantee the fastest output.
If you already own a computer, test it first. llama.cpp supports CPU execution, multiple accelerator backends, and CPU/GPU hybrid inference. A slower successful test teaches you more about the model you want than a speculative parts list.
4. Balance the rest of the machine
- System RAM: keep enough for the operating system, runtime, model loading, and any CPU-offloaded layers. More RAM expands what can load, but it does not turn system memory into fast VRAM.
- Storage: model files add up quickly. Use an SSD, leave working space for alternate quantizations, and avoid downloading every variant.
- CPU: important for CPU inference, tokenization, data preparation, and offloaded work. For a fully GPU-resident interactive model, it is rarely the first upgrade.
- Power and cooling: size the PSU for the complete system and the GPU manufacturer’s guidance. Confirm case clearance, connectors, airflow, and sustained noise—not only peak wattage.
5. Buy in the order that reduces uncertainty
- Run one smaller model on hardware you already have.
- Pick the largest model class you expect to use regularly, not once.
- Choose a quantization and context target.
- Confirm runtime support for the exact GPU and operating system.
- Price the accelerator, RAM, SSD, PSU, cooling, and case as a complete system.
If a less expensive build runs your regular model comfortably, a larger model must provide a measurable benefit to justify the extra hardware, power, heat, and maintenance.
6. Verify the finished rig
Load a representative model, use the context length you expect in real work, and watch where it is placed. Ollama’s ollama ps reports whether a loaded model is on the GPU, CPU, or split between both. Also test first-token delay, sustained generation, temperatures, noise, and stability. “It loaded” and “I enjoy using it” are different results.
Next, follow the first-model guide to choose a runtime and run a small, verifiable test.
Sources and review notes
- Hugging Face: optimizing LLMs for speed and memory — baseline weight-memory math.
- llama.cpp README — backends, quantization, and hybrid inference.
- Ollama GPU documentation and AMD ROCm compatibility — current platform support.
- Puget Systems hardware primer — testing-based discussion of memory headroom and context.
This is educational guidance, not a sponsored parts list. Hardware and software support changes; verify exact components before purchase.