The short answer

Choose LM Studio for a visual desktop workflow, Ollama for a simple command-line service and local API, or llama.cpp for direct control. Download a small model first, disconnect cloud features if privacy requires it, and confirm model placement with the runtime’s status tools.

1. Pick the runtime that removes friction

RuntimeBest first fitWhy choose it
LM StudioYou want a desktop interfaceModel search, downloads, chat, settings, and a local server in one visual app
OllamaYou want quick commands or an app backendSimple model management and a local API that other tools can call
llama.cppYou want direct control or broad hardware pathsGGUF focus, detailed options, CPU/GPU offload, and multiple accelerator backends

No runtime is universally best. Your operating system, accelerator, model format, and preference for a GUI or terminal matter more than popularity.

2. Start smaller than your maximum

Your first run is a systems check, not a contest. Choose a current instruction-tuned model in a modest size and a well-supported quantization. It should leave enough memory headroom that you can tell whether drivers, runtime support, and model loading work before testing the limit.

Use the model catalog for a curated starting point and My rig for a hardware-specific estimate. Check the model card and license before use, especially for commercial work.

3. Run it with the shortest path

Ollama

Install Ollama from its official site, choose a model from the official library, then run the command shown on that model’s page. A typical command has this form:

ollama run <model-name>

Ollama binds to 127.0.0.1:11434 by default. Do not expose the service to a network until you understand authentication, firewall, and proxy requirements.

LM Studio

Install the current app, search for a supported model, select a quantized file that fits your memory, load it, and open a chat. Use the Developer/Local Server tools only when another local application needs an API.

llama.cpp

Build or install a current release, download a compatible GGUF file from a trusted publisher, then invoke the CLI with that file. Exact binary names and flags can differ by release, so follow the current README rather than an old copied command.

4. Verify where it is running

With Ollama, run:

ollama ps

The processor column can show full GPU placement, full CPU placement, or a CPU/GPU split. For other runtimes, inspect the application’s load information, logs, and your operating system’s GPU memory/compute monitor. A GPU name appearing in the interface is not proof that every layer is accelerated.

  • Ask a repeatable prompt and note time to first token.
  • Watch accelerator memory, system RAM, and utilization.
  • Increase context toward a realistic workload and watch memory again.
  • Check temperatures and stability over more than one response.

5. Understand what “local” means

Local inference means prompts and generation can remain on your device, but the application may still contact the internet for model downloads, updates, or optional cloud features. Ollama documents a local-only mode using OLLAMA_NO_CLOUD=1 or its application settings. Review the current privacy and network settings of whichever tool you choose.

Keep the local API local

The default loopback address is safer than listening on every network interface. If you later expose a local AI service, add access controls and do not assume the model server provides them automatically.

6. Improve one variable at a time

If the first model works, change only one of model size, quantization, context, or offload settings at a time. That makes failures explainable. If it does not fit, read the VRAM guide; if it runs but is too slow, compare full GPU placement with offload and decide whether a hardware upgrade serves your real workload.

Sources and review notes

Install commands and compatibility can change. Use the linked official documentation for the latest release-specific steps.