Methodology

How results are produced.

Compatibility and speed are different questions. A model fitting in memory does not prove that it will be fast.

Data classifications

MEASURED

An observed benchmark with identifiable hardware, model variant, runtime, and environment.

ESTIMATED

A calculated range based on model size, usable memory, bandwidth, context, and conservative efficiency assumptions.

SPECIFICATION

A value taken from official hardware information, a model card, or runtime documentation.

UNKNOWN

There is not enough reliable information to make a useful claim.

Memory fit

Text-model estimates include quantized model weights and a working-memory allowance. Advanced calculations can also include KV cache, context length, runtime overhead, and partial CPU offloading. Dedicated VRAM and Apple unified memory are treated differently.

Performance estimates

When no benchmark is available, generation speed may be shown as a broad range derived from memory bandwidth and model size. These figures are planning estimates, not benchmark observations.

Limitations

Drivers, runtime versions, prompt length, thermals, power limits, background applications, model architecture, and quantization can materially change results. Business server presets are starting points, not production capacity guarantees.