HOW IT WORKS
A transparent estimate, not a promise.
Weights are estimated from parameter count × bytes per quantized parameter. KV cache uses 2 × layers × hidden size × context × concurrent sequences × 2 bytes, assuming FP16 keys and values.
Runtime overhead is estimated at 12% of weight memory with a 1.2 GB floor, then a 20% comfort margin is added. Real use varies by architecture, quantization format, backend, offloading and operating system.