VRAM math for any open LLM, from 7B to 405B
Three numbers and one ratio tell you whether a model fits on a given GPU. Here's the formula, with worked examples for Llama-3.3, Qwen 2.5, and DeepSeek across every common quant level.
Self-host
Sizing, GPU choice, networking, ops.
Three numbers and one ratio tell you whether a model fits on a given GPU. Here's the formula, with worked examples for Llama-3.3, Qwen 2.5, and DeepSeek across every common quant level.
There are eight open models worth considering in 2026. Picking the right one for your workload is mostly a function of three constraints: hardware, latency, and domain. Here's the decision matrix we use on engagements.