GPU VRAM guide for local AI
8–12GB
Small quantized chat models and light image gen. Cards: 5060, 5070, 3060 12GB, 4070, 3080 10/12GB.
16GB
Mid-size models. Cards: 5070 Ti, 5080, 5060 Ti 16GB, 4070 Ti Super, 4080, RX 9070 XT, 7800 XT.
20–24GB
Common consumer target for local AI. Cards: 3090, 4090, 7900 XTX, 7900 XT (20GB).
32GB
Consumer step above 24GB without a workstation SKU. Card: RTX 5090.
48GB+
Large models, multi-modal, fewer compromises. Cards: RTX A6000, RTX 6000 Ada. Thin volume, high ticket, eBay-heavy.
Quantization note
Q4/Q5 quantization shrinks memory needs. A 24GB card can run models that would not fit at full precision. Benchmarks change with model releases — re-check VRAM claims when a new open model ships.
FAQ
How much VRAM do I need for local AI?
More VRAM means larger models and longer context. For most hobby local LLM users, 24GB is the target. 16GB is workable. 12GB is entry-level. 48GB is workstation territory.