8GB vs 12GB vs 16GB VRAM: What Can You Actually Run?
Learn what local AI models, games, and apps you can realistically run on 8GB, 12GB, and 16GB VRAM GPUs, with benchmarks and setup commands.
The VRAM Wall: Why You Care
You just bought a GPU, or you are deciding which one to buy, and the specs say 8GB, 12GB, or 16GB of VRAM. The question is not whether the card is fast, but whether it can hold the model, the game, or the render in memory at all. If it cannot, you get an out-of-memory error, a crash, or a slow swap to system RAM that makes the card feel ten times slower.
This article gives you a practical map of what each VRAM tier can run today: local LLMs, image generation, video editing, and games. You get exact commands and config files to test your own card, so you can measure, not guess.
Before You Start: Know Your GPU
VRAM is not the only factor. Memory bandwidth, compute cores, and driver support matter. But for capacity decisions, VRAM is the hard limit. A model that needs 12GB simply will not fit in 8GB without quantization or offloading.
Find your GPU and its VRAM with a quick command.
nvidia-smi --query-gpu=name,memory.total --format=csv- If you have an AMD card, use rocminfo or clinfo to see VRAM.
- On Windows, Task Manager shows GPU memory in the Performance tab.
- For laptops, remember that VRAM is dedicated; shared memory is slower and not equivalent.
Local LLMs: The Model Size Reality
Large language models are the biggest driver of VRAM demand. A model's file size is a rough lower bound for the VRAM needed to load it. A 7B parameter model in 16-bit floats takes about 14GB, which fits only in 16GB cards. But you can quantize to 8-bit or 4-bit to cut that in half or more.
Here is a rough table of model sizes and VRAM requirements for inference with llama.cpp, the most common local runtime.
- 7B model, 16-bit: ~14GB, needs 16GB VRAM.
- 7B model, 8-bit: ~7GB, fits 8GB with tight margins.
- 7B model, 4-bit: ~4GB, runs on 8GB easily.
- 13B model, 8-bit: ~13GB, needs 16GB.
- 13B model, 4-bit: ~7GB, fits 8GB if you close everything else.
- 70B model, 4-bit: ~35GB, needs 48GB or offloading to RAM.
Step 1: Measure Your LLM Memory Footprint
While it runs, open another terminal and watch nvidia-smi. Note the memory used before and after loading. That is your real footprint.
watch -n 1 nvidia-smiStep 2: Quantize a Model to Fit Your Card
Now quantize to 4-bit to cut the memory requirement roughly in half.
./quantize ./models/llama-2-7b.Q8_0.gguf ./models/llama-2-7b.Q4_K_M.gguf Q4_K_M- Q4_K_M is a good balance of quality and size.
- You can also use Q5_K_M for slightly better quality at a small VRAM cost.
- After quantization, re-run the Python script to measure the new footprint.
Image Generation: Stable Diffusion and SDXL
For SDXL, you need to use the --lowvram flag on 8GB cards, which offloads to RAM and slows generation. On 12GB you can run SDXL without flags, but batch size should stay at 1. On 16GB you can run SDXL with a batch of 2 or more.
A quick test: generate a 1024x1024 image and watch VRAM with nvidia-smi.
watch -n 1 nvidia-smiVideo Editing and 3D Rendering
Video editing software like DaVinci Resolve uses VRAM for effects and color grading. 8GB is the minimum for 1080p timelines, but 4K projects will struggle. 12GB handles 4K with moderate effects. 16GB is comfortable for 4K with Fusion effects or 6K RAW.
For 3D rendering with Blender Cycles, VRAM holds the scene data and textures. An 8GB card can render small scenes, but a complex scene with high-res textures will spill to RAM and slow down drastically.
Test your card with a simple Blender benchmark scene: download the BMW benchmark scene and render it, watching VRAM usage.
blender -b bmw27.blend -o /tmp/render_ -F PNG -x 1 -t 0- The BMW scene is about 200MB and fits in 8GB.
- For a stress test, use the Classroom scene, which is larger and may exceed 8GB.
- If you see 'CUDA out of memory', reduce texture size or use the --cycles-device CPU fallback.
Gaming: What Fits in Each Tier
Games are less about raw model size and more about texture resolution and detail. At 1080p, 8GB is enough for most titles at high settings, but some modern games with high-res texture packs can exceed that. 12GB is a sweet spot for 1440p, and 16GB is future-proof for 4K and ray tracing.
A quick rule of thumb: check the recommended VRAM on the game's page. If it says 8GB, you are fine. If it says 12GB, you need at least 12GB. Here is a sample of common games and their VRAM usage at 1080p ultra (as of 2025).
- Cyberpunk 2077: 7-8GB at 1080p, 10GB at 1440p, 12GB+ at 4K.
- Hogwarts Legacy: 8GB at 1080p, 12GB at 1440p.
- Forza Horizon 5: 6GB at 1080p, 8GB at 1440p.
- Starfield: 8GB at 1080p, 10GB at 1440p.
- These numbers are from public benchmarks, but your mileage may vary with drivers and settings.
What I Would Do: Recommended Setup by VRAM
If you are buying a card today, here is my recommendation based on what you want to run. This is not a benchmark, but a practical guide from community experience and my own testing.
For 8GB: Use 4-bit quantized 7B models, SD 1.5 for images, and stick to 1080p gaming. Avoid SDXL unless you use --lowvram and are patient.
For 12GB: You can run 13B models at 4-bit, SDXL at 1024x1024 without lowvram, and 1440p gaming. This is the sweet spot for most hobbyists.
For 16GB: You can run 13B models at 8-bit, 7B models at 16-bit, SDXL with batch size 2, and 4K gaming. This is the entry point for serious local AI.
# Example: run a 13B 4-bit model on 12GB
./main -m models/llama-2-13b-chat.Q4_K_M.gguf -p "Explain VRAM" -n 128Troubleshooting Out-of-Memory Errors
You will hit OOM errors. Here is how to handle them.
First, check what is using VRAM. Use nvidia-smi to see processes. Kill anything you do not need.
nvidia-smi
# List processes using GPU
fuser -v /dev/nvidia*- Close browser tabs, especially Chrome, which can use hundreds of MB of VRAM.
- Lower the batch size in your AI script to 1.
- Use quantization or offloading flags like --lowvram in SD Web UI.
- For llama.cpp, reduce n_ctx (context length) to save memory.
- If all else fails, use CPU offloading with --n-gpu-layers to move some layers to system RAM.
FAQ
Answers to the questions that come up most often on this topic.
- Q: Can I run a 70B model on 16GB? A: Not without heavy quantization and offloading. A 4-bit 70B model is about 35GB, so it will not fit in 16GB. You would need to offload most layers to RAM, which is extremely slow.
- Q: Is VRAM more important than GPU speed? A: For local AI, yes. If the model does not fit, the card is useless. Speed matters only after it fits.
- Q: Can I upgrade VRAM on my GPU? A: No, VRAM is soldered on. You must buy a new card.
- Q: Does shared system RAM help? A: Some tools can offload to RAM, but it is 10-100x slower than VRAM. Use it only as a fallback.
Next Step: Run Your Own Test
Stop reading and run a test. Download a 7B quantized model and load it with llama.cpp. Measure your VRAM usage with nvidia-smi. That number is your real capacity. Then you will know exactly what you can run.
The command below is a quick smoke test for any card.
./main -m models/llama-2-7b-chat.Q4_K_M.gguf -p "Hello" -n 20Key takeaways
- Apply one concrete change from this post before collecting more reading.
- Prefer browser-side tools when the work involves secrets, tokens, or PII.
- Document the why next to the how so the next reviewer inherits context.
FAQ
- Who is this guide on gpu for?
- Working developers who need a practical take on 8gb vs 12gb vs 16gb vram: what can you actually run? — not a marketing overview. Skim the sections, apply one tip, then come back when you hit an edge case.
- Do I need an account to use the related tools?
- No. code.live tools run in your browser with no signup. Nothing you paste is uploaded to a server for the client-side utilities linked from this post.
- How often is this article updated?
- This post was published October 5, 2026. Fundamentals stay stable; check linked tool pages and official docs when version-specific behavior matters.