Why VRAM Matters More Than GPU Speed for Local AI
Learn why VRAM, not raw GPU speed, determines which local AI models you can run, and how to measure it with real commands.
The problem: your GPU is fast, but the model won't load
You just bought a shiny new GPU with impressive teraflops, and you try to run a 13B parameter model locally. The terminal spits out an error: CUDA out of memory. You have a fast card, but it only has 8GB of VRAM. The model needs at least 12GB just for weights.
This is the most common frustration in local AI. People shop for GPUs by clock speed or core count, but for running large language models, VRAM capacity is the hard constraint. Without enough VRAM, the model either won't load or runs at a crawl because it spills to system RAM.
This article shows you how to calculate VRAM requirements, measure your actual usage, and choose a GPU that fits the models you actually want to run.
- You will learn how to estimate VRAM needs for any model.
- You will get shell commands to check your GPU and measure memory usage.
- You will see a side-by-side comparison of speed vs. capacity in practice.
Why VRAM is the bottleneck
When you run a neural network, the entire model must live in memory that the GPU can access directly. That is VRAM. The GPU's compute units can only process data that is resident in VRAM; transferring data to and from system RAM over PCIe is orders of magnitude slower.
For inference, the model weights are read once per token generated. If the weights do not fit in VRAM, the system must page them in and out, and you get maybe 1-2 tokens per second instead of 30+. For training or fine-tuning, you also need room for gradients and optimizer states, which multiplies the requirement.
A fast GPU with insufficient VRAM is like a sports car with a tiny fuel tank: it can go fast, but you have to stop every few miles. A slower GPU with ample VRAM may be slower per token, but it can actually complete the job without thrashing.
Step 1: Check your current GPU and VRAM
Before you buy anything, know what you have. On Linux, the nvidia-smi command shows your GPU model, driver version, and current memory usage. On Windows, you can run the same command if you have the NVIDIA driver installed, or use Task Manager's Performance tab.
Run this to see your GPU and its total VRAM:
nvidia-smi --query-gpu=name,memory.total,memory.used --format=csv- The memory.total column is your VRAM capacity.
- If you see multiple GPUs, run with --id=0, --id=1, etc., to inspect each.
- On macOS, use system_profiler SPDisplaysDataType to see unified memory.
Step 2: Estimate VRAM for any model
You can estimate the VRAM needed for inference with a simple formula. For a model with P parameters, the weights take up roughly 2 bytes per parameter in FP16, or 4 bytes in FP32. Add about 20% overhead for activations and CUDA context.
For a 7B model in FP16: 7e9 * 2 bytes = 14 GB. Add overhead, and you need about 16 GB. For a 13B model, you need about 26 GB. This is why 8GB and 12GB cards struggle with anything above 7B.
You can also use quantization to reduce the footprint. A 4-bit quantized 7B model needs about 3.5 GB, which fits on many consumer cards. Here is a quick script to compute the numbers:
def vram_needed(params_billions, bits=16, overhead=0.2):
bytes_per_param = bits / 8
weights_gb = params_billions * 1e9 * bytes_per_param / 1e9
total_gb = weights_gb * (1 + overhead)
return round(total_gb, 1)
for p in [1, 3, 7, 13, 70]:
for bits in [16, 8, 4]:
print(f'{p}B, {bits}-bit: {vram_needed(p, bits)} GB')- Run this in any Python environment to see the numbers.
- The overhead factor varies; 0.2 is a safe starting point.
- For training, multiply by 3-4 for gradients and optimizer states.
Step 3: Measure real VRAM usage during inference
Estimation is fine, but measuring is better. When you run a model, you can watch VRAM usage in real time with nvidia-smi. Use the watch command to refresh every second:
Start your model server in one terminal, then run this in another:
watch -n 1 nvidia-smi --query-gpu=memory.used --format=csv- Keep the model loaded and generate a few tokens to see peak usage.
- Note the maximum memory.used value; that is your real requirement.
- If the process gets killed, you are out of VRAM.
Speed vs. capacity: a practical comparison
To demonstrate why capacity wins, here is a side-by-side scenario. GPU A is a fast card with 8GB VRAM; GPU B is a slower card with 24GB VRAM. Both try to run a 13B model in FP16, which needs about 26GB.
GPU A cannot fit the model. It falls back to CPU offloading, which is painfully slow. GPU B fits the model comfortably and generates tokens at full speed. The raw compute speed of GPU A is irrelevant because it never gets to use it.
You can simulate this with llama.cpp. On GPU A, you would run with partial offload:
./main -m model-13b.gguf -ngl 10 -p "Hello"./main -m model-13b.gguf -ngl 99 -p "Hello"- The -ngl flag controls how many layers go to the GPU.
- On GPU A, the CPU does most of the work, and you get single-digit tokens per second.
- On GPU B, you get 20-30 tokens per second. Same model, same GPU architecture, different VRAM.
Checklist: choosing a GPU for local AI
Use this checklist when evaluating a GPU for local AI. It is not about the highest clock speed or the most CUDA cores; it is about fitting the models you plan to run.
- List the largest model you want to run and its precision (FP16, 8-bit, 4-bit).
- Calculate the VRAM needed using the estimator above.
- Add 20-30% headroom for context length and overhead.
- Check the GPU's memory bandwidth; higher bandwidth helps token generation speed.
- Verify the GPU has enough PCIe lanes if you plan to run multiple cards.
- Read real-world reports from other users running the same model on the same card.
What I would do: recommended setups
Based on the math, here are my recommendations for common budgets. These are not absolute truths, but they reflect what actually works in practice today.
For 7B models, an RTX 4060 Ti 16GB is a sweet spot. It has enough VRAM for FP16 and room for longer contexts. For 13B models, you want at least 24GB, so an RTX 3090 or 4090 is the way to go. For 70B models, you are looking at 48GB or more, which means professional cards or multi-GPU setups.
If you are on a budget, consider used RTX 3090s. They offer 24GB of VRAM at a fraction of the cost of a new 4090, and the speed difference is acceptable for most inference tasks.
Here is a copy-paste command to check if a model fits on your GPU before you download it, using a simple Python script:
python -c "import torch; print(torch.cuda.get_device_properties(0).total_memory / 1e9, 'GB')"- Run this after installing PyTorch with CUDA support.
- If it prints an error, your PyTorch is CPU-only.
- Compare the output to your model's VRAM requirement.
Troubleshooting: out of memory errors
You will eventually hit CUDA out of memory. Here is how to handle it.
First, check if other processes are using VRAM. Use nvidia-smi to see what is running. Kill any leftover processes. Then, try reducing the context length or switching to a quantized version of the model.
If you are using Hugging Face Transformers, you can enable CPU offloading, but expect a massive slowdown. A better fix is to use llama.cpp with the right -ngl value.
nvidia-smi --query-compute-apps=pid,used_memory --format=csv- Kill unused processes with kill -9 PID.
- Use a 4-bit quantized model to halve VRAM usage.
- Reduce --ctx-size in llama.cpp to lower memory for attention.
- If all else fails, upgrade your GPU or rent a cloud instance.
FAQ
Answers to the questions that come up most often on this topic.
- Q: Can I run a model bigger than my VRAM? A: Yes, with CPU offloading, but it will be very slow. It is usable for batch jobs, not interactive chat.
- Q: Does more VRAM always mean faster? A: No, but it prevents slowdowns from offloading. A card with more VRAM and lower bandwidth might be slower per token than a smaller card that fits the model.
- Q: Is 8GB enough for local AI? A: Yes, for 1B-3B models or heavily quantized 7B models. You will be limited in model size and context length.
- Q: What about AMD or Intel GPUs? A: They are improving, but NVIDIA still has the best software support for AI frameworks.
Next action
Now that you know the math, run the estimator script for the model you want to run. If your current GPU cannot fit it, use that number to guide your next purchase. If you are just starting, pick a 7B model and a GPU with at least 16GB VRAM.
Your next command is to check your VRAM:
nvidia-smi --query-gpu=name,memory.total --format=csvKey takeaways
- Apply one concrete change from this post before collecting more reading.
- Prefer browser-side tools when the work involves secrets, tokens, or PII.
- Document the why next to the how so the next reviewer inherits context.
FAQ
- Who is this guide on gpu for?
- Working developers who need a practical take on why vram matters more than gpu speed for local ai — not a marketing overview. Skim the sections, apply one tip, then come back when you hit an edge case.
- Do I need an account to use the related tools?
- No. code.live tools run in your browser with no signup. Nothing you paste is uploaded to a server for the client-side utilities linked from this post.
- How often is this article updated?
- This post was published October 6, 2026. Fundamentals stay stable; check linked tool pages and official docs when version-specific behavior matters.