How Much VRAM Do You Actually Need to Run an AI Model?
Learn to calculate VRAM needs for local LLMs, estimate with formulas, and test real models on your GPU.
Before you start
You need Python 3.10 or newer, pip, and an NVIDIA GPU with at least 4 GB of VRAM to follow the hands-on parts. If you do not have a GPU, you can still run the estimation script, but the measurement sections require one.
The formula that matters
The minimum VRAM to load a model is roughly the number of parameters times bytes per parameter. For 16-bit (FP16) weights, that is 2 bytes per parameter. For 8-bit, it is 1 byte. For 4-bit, it is 0.5 bytes. Add about 20 percent for the inference engine, KV cache, and activations.
So a 7B model at FP16 needs about 14 GB just for weights. That is why you see 8 GB cards struggling with 7B models unless you use quantization.
python -c "
params = 7_000_000_000
bytes_per_param = 2 # FP16
weights_gb = params * bytes_per_param / 1e9
print(f'Weights: {weights_gb:.1f} GB')
print(f'With 20% overhead: {weights_gb * 1.2:.1f} GB')
"- FP32: 4 bytes per parameter
- FP16/BF16: 2 bytes per parameter
- INT8: 1 byte per parameter
- INT4: 0.5 bytes per parameter
- Add 20% overhead for the runtime
Step 1: Find your GPU memory
Before estimating, know what you have. On Linux, nvidia-smi gives you total and used memory. On Windows, you can use the same command if you have the NVIDIA driver tools installed, or use Task Manager.
nvidia-smi --query-gpu=name,memory.total,memory.free --format=csvStep 2: Estimate with a simple script
Use this Python script to estimate VRAM for any model size and precision. It is a back-of-the-envelope calculation, but it is accurate enough to decide which GPU to buy or whether to download a model.
def estimate_vram(params_b, precision_bits):
bytes_per_param = precision_bits / 8
weights_gb = params_b * bytes_per_param / 1e9
overhead_gb = weights_gb * 0.2
return weights_gb + overhead_gb
for params_b in [1, 3, 7, 13, 70]:
for bits in [16, 8, 4]:
vram = estimate_vram(params_b, bits)
print(f'{params_b}B model, {bits}-bit: ~{vram:.1f} GB')
print()- Run this script before downloading a model.
- It does not account for context length. Longer context increases KV cache, which can add gigabytes.
- It assumes no CPU offloading. If you offload layers to system RAM, you can run with less VRAM but slower.
Step 3: Measure real usage with llama.cpp
First, install llama.cpp and download a small model, for example a 1B parameter model in 8-bit. Then run it and read the memory usage from nvidia-smi.
git clone https://github.com/ggerganov/llama.cpp
cd llama.cpp
make -j4
# Download a small model (e.g., Qwen2.5-1.5B-Instruct GGUF)
wget https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct-GGUF/resolve/main/qwen2.5-1.5b-instruct-q8_0.gguf
# Run with 4 GB VRAM limit (or omit -ngl for CPU only)
./llama-cli -m qwen2.5-1.5b-instruct-q8_0.gguf -n 128 -ngl 99- The -ngl flag sets the number of layers to offload to GPU. 99 means all layers.
- Watch nvidia-smi in another terminal to see peak VRAM usage.
- If you get an out-of-memory error, reduce -ngl until it fits.
Step 4: Measure with Hugging Face transformers
For PyTorch users, the transformers library gives you a direct way to see memory usage. This snippet loads a model and prints the allocated VRAM.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "HuggingFaceTB/SmolLM2-1.7B-Instruct"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype=torch.float16).to("cuda")
# Trigger a forward pass to allocate memory
inputs = tokenizer("Hello", return_tensors="pt").to("cuda")
with torch.no_grad():
outputs = model(**inputs)
allocated = torch.cuda.memory_allocated() / 1e9
reserved = torch.cuda.memory_reserved() / 1e9
print(f"Allocated: {allocated:.2f} GB")
print(f"Reserved: {reserved:.2f} GB")- Run this with different precisions: float16, int8, int4 (using bitsandbytes).
- The reserved memory is what nvidia-smi shows, usually higher than allocated.
- You can see how quantization reduces the footprint.
Step 5: Compare quantization levels
In practice, a 7B model at 4-bit can run on 6 GB VRAM, while at 16-bit it needs around 16 GB. The quality drop from 8-bit to 4-bit is often negligible for many tasks, but you should test with your own prompts.
python -c "
params = 7_000_000_000
for bits in [16, 8, 4]:
bytes_per_param = bits / 8
weights_gb = params * bytes_per_param / 1e9
total_gb = weights_gb * 1.2
print(f'{bits}-bit: weights {weights_gb:.1f} GB, total ~{total_gb:.1f} GB')
"What I would do: Recommended setup
For most developers, I recommend a GPU with at least 12 GB of VRAM. That lets you run 7B models at 8-bit with a decent context length, or 13B models at 4-bit. If you are on a budget, 8 GB can work for 7B at 4-bit, but you will be limited to short contexts.
Here is a copy-paste setup for a 7B model at 8-bit using llama.cpp, which is the sweet spot for quality and speed.
# Install llama.cpp and download a 7B model in Q8_0
cd ~
git clone https://github.com/ggerganov/llama.cpp
cd llama.cpp
make -j4
wget https://huggingface.co/TheBloke/Llama-2-7B-Chat-GGUF/resolve/main/llama-2-7b-chat.Q8_0.gguf
# Run with all layers on GPU, 512 token context
./llama-cli -m llama-2-7b-chat.Q8_0.gguf -n 512 -ngl 99 --ctx-size 2048- If you have 8 GB, use a Q4_K_M quantized model instead of Q8_0.
- Set --ctx-size to a value that fits your VRAM. Larger context uses more memory.
- Use --no-mmap to avoid disk swapping if you have enough RAM.
Troubleshooting common VRAM errors
The most common error is CUDA out of memory. The fix is either reduce the context length, use a smaller quant, or offload fewer layers to the GPU.
# Example: run a 7B model with only 4 GB VRAM by offloading half the layers
./llama-cli -m llama-2-7b-chat.Q4_K_M.gguf -n 256 -ngl 20 --ctx-size 1024- CUDA out of memory: reduce -ngl or --ctx-size, or switch to a 4-bit quant.
- Model loads but is very slow: you are likely offloading to CPU. Check nvidia-smi to see how many layers are on GPU.
- llama.cpp reports insufficient memory for mmap: use --no-mmap.
- PyTorch gives an error about mixed precision: set torch_dtype consistently.
FAQ
Q: Can I run a 70B model on 8 GB VRAM?
A: Only with aggressive quantization (2-bit) and heavy CPU offloading, which is extremely slow. Not practical for real use.
Q: Does more VRAM always mean faster inference?
A: Not always. If the model fits in VRAM, speed is mostly limited by memory bandwidth. A 4090 with 24 GB is faster than a 3090 with 24 GB because of bandwidth.
Q: Is 16 GB enough for coding assistants like CodeLlama 34B?
A: At 4-bit, yes, but with a small context. For full 34B at 8-bit, you need 32 GB.
Q: What about Apple Silicon unified memory?
A: The same math applies, but the memory is shared with the CPU. An M1 Max with 64 GB can run 70B models at 4-bit because the OS can swap.
Next step
Now that you know your VRAM budget, run the estimation script for the model you want to use. Then download a GGUF quant that fits, and measure real usage with nvidia-smi. That will tell you the exact numbers for your setup.
python -c "
params = int(input('Model size in billions: '))
bits = int(input('Precision bits (4, 8, 16): '))
bytes_per_param = bits / 8
weights_gb = params * 1e9 * bytes_per_param / 1e9
print(f'Weights: {weights_gb:.1f} GB')
print(f'Estimated total: {weights_gb * 1.2:.1f} GB')
"Key takeaways
- Apply one concrete change from this post before collecting more reading.
- Prefer browser-side tools when the work involves secrets, tokens, or PII.
- Document the why next to the how so the next reviewer inherits context.
FAQ
- Who is this guide on vram for?
- Working developers who need a practical take on how much vram do you actually need to run an ai model? — not a marketing overview. Skim the sections, apply one tip, then come back when you hit an edge case.
- Do I need an account to use the related tools?
- No. code.live tools run in your browser with no signup. Nothing you paste is uploaded to a server for the client-side utilities linked from this post.
- How often is this article updated?
- This post was published October 4, 2026. Fundamentals stay stable; check linked tool pages and official docs when version-specific behavior matters.