NVIDIA · 32GB VRAM
What AI models can a GeForce RTX 5090 run?
A GeForce RTX 5090 has 32GB of VRAM. Checked against 228 models from the Ollama library, 194 run fully on the GPU at 4K context and 10 more run with part of the model in system RAM.
194
run well
10
with trade-offs
228
models tested
Best AI models for a GeForce RTX 5090
Largest first — these fit entirely in 32GB of VRAM at 4K context, so the whole model is GPU-accelerated.
| Model | Build | Memory needed | Verdict |
|---|---|---|---|
| laguna-xs-2.1Ollama Library | latest | ~27.4 GB | Good fit |
| codeboogaOllama Library | 34B | ~27.3 GB | Good fit |
| phind-codellamaMicrosoft | 34B | ~27.3 GB | Good fit |
| qwqAlibaba Cloud | 32B | ~26.8 GB | Good fit |
| command-rCohere | 35B | ~26.8 GB | Good fit |
| olmo-3.1Allen Institute for AI | 32B | ~26.3 GB | Good fit |
| glm-4.7-flashZhipu AI | latest | ~25.7 GB | Good fit |
| north-mini-code-1.0Ollama Library | latest | ~25.2 GB | Good fit |
| qwen3-coderAlibaba Cloud | 30B | ~25.1 GB | Good fit |
| qwen3.6Alibaba Cloud | 27B | ~23.6 GB | Good fit |
| qwen3.8Alibaba Cloud | 27B | ~22.8 GB | Excellent fit |
| muse-glimmerOllama Library | 30B | ~22.7 GB | Excellent fit |
| mistral-small3.1Mistral AI | 24B | ~21.0 GB | Excellent fit |
| devstral-small-2Mistral AI | 24B | ~20.6 GB | Excellent fit |
| mistral-small3.2Mistral AI | 24B | ~20.6 GB | Excellent fit |
| lfm2Ollama Library | 24B | ~19.6 GB | Excellent fit |
| devstralMistral AI | 24B | ~19.5 GB | Excellent fit |
| magistralMistral AI | 24B | ~19.5 GB | Excellent fit |
| gpt-ossOpenAI | 20B | ~18.8 GB | Excellent fit |
| gpt-oss-safeguardOpenAI | 20B | ~18.8 GB | Excellent fit |
| mistral-smallMistral AI | 22B | ~18.2 GB | Excellent fit |
| codestralMistral AI | 22B | ~18.2 GB | Excellent fit |
| solar-proUpstage | 22B | ~18.1 GB | Excellent fit |
| phi4-reasoningMicrosoft | 14B | ~15.2 GB | Excellent fit |
| deepseek-coder-v2DeepSeek AI | 16B | ~14.2 GB | Excellent fit |
Showing the 25 largest of 194 models that run well.
How these results were calculated
Each model is evaluated at 4K context against 32GB of VRAM, counting model weights, KV cache, compute buffers and runtime overhead, minus a reserve for the display and operating system. Download sizes come from the official Ollama registry.
These figures assume 32GB of system RAM and working GPU drivers. Your own machine may differ — RAM, free disk space and whether an acceleration backend is actually installed all change the answer, which is what the PC scan measures directly.
GeForce RTX 5090 local AI — frequently asked questions
- How many AI models can a GeForce RTX 5090 run?
- Out of 228 models in the Ollama library, 194 run fully on a GeForce RTX 5090's 32GB of VRAM at 4K context, and a further 10 run with part of the model offloaded to system RAM, more slowly.
- What is the largest AI model a GeForce RTX 5090 can run?
- The heaviest build that fits entirely in VRAM is laguna-xs-2.1:latest, needing about 27.4 GB of memory from a 18.9 GB download.
- Is 32GB of VRAM enough for local AI?
- 32GB is enough for 194 of the 228 models tested here, which covers most general-purpose and coding assistants. Very large models still need either partial CPU offload or a card with more memory.