Find the best AI models your hardware can host
Check local AI compatibility based on your GPU, VRAM, RAM and CPU. Get practical model recommendations, estimated memory usage and runtime guidance — without guessing.
Check your hardware
Your system
Not sure about a value? A close estimate still produces a useful compatibility analysis.
Recommended for your PC
Prefer reading first? Browse the guides to understand what VRAM, quantization and fit levels mean before you check.
Learn before you choose
Compatibility is more than a yes/no answer. These guides explain the trade-offs behind every result.
Why VRAM matters
Model weights must live somewhere. Dedicated GPU memory is the fastest place to put them — and the first resource to run out.
Read guideWhat quantization means
Storing weights at 4 bits instead of 16 changes everything: what fits, what runs fast, and how much quality you trade.
Read guideWhat fit levels mean
Perfect, good, tight, too tight — how llmfit grades each model against your hardware, and how to read the estimates.
Read guideExplore by GPU
Start from the graphics card: see what model classes fit and how VRAM changes your options.
NVIDIA GeForce RTX 4050 Laptop GPU
6 GB VRAM3B (Q8/FP16) · 7B (Q4/Q5) · 8B (Q4)
NVIDIA GeForce RTX 4060
8 GB VRAM7B (Q5/Q6) · 8B (Q4/Q5) · 12B (Q3/Q4) · 14B (Q3)
NVIDIA GeForce RTX 4070
12 GB VRAM7B (Q8) · 8B (Q6/Q8) · 12B (Q5) · 14B (Q4/Q5)
NVIDIA GeForce RTX 4090
24 GB VRAM14B (FP16) · 32B (Q4/Q5) · 70B (Q2/Q3 with offload)
AMD Radeon RX 7900 XTX
24 GB VRAM14B (FP16) · 32B (Q4/Q5) · 70B (Q2/Q3 with offload)
Apple M4 Pro
27 GB unified7B (FP16) · 14B (Q8) · 32B (Q4)
Popular guides
Straight answers to the questions everyone asks when starting with local AI.
How much VRAM do I need for local LLMs?
A practical breakdown by model size, quantization and context — from 4 GB laptops to 24 GB desktops.
Read guideVRAM vs RAM for local AI
Where model weights live, when system RAM helps, and why VRAM is usually the first wall you hit.
Read guideWhat is GGUF?
The file format behind most local models — how quantized builds pack weights and what the labels mean.
Read guideCPU vs GPU for LLM inference
Why inference speed tracks memory bandwidth, and when a CPU-only setup is genuinely usable.
Read guideReady to check your PC?
The checker is free, needs no account, and gives you a ranked list of models that fit your GPU, RAM and CPU in seconds.
Check my hardware