The full local AI model catalogue
Nearly 7,000 models with parameters, quantization, disk size, estimated speed and a fit rating computed against your hardware profile. Search by name or provider, sort by score or speed, and open a model for a detailed estimate of what it takes to run it.
| Model | Provider | Params | Score | tok/s | Quant | Disk | Mode | Ctx | Use case | Fit |
|---|
How to read this catalogue
The Score column is the llmfit engine's overall rating for your profile — it blends fit, expected speed, quality and usable context. tok/s is the estimated generation speed for a reference machine (6 GB VRAM, 24 GB RAM, 12 threads); your numbers will differ, which is what the hardware checker is for. Fit grades memory: green means the model fits comfortably, amber means it will partially offload, red means it will not load usefully.
Thousands of these entries are community fine-tunes and quantization variants of the same base families. The names worth learning first are the canonical ones — Llama, Qwen, Gemma, Mistral, Phi — and our VRAM guide explains how to read the sizes. Estimates, not benchmarks: the methodology page documents exactly how each number is produced and where it can be wrong.
Start with the classics
Detailed editorial pages for the most popular model families.