Which AI Models Fit Your Laptop? One Free Command Answers
“Can my laptop run AI locally?” is the most common question I get from students — and the usual answer is a shrug and a 40-minute download that turns out too slow to use. There is a one-command way to know before you download anything.
The tool: llmfit
llmfit detects your machine’s RAM, CPU and GPU, then scores every open model in its database for fit, speed and quality — right-sizing models to your hardware. It runs locally, needs no API key, and works on Apple Silicon, NVIDIA and AMD.
Install (one line)
brew install llmfit
(Or the cross-platform installer on the repo.)
See what your machine can run
First, what llmfit sees:
llmfit system
It prints your CPU, total and available RAM, memory bandwidth, and GPU. Then the ranking:
llmfit
You get a table of models, each marked 🟢 Perfect / 🟡 Good / red (won’t fit), with estimated tokens/sec, memory use and quantization. Green means download it today. Red means don’t waste the afternoon.
Checking a specific laptop (e.g. 8 GB)
llmfit lets you override the detected hardware, so you can plan for a machine you don’t own yet:
llmfit --ram 8G
On an 8 GB laptop, the top fits are small reasoning models — DeepSeek-R1-Distill-Qwen-7B, DeepSeek-R1-0528-Qwen3-8B, and Qwen3-4B — all comfortably green. Filter by name (press / in the interactive view and type deepseek) and llmfit shows exactly how many of that family fit.
Running models locally is step one — building with them is the skill
DeployU teaches you to build and deploy real AI systems on real infrastructure, not just run a chatbot.
My 3 starter picks by RAM
- 8 GB: DeepSeek-R1-Distill-Qwen-7B (reasoning that actually fits) or Qwen3-4B for speed.
- 16 GB: an 8B model at higher quantization — DeepSeek-R1-0528-Qwen3-8B is a strong daily driver.
- ≥32 GB / good GPU: step up to 14B+ models; run
llmfitand take the top 🟢 rows.
Pair any of these with Ollama (ollama run <model>) and you have a private, offline AI on your own machine.
The honest limits
- Fit ≠ fast enough. llmfit estimates tokens/sec; a model can “fit” and still feel slow. Check the tok/s column, not just the green dot.
- It’s an estimate. Real speed depends on quantization, context length and background load.
- It ranks open models, not closed APIs — this is about what you can run yourself, free.
The next time someone asks “will this run on my laptop?”, you don’t guess — you run one command and know.
From running models to shipping AI
DeployU turns “I ran a local model” into deployable, portfolio-ready AI projects on real cloud accounts.
