An investigation · 2023 → 2026

Small local models, year over year

Open-weight LLMs that fit on consumer hardware went from “interesting toy” to “GPT-4-class in 24 GB” in three years. This dashboard traces that curve across three intelligence benchmarks: MMLU, HumanEval, and Chatbot Arena Elo.

Models tracked
47
Best MMLU, 2023
35.1
LLaMA 7B
Best MMLU, 2026
89
Qwen 3.6 27B
First ≤4B > 70 MMLU

Intelligence over time

Each dot is one open-weight model release. Size ∝ parameters. Line traces the best score achieved to date.

What could I run on 24 GB?

Best MMLU achievable in a given VRAM budget, quarter by quarter.

Today's champion at ≤ 24 GB
Qwen 3.6 27B
89% MMLU
27.8B params · 24 GB @ Q6 · released Apr 2026
QuarterBest model (≤ 24 GB)ParamsMMLU
2023-Q1LLaMA 33B33B57.8%
2023-Q3Mistral 7B7.3B60.1%
2023-Q4Mistral 7B7.3B60.1%
2024-Q1Gemma 7B7B64.3%
2024-Q2Phi-3 medium 14B14B78%
2024-Q3Qwen2.5 32B Instruct32.5B83.3%
2024-Q4Phi-4 14B14B84.8%
2025-Q1Phi-4 14B14B84.8%
2025-Q2Qwen3 32B32.8B85.1%
2025-Q3Qwen3 32B32.8B85.1%
2026-Q1GLM-4.7-Flash31.2B86%
2026-Q2Qwen 3.6 27B27.8B89%

Milestones

  1. Sep 2023
    Mistral-7B matches Llama-2-13B
    First widely-used 7B to leap past a 13B on MMLU (60.1 vs 54.8) — the 'smaller can win' era begins.
  2. Apr 2024
    Phi-3-mini fits Llama-2-70B on a phone
    3.8B parameters hit MMLU 68.8, matching Meta's 70B from nine months earlier at ~20× less VRAM.
  3. Sep 2024
    Qwen2.5-7B passes GPT-3.5-class coding
    HumanEval 84.8 on a 5 GB laptop model — coding capability once reserved for API-only frontier models.
  4. Dec 2024
    Phi-4 14B crosses MMLU 85
    A 14B dense model reaches accuracy Meta needed 70B for in April, on the same 9 GB VRAM budget.
  5. Jan 2025
    DeepSeek-R1 distills bring o1-class reasoning to 32B
    Chain-of-thought reasoning distilled from a 671B teacher into a 32B student runnable on a single 24 GB card.
  6. Aug 2025
    gpt-oss-20B rivals proprietary mid-tier
    OpenAI-released open weights: 20B MoE (3.6B active) at MMLU 82.5, HumanEval 89.5 — proprietary quality, apache-2.0 license.
  7. Apr 2026
    Qwen 3.6 27B: frontier fits in 24 GB
    A single consumer GPU now runs a model that would have been GPT-4-class two years earlier.

Size class comparison

Compare tiers by:

By size class · MMLU

Running best score for dense models in each parameter tier. Smaller tiers catch up to yesterday's giants.

≤ 3B
56.7%
≤ 8B
74.2%
≤ 14B
84.8%
≤ 32B
89%

Full dataset

47 model releases. Click a column to sort.

Date ModelFamilyParams VRAM MMLU HumanEval Arena Elo License
2026-04-21Qwen 3.6 27BQwen 327.8B24 GB89941395apache-2.0
2026-03-11Gemma 4 26B-A4B ITGemma26.5 (3.8a)B28 GB87.4931370apache-2.0
2026-03-11Gemma 4 31B ITGemma32.7B29 GB88.193.41381apache-2.0
2026-01-19GLM-4.7-FlashGLM31.2 (12a)B23 GB8692.51352mit
2025-09-09Qwen3-Next 80B-A3BQwen 381.3 (3a)B53 GB85.891.81348apache-2.0
2025-08-04gpt-oss 20Bgpt-oss21.5 (3.6a)B23 GB82.589.51315apache-2.0
2025-08-04gpt-oss 120Bgpt-oss120.4 (5.1a)B51 GB87.292.31361apache-2.0
2025-07-20GLM-4.5-AirGLM110.5 (12a)B60 GB84.791.51341mit
2025-04-27Qwen3 14BQwen 314.8B12 GB82.488.41288apache-2.0
2025-04-27Qwen3 32BQwen 332.8B20 GB85.190.21332apache-2.0
2025-04-27Qwen3 30B-A3BQwen 330.5 (3a)B32 GB82.789.61304apache-2.0
2025-04-02Llama 4 Scout 17B-16ELlama 4109 (17a)B60 GB79.679.61272llama-4
2025-03-12Gemma 3 27B ITGemma27B17 GB78.687.81338gemma
2025-03-05QwQ 32BQwen32.8B20 GB82.5891314apache-2.0
2025-01-20DeepSeek-R1-Distill-Qwen 32BDeepSeek32.8B20 GB82.988.41305mit
2024-12-12Phi-4 14BPhi14B9 GB84.882.61229mit
2024-11-11Qwen2.5-Coder 32BQwen32.5B20 GB75.192.7apache-2.0
2024-11-01SmolLM2 1.7BSmolLM1.7B1 GB50.323.3apache-2.0
2024-09-25Llama 3.2 1BLlama 31.2B1 GB32.2llama-3.2
2024-09-25Llama 3.2 3BLlama 33.2B2 GB63.41103llama-3.2
2024-09-19Qwen2.5 7B InstructQwen7B5 GB74.284.81189apache-2.0
2024-09-19Qwen2.5 14B InstructQwen14B9 GB79.783.51207apache-2.0
2024-09-19Qwen2.5 32B InstructQwen32.5B20 GB83.388.41257apache-2.0
2024-09-19Qwen2.5 72B InstructQwen72B42 GB86.186.61268tongyi-qianwen
2024-09-17Mistral Small 22BMistral22B13 GB72421119mrl
2024-07-23Llama 3.1 8B InstructLlama 38B6 GB7372.61176llama-3.1
2024-07-23Llama 3.1 70B InstructLlama 370B40 GB83.680.51247llama-3.1
2024-07-18Mistral Nemo 12BMistral12.2B8 GB6840.31073apache-2.0
2024-06-27Gemma 2 9BGemma9B6 GB71.340.21189gemma
2024-06-27Gemma 2 27BGemma27B17 GB75.251.81218gemma
2024-06-06Qwen2 7B InstructQwen7B5 GB70.379.91141apache-2.0
2024-06-06Qwen2 72B InstructQwen72B42 GB82.3861187tongyi-qianwen
2024-05-21Phi-3 medium 14BPhi14B9 GB7862.2mit
2024-04-23Phi-3 mini 3.8BPhi3.8B3 GB68.858.51037mit
2024-04-18Llama 3 8B InstructLlama 38B6 GB68.462.21152llama-3
2024-04-18Llama 3 70B InstructLlama 370B40 GB8281.71206llama-3
2024-02-21Gemma 7BGemma7B5 GB64.332.31037gemma
2023-12-13Phi-2 2.7BPhi2.7B2 GB56.747mit
2023-12-11Mixtral 8x7BMistral46.7 (12.9a)B28 GB70.640.21114apache-2.0
2023-09-27Mistral 7BMistral7.3B5 GB60.130.51071apache-2.0
2023-07-18Llama 2 7BLlama 27B5 GB45.312.81037llama-2
2023-07-18Llama 2 13BLlama 213B8 GB54.818.31063llama-2
2023-07-18Llama 2 70BLlama 270B40 GB68.929.91093llama-2
2023-02-24LLaMA 7BLLaMA7B14 GB35.110.5llama-1
2023-02-24LLaMA 13BLLaMA13B26 GB46.915.8llama-1
2023-02-24LLaMA 33BLLaMA33B20 GB57.821.7llama-1
2023-02-24LLaMA 65BLLaMA65B38 GB63.423.7llama-1