Large Language Models

A comparative timeline of frontier & open-weight models · 2024–2026 · Hover dots for details

Last updated: Jul 18, 2026
Comparison
Timeline Range Jul 2025 – Jul 2026
Jan 2024 Aug 2026

Model Name

Model Name

Tooltip – Model Info
Release date
Parameter count
Context window
Knowledge cutoff
AA Ratings (0–10 segments)
CD — Cognitive Depth
CO — Coding
AG — Agentic Reliability
SC — Scientific Reasoning
UX — UX & Empathy
Timeline Dots
Flagship (proprietary)
Release (proprietary)
Flagship (open weights)
Release (open weights)
Not yet released (leak / teaser)
Modality Badges
LLM Text / Language
VIS Image understanding
IMG Image generation
AUD Audio understanding
VID Video understanding
VGEN Video generation
Index Methodology
Cognitive Depth
Direct AA Intelligence Index score, normalised so the current best model = 10 segments.
Coding
Weighted composite of two AA coding benchmarks, normalised to best model = 10 segments.
Agentic Reliability
Weighted composite of two AA agentic benchmarks, normalised to best model = 10 segments.
Scientific Reasoning
Weighted composite of three AA science benchmarks, normalised to best model = 10 segments.
HLE 50%
CritPt 25%
UX & Empathy
Weighted composite of four EQ-Bench 3 sub-scores, min-max normalised across all models to 1–10 segments.