Somewhere around the ninth “OpenAI just announced” Slack message in a single week, I gave up trying to hold the timeline in my head. Model names stopped mapping to release dates, release dates stopped mapping to what actually changed, and “the new one” became ambiguous across at least four labs simultaneously. That’s not a personal failing so much as a description of the last two years: the industry went from a handful of frontier labs shipping a model every few months to dozens of labs (frontier and open-weight alike) shipping something worth knowing about most weeks.
So I built the timeline I wished existed: every model below is placed in the era it actually shaped, carries the specs a developer would check before reaching for it (parameters, context window, license, how to run it), and links back to a first-party source. Nothing here is guessed. Where a spec wasn’t publicly disclosed, the card just doesn’t show it, rather than making one up.
Four eras structure the scroll:
The Chat Era (through August 2024): scale and instruction-tuning as the whole story, one-shot assistants competing on breadth.
The Reasoning Turn (September 2024 to June 2025): test-time compute becomes a first-class feature, chain-of-thought moves from prompt trick to trained behavior.
The Agentic Era (July 2025 to February 2026): long-horizon tool use, coding agents, and multi-step planning become the headline capability.
Efficiency & Hybrids (March 2026 onward): sparse MoE and linear-attention hybrids push frontier-adjacent capability down into models small enough to run locally.
Browse chronologically within an era, or jump straight to one from the rail below the search bar. Click any card for the full writeup: specs, how to access it, and the one sentence on why it mattered. The search box on top isn’t just a text filter either: it understands org:, era:, license:, open:true, and numeric filters like params:>100b, so you can slice the whole list by whatever axis you actually care about.
∅
No models found
Try a broader term or clear the active filter.
Reset search
OpenAI
Mar 2023
GPT-4
OpenAI's first large multimodal flagship, setting the bar for reasoning and exam performance that the whole industry chased for a year.
Anthropic
Mar 2023
Claude 1
Anthropic's first publicly available model, launched the same week as GPT-4 and establishing Claude as a frontier-lab contender.
Google
May 2023
PaLM 2
Google's pre-Gemini flagship, powering Bard and 25+ products; the last major PaLM-line model before the Gemini unification.
TII
May 2023
Falcon 40B
One of the first genuinely commercially-usable open-weight frontier-scale models, briefly topping the Open LLM Leaderboard.
Anthropic
Jul 2023
Claude 2
Anthropic's first model available to the general public rather than a waitlist, with a then-huge 100K context window.
Meta
Jul 2023
Llama 2
Free for commercial use and co-released with Microsoft, it kicked off the modern open-weights ecosystem and countless fine-tunes.
Mistral AI
Sep 2023
Mistral 7B
Punched far above its weight class, beating Llama 2 13B, and put Mistral AI on the map as a serious open-weights lab.
Google
Dec 2023
Gemini 1.0
Google's first natively multimodal model family (Nano/Pro/Ultra), unifying its LLM efforts under the Gemini brand for the first time.
Google
Feb 2024
Gemini 1.5 Pro
Broke the long-context ceiling with a 1M-token window, five times larger than anything publicly available at the time.
Google
Feb 2024
Gemma
Google's first open-weight release built on Gemini research, marking its entry into the open-model ecosystem after years of closed models.
Anthropic
Mar 2024
Claude 3 Opus
Top of the Claude 3 family and Anthropic's first model to credibly outperform GPT-4 on several public benchmarks.
Anthropic
Mar 2024
Claude 3 Haiku
Anthropic's fastest, cheapest model at launch, establishing the small-fast tier that every lab now ships alongside its flagship.
Meta
Apr 2024
Llama 3
Meta's 'best open model available' claim held up; Llama 3 70B closed most of the gap to closed frontier models.
OpenAI
May 2024
GPT-4o
First truly natively omni-modal flagship (text, vision, audio in one model) and OpenAI's first free-tier GPT-4-class model.
Anthropic
Jun 2024
Claude 3.5 Sonnet
A mid-tier model that outperformed Anthropic's own flagship Opus while being far cheaper, reshaping the price/performance curve industry-wide.
Meta
Jul 2024
Llama 3.1 405B
The first openly available model to rival closed frontier labs on general capability, at a scale (405B) no open lab had shipped before.
xAI
Aug 2024
Grok 2
xAI's first model to claim competitiveness with GPT-4o and Claude 3.5, plus built-in image generation via FLUX.
OpenAI
Sep 2024
OpenAI o1-preview
The model that opened the reasoning era: the first widely-available LLM built to spend inference-time compute 'thinking' before answering.
NVIDIA
Oct 2024
Llama-3.1-Nemotron-70B-Instruct
An RLHF/DPO-tuned Llama 3.1 that briefly topped alignment leaderboards ahead of GPT-4o and Claude 3.5, showing post-training alone could close the gap to frontier labs.
Anthropic
Nov 2024
Claude 3.5 Haiku
Matched the prior generation's flagship (Claude 3 Opus) on many evals while running at small-model speed and cost.
Alibaba
Nov 2024
Qwen2.5-Coder-32B
The strongest open-weight coding model of its era, claimed to match GPT-4's coding ability at a fraction of the size.
Mistral AI
Nov 2024
Pixtral Large
Mistral's first frontier-class multimodal model, built on Mistral Large 2, capable of parsing 30+ high-resolution images in context.
DeepSeek
Nov 2024
DeepSeek-R1-Lite-Preview
DeepSeek's first public reasoning model, matching OpenAI o1-preview on math and coding benchmarks and foreshadowing the R1 shock two months later.
OpenAI
Dec 2024
OpenAI o1
Full production release of o1 with a 34% reduction in major errors over the preview, plus image understanding; also debuted the $200/mo o1 pro tier.
Meta
Dec 2024
Llama 3.3 70B
Delivered near-405B performance in a 70B footprint, showing post-training gains could substitute for raw parameter count.
Google
Dec 2024
Gemini 2.0 Flash Experimental
Google's first model pitched explicitly for 'the agentic era', combining native tool use with multimodal output at Flash-tier speed.
Microsoft
Dec 2024
Phi-4
Matched or beat much larger models on STEM reasoning via aggressive synthetic-data curation rather than scale, Microsoft's clearest small-model bet.
DeepSeek
Dec 2024
DeepSeek-V3
A 671B open-weight MoE trained for a fraction of frontier-lab budgets that matched GPT-4o and Claude 3.5 Sonnet, a preview of the R1 shock to come.
DeepSeek
Jan 2025
DeepSeek-R1
Open-sourced a full o1-class reasoning model with weights and a technical report under MIT license, triggering a market shock and a wave of distilled derivatives.
Moonshot AI
Jan 2025
Kimi k1.5
Moonshot's o1-class multimodal reasoning model, notable for scaling reinforcement learning without relying on Monte Carlo tree search or process reward models.
Alibaba
Jan 2025
Qwen2.5-Max
Alibaba's largest MoE flagship, trained on 20T+ tokens and claimed to outperform GPT-4o, DeepSeek-V3 and Llama 3.1 405B across public benchmarks.
OpenAI
Jan 2025
OpenAI o3-mini
Rushed out days after the DeepSeek R1 shock, offering three selectable reasoning-effort tiers and free-tier ChatGPT access to a reasoning model for the first time.
xAI
Feb 2025
Grok 3
Trained on the 100K-GPU Colossus cluster (10x the compute of Grok 2) and launched with a built-in DeepSearch agentic research tool.
Anthropic
Feb 2025
Claude 3.7 Sonnet
Billed as the first 'hybrid reasoning' model, letting one model switch between instant answers and visible extended thinking; shipped alongside Claude Code.
OpenAI
Feb 2025
GPT-4.5
OpenAI's largest non-reasoning pretrained model ('Orion'), pitched on conversational warmth and lower hallucination rates rather than benchmark gains, and later deprecated in favor of GPT-4.1.
Alibaba
Mar 2025
QwQ-32B
A 32B open reasoning model that rivaled DeepSeek-R1-level results on math and coding at a fraction of the parameter count.
Google
Mar 2025
Gemma 3 27B
Added multimodality and 140+ language support to Gemma while beating Gemini 1.5 Pro on several benchmarks in a fully open-weight package.
Cohere
Mar 2025
Command A
Cohere's enterprise-focused flagship, running on just two GPUs while delivering 150% higher throughput than its predecessor Command R+.
Baidu
Mar 2025
ERNIE 4.5
Baidu's flagship multimodal-native model family, later open-sourced across ten variants in June 2025 after initially launching as a paid API.
Google
Mar 2025
Gemini 2.5 Pro
Made 'thinking' the default across the whole Gemini lineup for the first time and debuted at #1 on the LMArena leaderboard.
Meta
Apr 2025
Llama 4 Maverick
Meta's first MoE-architecture Llama, and its first natively multimodal open flagship, though it drew controversy over benchmark-version discrepancies.
Meta
Apr 2025
Llama 4 Scout
Shipped with an unprecedented 10-million-token context window while remaining runnable on a single H100 GPU.
OpenAI
Apr 2025
GPT-4.1
A developer-focused, API-first release emphasizing coding and long-context reliability, shipped alongside 4.1 mini and nano tiers.
OpenAI
Apr 2025
OpenAI o3
OpenAI's most advanced reasoning model at launch, the first to natively chain tool use (web browsing, code execution, image generation) into its reasoning trace.
OpenAI
Apr 2025
OpenAI o4-mini
A cost-efficient sibling to o3 that reached full ChatGPT free-tier rollout within a week, mainstreaming agentic tool-using reasoning.
Alibaba
Apr 2025
Qwen3
Alibaba's first hybrid-reasoning model family, spanning eight sizes from 0.6B to a 235B MoE flagship, all open-sourced simultaneously.
Alibaba
Apr 2025
Qwen3-235B-A22B
Qwen3's flagship MoE introduced a single model that toggles between 'thinking' and 'non-thinking' modes, competing with DeepSeek-R1, o1 and Gemini 2.5 Pro under Apache 2.0.
Alibaba
Apr 2025
Qwen3-32B
The dense flagship of the Qwen3 launch, offering hybrid reasoning at a size runnable on a single high-end GPU.
Alibaba
Apr 2025
Qwen3-30B-A3B
A small, fast MoE variant of Qwen3 that made hybrid reasoning practical on consumer hardware.
Microsoft
Apr 2025
Phi-4 Reasoning Plus
Showed a 14B dense model, fine-tuned on 1.4M STEM/coding questions plus an RL stage, could approach much larger reasoning models.
Microsoft
Apr 2025
Phi-4 Mini Reasoning
A sub-4B reasoning model targeted at latency- and memory-constrained edge deployments, trained on 150B tokens.
Mistral AI
May 2025
Devstral
An open-weight software-engineering agent model, built with All Hands AI, that led open models on SWE-bench Verified at launch.
Mistral AI
May 2025
Mistral Medium 3
Mistral's pitch that mid-sized models could hit frontier-class performance at 'an order of magnitude less' cost than giant proprietary models.
Anthropic
May 2025
Claude Sonnet 4
Anthropic's mid-tier Claude 4 model, later extended to a 1M-token context window, became the default workhorse for Claude Code.
Anthropic
May 2025
Claude Opus 4
Anthropic's flagship Claude 4 launch, billed as its best coding model at the time and the debut of extended hybrid reasoning at the top tier.
OpenAI
Jun 2025
o3-pro
OpenAI's most advanced reasoning model at the time, replacing o1-pro for ChatGPT Pro/Team users and the API.
Mistral AI
Jun 2025
Magistral Medium
Mistral's first dedicated reasoning model family, emphasizing transparent, verifiable chain-of-thought across professional domains and languages.
Google
Jun 2025
Gemma 3n
A mobile-first multimodal model inheriting Gemini Nano's architecture, able to process text, image, video and audio fully offline.
Tencent
Jun 2025
Hunyuan-A13B
Tencent's first open-source hybrid fast/slow-thinking MoE, scoring on par with o1 and DeepSeek on several benchmarks while running on one accelerator card.
Baidu
Jun 2025
ERNIE-4.5-300B-A47B
Baidu's reversal from closed to open source, releasing its flagship ERNIE MoE family and claiming wins over DeepSeek-V3 on 22/28 benchmarks.
Liquid AI
Jul 2025
LFM2
Liquid AI's second-generation on-device foundation models, claiming 200% faster CPU decode/prefill than Qwen3 and Gemma 3 at comparable sizes.
Moonshot AI
Jul 2025
Kimi K2
The first trillion-parameter open-weight model widely available under a permissive license, notable for strong agentic and coding performance and the novel MuonClip optimizer.
xAI
Jul 2025
Grok 4
xAI's frontier launch that paired a $300/month subscription tier with strong reasoning-benchmark results, including a Humanity's Last Exam record via Grok 4 Heavy.
xAI
Jul 2025
Grok 4 Heavy
A multi-agent variant of Grok 4 that runs several reasoning agents in parallel, launched as the top tier of xAI's $300/month plan.
LG AI Research
Jul 2025
EXAONE 4.0
Korea's first open-weight hybrid reasoning model, notable for passing written exams for six national professional qualifications.
Google
Jul 2025
Gemini 2.5 Flash-Lite
Google's cheapest, fastest 2.5-series model, aimed at high-volume classification, routing and translation workloads.
Alibaba
Jul 2025
Qwen3-Coder (480B-A35B-Instruct)
Alibaba's most powerful open agentic coding model, purpose-built for multi-file, tool-using software engineering workflows.
Z.ai
Jul 2025
GLM-4.5
Z.ai's flagship agentic reasoning model, released under a fully auditable MIT license and positioned to match Claude/DeepSeek-tier performance.
Z.ai
Jul 2025
GLM-4.5-Air
A compact, self-hostable companion to GLM-4.5 that made Z.ai's agentic reasoning stack practical on smaller GPU budgets.
Google
Aug 2025
Gemini 2.5 Pro Deep Think
An extended-reasoning Gemini variant that achieved gold-medal-standard performance at the 2025 International Mathematical Olympiad.
Z.ai
Aug 2025
GLM-4.5V
Extended GLM-4.5's agentic reasoning stack to vision-language tasks, open-weighted under MIT.
Anthropic
Aug 2025
Claude Opus 4.1
An incremental but notable Opus refresh that pushed agentic coding performance from 72.5% to 74.5% on SWE-bench Verified.
OpenAI
Aug 2025
gpt-oss-120b
OpenAI's first open-weight model release since GPT-2, marking a strategic return to open weights with a configurable reasoning-effort MoE.
OpenAI
Aug 2025
gpt-oss-20b
The smaller sibling in OpenAI's open-weight return, sized to run on consumer hardware while retaining strong tool-use and reasoning performance.
OpenAI
Aug 2025
GPT-5
OpenAI's unified flagship merging chat and reasoning models into a single system with automatic routing, though the rocky rollout forced a partial rollback of GPT-4o's removal.
NVIDIA
Aug 2025
NVIDIA Nemotron Nano 2 (9B)
Replaced most self-attention layers with Mamba-2 state-space layers, delivering up to 6x higher reasoning throughput than same-sized Qwen3-8B.
DeepSeek
Aug 2025
DeepSeek-V3.1
Merged V3 and R1 into a single hybrid model that switches between fast direct answers and R1-style chain-of-thought via one chat template.
Google
Aug 2025
Gemini 2.5 Flash Image (Nano Banana)
Google's state-of-the-art native image generation/editing model ('Nano Banana'), notable for character-consistent multi-image blending.
Other
Aug 2025
Hermes 4
Nous Research's frontier-scale open-weight family adding toggleable hybrid reasoning on top of Llama 3.1, released with a 94-page technical report.
xAI
Aug 2025
Grok Code Fast 1
xAI's first dedicated coding model, a from-scratch architecture launched free across GitHub Copilot, Cursor, and other coding tools.
xAI
Sep 2025
Grok 4 Fast
A cost-efficiency play matching Grok 4 quality with 40% fewer reasoning tokens and a 98% price cut, extending context to 2M tokens.
Other
Sep 2025
Apertus-70B
A fully transparent, fully open (data, weights, and training recipe) multilingual LLM from ETH Zurich, EPFL and CSCS, trained on 15T tokens across 1,000+ languages.
Alibaba
Sep 2025
Qwen3-Next-80B-A3B
Previewed a new hybrid linear-attention + ultra-sparse MoE architecture aimed at drastically cutting long-context inference cost.
DeepSeek
Sep 2025
DeepSeek-V3.1-Terminus
A reliability-focused refresh of V3.1 fixing language-mixing artifacts and improving agentic tool-use consistency.
Alibaba
Sep 2025
Qwen3-Max
Alibaba's largest proprietary flagship to date, trained on 36T tokens, positioned to compete directly with GPT-5 and Claude Opus on LMArena.
Anthropic
Sep 2025
Claude Sonnet 4.5
Shipped alongside Claude Code 2.0 and billed by Anthropic as the best coding model available, able to sustain autonomous operation for 30+ hours.
DeepSeek
Sep 2025
DeepSeek-V3.2-Exp
Debuted DeepSeek Sparse Attention, a fine-grained sparse-attention mechanism cutting long-context inference cost, paired with a 50%+ API price cut.
Z.ai
Sep 2025
GLM-4.6
Z.ai's coding-focused GLM refresh, expanding context to 200K and improving tool-use/search-agent performance over GLM-4.5.
InclusionAI
Oct 2025
Ring-1T
First open-source trillion-parameter 'thinking' model, scaling MoE reinforcement learning from tens of billions to a full trillion parameters.
MiniMax AI
Oct 2025
MiniMax-M2
Open-sourced agentic/coding MoE model priced at roughly 8% of Claude while running about twice as fast, kicking off MiniMax's rapid M-series cadence.
Moonshot AI
Oct 2025
Kimi-Linear-48B-A3B-Instruct
Introduced Kimi Delta Attention, a hybrid linear-attention mechanism that outperforms full attention while cutting KV cache usage up to 75% and giving 6x decoding throughput at 1M context.
IBM
Oct 2025
Granite 4.0
First major enterprise model family to ship a production hybrid Mamba-2/Transformer stack, cutting serving memory over 70% versus comparable transformers.
Other
Oct 2025
Tiny Recursive Model (TRM)
A 7M-parameter recursive reasoner from Samsung SAIT Montreal that outperformed DeepSeek-R1, o3-mini-high and Gemini 2.5 Pro on ARC-AGI puzzles, showing recursion can substitute for scale on narrow tasks.
Anthropic
Oct 2025
Claude Haiku 4.5
Anthropic's smallest, fastest model matched Sonnet 4 performance at a third of the cost, and was the first Haiku with an extended-thinking mode.
xAI
Nov 2025
Grok 4.1
Rolled out to all users after a silent production A/B test in which it was preferred 64.8% of the time over the prior model, taking the top spot on LMArena's Text Arena.
Moonshot AI
Nov 2025
Kimi K2 Thinking
Open-source 'thinking agent' able to execute 200-300 sequential tool calls autonomously, reportedly trained for only $4.6M, with native INT4 for 2x inference speed.
OpenAI
Nov 2025
GPT-5.1
Added adaptive 'thinking time' based on task complexity, customizable personalities, and a no-reasoning fast mode, refining GPT-5's balance of speed and intelligence.
Google
Nov 2025
Gemini 3 Pro (preview)
Google's flagship multimodal and agentic/vibe-coding model, positioned as state of the art on release and triggering an OpenAI 'Code Red' response.
OpenAI
Nov 2025
GPT-5.1-Codex-Max
First OpenAI coding model natively trained to work coherently across compaction boundaries, running unattended for over 24 hours in internal tests.
Ai2
Nov 2025
Olmo 3
First fully open 32B 'thinking' model shipped with the complete model flow: data, code, checkpoints, and training pipeline, not just weights.
Anthropic
Nov 2025
Claude Opus 4.5
First model to break 80% on SWE-bench Verified, paired with a 67% price cut from prior Opus models, making frontier-tier coding broadly affordable.
Prime Intellect
Nov 2025
INTELLECT-3
Prime Intellect's fully open recipe (weights, RL frameworks, environments, and evals) for large-scale post-training RL, beating its GLM-4.5-Air base by 8 points on LiveCodeBench.
DeepSeek
Dec 2025
DeepSeek-V3.2
First model to integrate thinking directly into tool-use, and introduced DeepSeek Sparse Attention to cut long-context training/inference cost.
DeepSeek
Dec 2025
DeepSeek-V3.2-Speciale
Reasoning-maxed variant of V3.2 that reached gold-medal level on IMO, CMO, ICPC World Finals and IOI 2025, rivaling Gemini 3 Pro on hard reasoning tasks.
Mistral AI
Dec 2025
Mistral Large 3
Mistral's most capable model to date and, at 675B total parameters, the largest open-weight MoE model released by a major Western lab.
Amazon
Dec 2025
Amazon Nova 2 Pro
Amazon's first reasoning-capable Nova flagship, adding adjustable thinking intensity, built-in code interpreter/web grounding, and a 1M-token context on Bedrock.
Mistral AI
Dec 2025
Ministral 3 14B
Largest of the Ministral 3 small-model family (3B/8B/14B), delivering Mistral Small 3.2 24B-class performance while fitting in 24GB VRAM for local deployment.
Z.ai
Dec 2025
GLM-4.6V
Open vision-language model with native multimodal function calling, letting images and screenshots pass directly as tool parameters for multimodal agents.
Mistral AI
Dec 2025
Devstral 2
Open-weight agentic coding model launched with Mistral Vibe CLI, claiming near-frontier SWE-bench performance at a fraction of Claude Sonnet's cost.
OpenAI
Dec 2025
GPT-5.2
Rushed out under an internal 'Code Red' after Gemini 3 Pro's launch, adding Instant/Thinking/Pro tiers and a dedicated GPT-5.2-Codex variant.
Google
Dec 2025
Gemini 3 Flash
Became the new default model in the Gemini app, bringing Gemini 3-generation reasoning to a low-latency, low-cost workhorse tier.
Z.ai
Dec 2025
GLM-4.7
Ranked the top open-weight model across SWE-bench, tau2-bench and LiveCodeBench, ahead of DeepSeek-V3.2, cementing Z.ai's coding-agent focus.
MiniMax AI
Dec 2025
MiniMax-M2.1
Follow-up to M2 with large multilingual coding gains, approaching Claude Opus 4.5 on real-world dev tasks while keeping the 10B-active MoE efficiency.
Naver
Dec 2025
HyperCLOVA X SEED Think 32B
One of the strongest South Korean open-weight reasoning models, pairing a unified vision-language backbone with a reasoning-centric training recipe, outperforming EXAONE 4.0 32B.
TII
Jan 2026
Falcon-H1R-7B
A 7B hybrid Transformer/Mamba2 reasoning model that matched or beat models 2-7x its size, showing hybrid state-space architectures scale down efficiently for reasoning.
Liquid AI
Jan 2026
LFM2.5
Launched Liquid AI's next-generation on-device model family, later expanding into vision, audio, encoder, and agent-focused variants across 2026.
Alibaba
Jan 2026
Qwen3-Max-Thinking
Alibaba's flagship reasoning model, trained on 36T tokens with adaptive tool invocation, matching GPT-5.2-Thinking, Claude Opus 4.5 and Gemini 3 Pro on 19 benchmarks.
Moonshot AI
Jan 2026
Kimi K2.5
Introduced 'Agent Swarm', splitting complex tasks across up to 100 orchestrated sub-agents per prompt, cutting execution time 4.5x versus single-agent runs.
Alibaba
Feb 2026
Qwen3-Coder-Next
Small hybrid MoE coding model trained at scale on executable task synthesis and RL, matching models with 10-20x more active parameters for local agentic coding.
Anthropic
Feb 2026
Claude Opus 4.6
Improved planning, code review and large-codebase reliability, topping the Finance Agent benchmark and shipping agent-teams features for 'vibe working'.
Z.ai
Feb 2026
GLM-5
Z.ai's frontier flagship, released weeks after its Hong Kong IPO, reaching #1 open-weight model on Artificial Analysis and LMArena's Text Arena; trained entirely on Huawei Ascend chips with zero NVIDIA dependency.
ByteDance
Feb 2026
Seed 2.0 Pro
ByteDance's next-generation general-purpose agent model family (Pro/Lite/Mini plus a dedicated Code model), powering the Doubao assistant lineup.
Alibaba
Feb 2026
Qwen3.5-397B-A17B
Flagship open-weight launch of the Qwen3.5 family, outperforming the much larger trillion-parameter Qwen3-Max while supporting 201 languages and 1M-token context.
Alibaba
Feb 2026
Qwen3.5-Plus
Hosted flagship variant of Qwen3.5, cutting deployment memory 60% and boosting long-context throughput up to 19x versus the prior Qwen3 generation.
xAI
Feb 2026
Grok 4.2
Shipped as a public-beta 'rapid learning' model with weekly capability updates instead of a static release, a departure from xAI's prior monolithic training cadence.
Anthropic
Feb 2026
Claude Sonnet 4.6
Anthropic's new default model, adding a 1M-token context beta and 30-50% faster inference than Sonnet 4.5 while holding frontier coding and agent-planning scores.
Alibaba
Feb 2026
Qwen3.5-Flash
Production-grade low-cost Qwen3.5 tier combining linear attention with sparse MoE for higher inference efficiency, part of the broader Qwen3.5 Medium series launch.
OpenAI
Mar 2026
GPT-5.3 Instant
Became ChatGPT's new default everyday model, cutting hallucinations (26.8% with web, 19.7% offline) and reducing unnecessary refusals and caveats.
OpenAI
Mar 2026
GPT-5.4
First general-purpose OpenAI model with native, state-of-the-art computer-use built into Codex and the API.
Meta
Apr 2026
Muse Spark
First model from newly formed Meta Superintelligence Labs, Meta's flagship reset after Llama, positioned for multimodal reasoning and coding.
Alibaba
Apr 2026
Qwen3.6-Plus
Flagship agentic-coding update to Qwen3, adding native front-end generation from screenshots/design drafts and agentic benchmarks near Claude Opus 4.5.
Z.ai
Apr 2026
GLM-5.1
Open-weight coding model able to review and fundamentally change its own strategy across hundreds of iterations rather than getting stuck; edged out GPT-5.4 and Claude Opus 4.6 on SWE-Bench Pro.
Anthropic
Apr 2026
Claude Opus 4.7
Notable gains in advanced software engineering and self-verification, plus higher-resolution vision; Anthropic conceded it trails its unreleased 'Mythos' tier model.
Moonshot AI
Apr 2026
Kimi K2.6
Introduced a 300-agent Agent Swarm system executing up to 4,000 coordinated steps in a single autonomous run, claimed to outperform GPT-5.4 and Claude Opus 4.6.
InclusionAI
Apr 2026
Ling-2.6-flash
Ant Group's efficiency-focused agent model using roughly one-tenth the token usage of Nemotron 3 Super for comparable agentic tasks.
OpenAI
Apr 2026
GPT-5.5
Shipped just six weeks after GPT-5.4, an unusually fast cadence, with better tool use, self-checking, and task persistence.
DeepSeek
Apr 2026
DeepSeek V4 Pro
DeepSeek's next flagship after V3.2, shipped open-weight with low/high/max thinking-effort levels and performance rivaling closed frontier labs.
Mistral AI
Apr 2026
Mistral Medium 3.5
Merged Mistral's three separate model lines (Medium, Magistral reasoning, Devstral coding) into a single dense open-weight model with configurable reasoning effort.
IBM
Apr 2026
Granite 4.1
IBM's small-model refresh where the dense 8B instruct model matched the prior 32B Granite 4.0 MoE, showing continued gains from architecture over scale.
xAI
Apr 2026
Grok 4.3
Added native video input and document generation, with reasoning made a permanent always-on state for every query rather than a toggle.
Google
May 2026
Gemini 3.1 Flash-Lite
Google's fastest, cheapest Gemini 3 tier reached GA, built for high-volume agentic tool-calling and orchestration pipelines.
Baidu
May 2026
ERNIE 5.1
Reached leading performance at roughly 6% of the pre-training compute cost of comparable models by reusing ERNIE 5.0's elastic sub-model matrix instead of training from scratch.
Cursor
May 2026
Composer 2.5
Cursor's in-house agentic coding model, tuned with 25x more synthetic training tasks and textual-feedback RL, running 10-60x cheaper per task than Opus 4.7 or GPT-5.5.
Google
May 2026
Gemini 3.5 Flash
First release of Google's Gemini 3.5 family, combining frontier intelligence with 'real-world action'; strongest Gemini yet for agents and coding at launch.
Cohere
May 2026
Command A+
Cohere's first open-weight Command release, runnable on as few as two H100 GPUs, more than doubling official-language coverage to 48 languages.
Alibaba
May 2026
Qwen3.7-Max
Alibaba's proprietary flagship claimed to run autonomously for 35 hours without degradation, topping Claude Opus 4.6 Max on Terminal-Bench 2.0 and SWE-Bench Pro.
Anthropic
May 2026
Claude Opus 4.8
Shipped just 41 days after Opus 4.7, the shortest Opus point-release gap to date, emphasizing improved honesty (avoiding unsupported claims).
MiniMax AI
Jun 2026
MiniMax M3
First open-weight model to combine frontier coding performance, 1M-token context, and native multimodality in one release.
Anthropic
Jun 2026
Claude Fable 5
Anthropic's first publicly released 'Mythos-class' model, state of the art on nearly all tested benchmarks at launch across coding, knowledge work and vision.
Moonshot AI
Jun 2026
Kimi-K2.7-Code
Coding-specialized checkpoint in Moonshot's Kimi K2.x line, later used as the base for Cursor's Composer 2 agentic coding model.
NVIDIA
Jun 2026
NVIDIA Nemotron 3 Ultra
A fully open (weights, data, recipes) hybrid Mamba-Transformer reasoning model purpose-built for million-token long-running agents.
NVIDIA
Jun 2026
NVIDIA Nemotron 3 Super
Mid-size sibling in NVIDIA's open Nemotron 3 hybrid Mamba-Transformer family released alongside Nemotron 3 Ultra.
Microsoft
Jun 2026
MAI-Thinking-1
Microsoft's first flagship in-house reasoning model built from scratch, unveiled at Build 2026, preferred by human raters over Claude Sonnet 4.6 in blind comparisons.
Z.ai
Jun 2026
GLM-5.2
Cut per-token FLOPs 2.9x at 1M context via IndexShare attention; widely regarded as the most capable open-weight model at release, on par with GPT-5.2.
Weibo
Jun 2026
VibeThinker-3B
A 3B open-weight model claiming reasoning parity with 671B-1T models like DeepSeek V3.2 and Kimi K2.5 on math/code benchmarks, reigniting debate over small-model post-training gains versus scale.
Anthropic
Jun 2026
Claude Sonnet 5
Replaced Sonnet 4.6 as the default free/Pro model the day it launched, targeting Opus-adjacent performance with major gains in agentic planning and tool use at lower cost.
xAI
Jul 2026
Grok 4.5
First xAI release after acquiring Cursor, built on a 1.5T-parameter V9 foundation with Cursor coding data mixed into training; Musk called it 'Opus-class but faster.'
OpenAI
Jul 2026
GPT-5.6 (Sol / Terra / Luna)
Restructured OpenAI's coding lineup into durable capability tiers (Sol/Terra/Luna) and merged ChatGPT and Codex into one desktop app; Sol claimed 54% better token efficiency on coding tasks.
OpenAI
Jul 2026
GPT-5.6 Sol
Flagship of OpenAI's first three-tier model family (Sol/Terra/Luna), each tier advancing on its own release cadence; Sol outscored Claude Fable 5 on Agents' Last Exam.
OpenAI
Jul 2026
GPT-5.6 Terra / Luna
Balanced (Terra) and fastest/cheapest (Luna) tiers of OpenAI's first durable-cadence model family, later repriced (-20%/-80%) after Sol autonomously optimized inference infrastructure.
KwaiKAT
Jul 2026
KAT-Coder-Pro V2.5
Kuaishou's KwaiKAT team shipped China's first agentic coding model claimed capable of running full software-engineering projects end-to-end, beating Claude Opus 4.8 on PinchBench.
Moonshot AI
Jul 2026
Kimi K3
Largest open-weight model publicly released to date; Moonshot claimed performance competitive with Claude Fable 5 and ahead of Opus 4.8, driving 6x revenue growth for the company.
Google
Jul 2026
Gemini 3.6 Flash
Google's 'workhorse' model, using ~17% fewer output tokens than 3.5 Flash while improving on coding, long-context and computer-use benchmarks; Google teased Gemini 4 alongside it.
Anthropic
Jul 2026
Claude Opus 5
Positioned as Anthropic's everyday enterprise model approaching Fable 5-level intelligence at half the price, with a low/medium/high reasoning-effort toggle; Anthropic's fourth Claude 5 release in under two months.
DeepSeek
Jul 2026
DeepSeek V4 Flash
Fast, economical open-weight sibling to V4 Pro, giving developers a cheaper on-ramp to DeepSeek's fourth-generation architecture.
InclusionAI
Aug 2026
Ling 3.0 Flash
Activated only 5.1B parameters per token while matching or surpassing InclusionAI's previous trillion-parameter flagship on key reasoning and agent benchmarks.
Alibaba
Aug 2026
Qwen3.8-Max / 2.4T-A95B
Brought a Qwen Max-class checkpoint to open release for the first time, exposing 2.4T total parameters and 95B active parameters for long-horizon coding and agentic work.
Liquid AI
Aug 2026
LFM2.5-2.6B
Fully on-device agentic model that plans and calls tools with under 2.5GB memory, reported at 220 tok/s on Apple M5 Max, extending agentic capability to edge hardware.
Alibaba
Aug 2026
Qwen3.8-27B
Packed Qwen3.8's multimodal understanding, adjustable reasoning effort and long-horizon agent capabilities into a deployment-friendly dense 27B model.
InclusionAI
Aug 2026
Ling 3.0 Tiny
Delivered hybrid reasoning and agentic capabilities with only 1.3B active parameters, running in about 8.34 GiB and reaching 86-90 tok/s on an M4 Pro MacBook.
xAI
Aug 2026
Grok 4.6
Built for long-running agentic and coding work, matched GPT-5.6 Sol on the Artificial Analysis Intelligence Index and led on GDPVal-AA v2 at launch.
Liquid AI
Aug 2026
LFM2.5-VL-3B
On-device vision-language model with screen understanding, object grounding, and function calling, matching InternVL-3.5-4B despite running on phones and laptops.
DeepSeek
Aug 2026
DeepSeek V4 Pro 0813
Superseded the V4 Pro preview with sharply improved production agent performance and integrated DSpark speculative decoding for faster inference.