LLMWatch
Issue #3 · June 18, 2026
|
|
Weekly Briefing
GLM-5.2 Seizes the Open-Weight Frontier
GLM-5.2 (Z.ai, AA Intelligence Index v4.1: 51) is the new open-weight leader — 7 points ahead of DeepSeek V4 Pro and MiniMax M3 (tied at 44). It launched one day after the US Commerce Department forced Anthropic to globally disable Claude Fable 5 / Mythos 5, the first-ever government-ordered takedown of a deployed frontier model.
|
|
|
Key Takeaways
|
1. The US government forced Claude Fable 5 offline on June 12 — the first-ever takedown of a deployed frontier model.
|
|
2. GLM-5.2 (Z.ai, MIT, 1M context) is the new open-weight leader at AA v4.1: 51, leading the field by 7 points.
|
|
3. The open-to-closed gap narrowed from ~10 points to ~5 points (Opus 4.8 56 vs. GLM-5.2 51).
|
|
4. The new AA v4.1 index reweights toward agentic tasks; DeepSeek V4 Pro leads on cost at $0.04 per task.
|
|
5. Three open models now operate at or near 1 trillion parameters — Kimi K2.7 Code, InclusionAI Ling/Ring 2.6, and DeepSeek V4 Pro.
|
|
|
|
What Changed Since Last Issue
June 13 → June 17, 2026
New This Issue
| GLM-5.2 |
Z.ai · MIT |
AA v4.1: 51 (#1 open) |
| Kimi K2.7 Code |
Moonshot · Modified MIT |
~1T/32B MoE |
| InclusionAI Ling/Ring 2.6 |
Ant Group · MIT |
Two 1T checkpoints |
| Olmo Hybrid |
AI2 · Apache 2.0 |
2× data efficiency |
| Nemotron-Labs Diffusion |
NVIDIA · NVIDIA Open Model License |
First diffusion LM family |
| Grok 4 Open 100B-A20B |
xAI · xAI Custom License |
xAI's first open weights |
Score & Ranking Moves
| AA Index (open leader) |
MiniMax M3 55 → GLM-5.2 51 ↓ v4.0→v4.1 |
| AA Index (MiniMax M3) |
55 → 44 ↓ recalibration |
| BenchLM Open Overall (DeepSeek V4 Pro) |
87 → 86 ↓ |
| SWE-bench Verified (Opus 4.8) |
~82% → 88.6% ↑ |
| BenchLM M3 rank (MiniMax M3) |
#29/119 → #23/124 ↑ |
| Open-to-closed gap |
~10 pts → ~5 pts ↑ |
Dropped
LightningLM, Macaron-V1-Preview, Quasar-Preview, Gemma 4 12B, Magistral Small, GLM-5.1, Nemotron 3 Super — superseded or not deployable.
|
|
|
By the Numbers
|
51
AA v4.1 open leader
|
86
BenchLM open anchor
|
$0.04
Cost-per-Task cheapest open
|
1M
Baseline Context
|
|
|
|
The Bigger Picture
The LLM market shifted on two fronts this week. The closed ceiling dropped sharply. Claude Fable 5 was the overall #1 model on the AA Intelligence Index. Then the US Commerce Department ordered it disabled globally on June 12. All customers lost access, not just foreign nationals.
At the same time, the open frontier rose. GLM-5.2 from Z.ai scored 51 on the new AA v4.1 index. That makes it the open-weight leader, 7 points ahead of DeepSeek V4 Pro and MiniMax M3. It also took the #1 open agentic spot on WhatLLM. The open-to-closed gap is now just 5 points.
Architecture is diversifying fast. Every major open release still uses sparse MoE design. But hybrid and diffusion models are now joining the field. Olmo Hybrid doubles data efficiency. NVIDIA and Google both shipped diffusion language models. Licensing is also splitting: MIT and Apache 2.0 strengthened, even as some infrastructure went closed.
|
Core Signal This Week
Three forces converged this week. The closed frontier was capped by government action (Fable 5 offline). The open frontier surged upward (GLM-5.2 at AA 51). And the methodology shifted toward agentic evaluation (AA v4.1). For the first time, the best open and closed models are within 5 composite points.
|
|
|
|
Top Releases
|
Kimi K2.7 Code
Jun 12 · Modified MIT · ~1T/32B MoE · 256K context
|
BREAKOUT
|
Kimi K2.7 Code posts +21.8% on Kimi Code Bench v2 over K2.6. It beats Opus 4.8 on tool invocation (MCP Mark Verified 81.1 vs 76.4). All benchmarks are vendor-reported with zero independent confirmation.
|
|
GLM-5.2
Jun 13 · MIT · 744B/40B active · MoE · 1M context
|
GLM-5.2 is the new open-weight leader at AA v4.1: 51. It also took #1 open-weight agentic on WhatLLM (75.9). MIT license and 1M context make it a strong long-context coding alternative.
|
|
InclusionAI Ling-2.6 / Ring-2.6
Jun 13 · MIT · ~1T/63B active (Ling/Ring-1T) · 256K–262K context
|
Ant Group's InclusionAI shipped three trillion-scale checkpoints under MIT. The accompanying arXiv paper focuses on instant agentic intelligence at trillion-parameter scale.
|
|
Olmo Hybrid
Jun 16 (week of) · Apache 2.0 · 7B · Hybrid transformer + linear recurrent
|
Olmo Hybrid pairs gated DeltaNet with attention in a 3:1 pattern. It matches Olmo 3 accuracy with 49% fewer training tokens. Weights, data, code, and checkpoints are all released.
|
|
Nemotron-Labs Diffusion
Jun 16 · NVIDIA Nemotron Open Model License (text) · 3B / 8B / 14B
|
NVIDIA's first diffusion language model family. It supports AR, diffusion, and self-speculation modes. Training code is released via NVIDIA Megatron Bridge.
|
|
DiffusionGemma
Jun 10 · Apache 2.0 · 26B/3.8B active · MoE with diffusion
|
DiffusionGemma hits 1,000+ tok/s on H100 via non-autoregressive diffusion. It is experimental and lower quality than Gemma 4, but up to 4× faster. Useful for code infilling, math graphs, and Sudoku-style tasks.
|
|
|
|
Key Trends
|
Fable 5 forced offline
The US Commerce Department ordered Anthropic to disable Claude Fable 5 / Mythos 5 on June 12. It was the first-ever government takedown of a deployed frontier model. The effective closed leader is now Opus 4.8 (AA v4.1: 56).
|
|
GLM-5.2 takes open crown
GLM-5.2 (Z.ai, MIT, 1M context) scored 51 on AA v4.1 — the open-weight leader by 7 points. It launched one day after the Fable 5 shutdown, positioned as a long-context coding alternative.
|
|
AA v4.1 agentic shift
Artificial Analysis reweighted its index toward agentic workloads. The new Cost-per-Task metric quantifies the open-value story: DeepSeek V4 Pro at $0.04/task vs. Opus 4.8 at $1.78/task.
|
|
Trillion-scale open flood
Three open releases now operate at or near 1T parameters: Kimi K2.7 Code, InclusionAI Ling/Ring 2.6, and DeepSeek V4 Pro. Combined with GLM-5.2 and MiniMax M3, the open frontier is in the trillion-parameter regime.
|
|
MIT license strengthens
GLM-5.2, InclusionAI Ling/Ring 2.6, Olmo Hybrid, DeepSeek V4 Pro/Flash, and MiMo-V2.5-Pro all ship under fully permissive MIT. This shifts the frontier toward unrestricted commercial use.
|
|
Open infra trust erodes
Google will replace its Apache-2.0 Gemini CLI with a closed-source Antigravity binary on June 18. Separately, Meta's Alexandr Wang said the old open-source playbook "didn't work." Muse Spark stays closed.
|
|
DeepSeek V4 value dominance
DeepSeek V4 Pro combines the highest independently confirmed open composite (BenchLM 86) with the lowest Cost-per-Task ($0.04). It is the standout value story even though GLM-5.2 now leads the AA composite.
|
|
Hybrid & diffusion emerge
Olmo Hybrid pairs gated DeltaNet with attention for 2× data efficiency. NVIDIA's Nemotron-Labs Diffusion family and Google's DiffusionGemma bring non-autoregressive text generation into the open ecosystem.
|
|
|
|
Benchmark Snapshot
| AA Intelligence Index v4.0 |
Claude Fable 5 — 64.9 [OFFLINE] |
| AA Intelligence Index v4.1 |
Claude Fable 5 — 60 [OFFLINE] |
| AA Intelligence Index v4.1 (available) |
Claude Opus 4.8 — 56 |
| AA Intelligence Index v4.1 (open) |
GLM-5.2 — 51 |
| AA Intelligence Index v4.1 (open, 2nd) |
DeepSeek V4 Pro / MiniMax M3 — 44 (tied) |
| WhatLLM Agentic Index (overall) |
Claude Fable 5 — 80.6 [OFFLINE] |
| WhatLLM Agentic Index (open) |
GLM-5.2 — 75.9 |
| BenchLM Coding (overall) |
Claude Opus 4.8 — 76.4 |
| BenchLM Coding (open) |
DeepSeek V4 Pro (Max) — 75.9 |
| BenchLM Overall (open, confirmed) |
DeepSeek V4 Pro (Max) — 86 |
| BenchLM Overall (open, 2nd) |
DeepSeek V4 Pro (High) — 82 |
| BenchLM Overall (open, 3rd) |
DeepSeek V4 Flash (Max) — 74 |
| SWE-bench Verified (closed) |
Claude Opus 4.8 — 88.6% |
| SWE-bench Verified (open, confirmed) |
DeepSeek V4 Pro — ~88% (favoured) |
| SWE-bench Verified (open, Qwen 4) |
Qwen 4 — 78% |
| SWE-bench Verified (open coder, best) |
Qwen 4 Coder 32B-A3B — 82% |
| GDPval-AA v2 Elo (v4.1) |
Claude Fable 5 — 1818 [OFFLINE] |
| GDPval-AA (open, WhatLLM) |
GLM-5.2 — 1524 |
| HLE |
Claude Fable 5 — 53% [OFFLINE] |
| Kimi Code Bench v2 |
Kimi K2.7 Code — 62.0 [vendor] |
| LLMCheck Score (open) |
Qwen 4 — 75 |
| RULER @1M |
Nemotron 3 Ultra — 95% |
| IOI 2025 |
Nemotron 3 Ultra — 570.0 |
| LiveCodeBench v6 |
Nemotron 3 Ultra — 89.0 |
| AIME 2024 (open) |
Magistral Small — 70.7% |
| BenchLM M3 (provisional) |
MiniMax M3 — #23 of 124 |
| Cost-per-Task (cheapest) |
DeepSeek V4 Pro (max) — $0.04/task |
| Largest Context (open) |
Llama 4 Scout — 10M |
| Largest Context (closed) |
Grok 4 Fast — 2.0M |
| Cheapest Input (open) |
MiMo-V2.5 — $0.18/M |
|
|
|
Price-Quality Rankings
Source: WhatLLM Agentic Index — top models ranked by agentic quality score
| # |
Model |
Quality |
Price/M |
Speed |
Ctx |
| 1 |
Claude Fable 5 [OFFLINE] |
80.6 |
$10/$50 |
— |
N/A |
| 2 |
Claude Opus 4.8 |
77.8 |
$5/$25 |
— |
— |
| 3 |
GLM-5.2 (max) |
75.9 |
$1.40/$4.40 |
— |
1M |
| 4 |
GPT-5.5 (xhigh) |
74.1 |
$5/$30 |
— |
— |
| 5 |
Claude Opus 4.7 |
71.3 |
— |
— |
— |
| 6 |
Gemini 3.5 Flash |
70.3 |
$1.50/$9 |
— |
— |
| 7 |
MiniMax-M3 |
68.6 |
$0.30/$1.20 |
— |
1M |
Speed values marked "—" are single-source (WhatLLM) and could not be re-verified from the live JS-rendered page. The Agentic Index has only 7 ranked models — all rendered; top-10 requested but only 7 exist in source.
|
|
|
Best Picks by Use Case
|
Frontier Coding (API)
Main: GLM-5.2
Runner-up: DeepSeek V4 Pro
GLM-5.2 leads open-weights on AA v4.1 (51) and WhatLLM Agentic (75.9), with MIT license, 1M context, and $1.40/$4.40 per M. DeepSeek V4 Pro is the strongest independently confirmed coder (BenchLM Coding 75.9, SWE-Verified ~88%).
|
|
|
Frontier Coding (self-hosted)
Main: DeepSeek V4 Pro
Runner-up: GLM-5.2
DeepSeek V4 Pro is the safest benchmark anchor under MIT (BenchLM 86, SWE-Verified ~88%). GLM-5.2 adds 1M context and AA v4.1 leadership but requires ~744B parameters to host.
|
|
|
Cost-Efficient Production
Main: DeepSeek V4 Flash
Runner-up: MiMo-V2.5
DeepSeek V4 Flash offers $0.14/$0.28 per M with BenchLM 74 and 1M context — the best value-to-quality ratio. MiMo-V2.5 is cheaper ($0.18/M blended) with 210 tok/s and 1M context.
|
|
|
Single-GPU Coding
Main: Devstral Small 2
Runner-up: Qwen 4 Coder 32B-A3B
Devstral Small 2 (24B dense, Apache 2.0) is purpose-built for consumer-hardware coding at 68.0% SWE-bench Verified. Qwen 4 Coder (32B/3B active, Apache 2.0) hits 82% SWE-Verified as the best open-source Mac coder.
|
|
|
Edge / Small Devices
Main: LFM2.5-8B-A1B
Runner-up: Gemma 4.5 12B
LFM2.5-8B-A1B (8.3B/1.5B active) is explicitly designed for edge deployment (CPU/NPU/GPU) with τ²-Telecom 88.07 and IFEval 91.84. Gemma 4.5 12B (Apache 2.0, 1M context) is the best multimodal small-model alternative.
|
|
|
RAG / Routing / Sub-Agents
Main: Mellum2
Runner-up: Command A+
Mellum2 (12B/2.5B active, Apache 2.0) is explicitly optimized for routing and sub-agent workflows with low latency. Command A+ adds enterprise-grade multilingual and multimodal support (Apache 2.0).
|
|
|
Computer Use Agents
Main: GLM-5.2
Runner-up: Kimi K2.6
GLM-5.2 is the #1 open-weight agentic model (WhatLLM Agentic 75.9, GDPval-AA 1524) with MIT license and 1M context. Kimi K2.6 adds 300 parallel sub-agents and BrowseComp 86.3% under Modified MIT.
|
|
|
Long Context (10M+)
Main: Llama 4 Scout
Runner-up: GLM-5.2 / DeepSeek V4 Pro
Llama 4 Scout is the open-weight long-context landmark at 10M tokens. For 1M-context use cases, GLM-5.2, DeepSeek V4 Pro, and Nemotron 3 Ultra all offer it.
|
|
|
Reasoning / Math
Main: GLM-5.2
Runner-up: Nemotron 3 Ultra
GLM-5.2 (AA v4.1: 51, MIT) is the strongest open composite including reasoning. Nemotron 3 Ultra (GPQA 87.0%, IOI 2025 570.0, OpenMDW-1.1) is the specialist for math/code competition.
|
|
|
Multilingual
Main: Command A+
Runner-up: GLM-5.2
Command A+ (Apache 2.0) emphasizes multilingual enterprise use. GLM-5.2 (MIT) adds strong Chinese-English bilingual capabilities from Z.ai.
|
|
|
Physical AI / Robotics
Main: Cosmos 3
Runner-up: Qwen-Robot Suite
Cosmos 3 (NVIDIA, OpenMDW-1.1) is an omnimodal foundation model covering text/images/video/sound/actions. Qwen-Robot Suite (Alibaba) adds VLA, VLN, and video world model variants — pilot testing underway.
|
|
|
Speed-Optimized Generation
Main: DiffusionGemma
Runner-up: MiMo-V2.5-Pro-UltraSpeed
DiffusionGemma achieves 1,000+ tok/s on H100 via non-autoregressive diffusion under Apache 2.0 — but is experimental and lower quality than Gemma 4. MiMo UltraSpeed offers 1,000+ tok/s on a 1T-param model.
|
|
|
|
Licensing Landscape
| License |
Representative Models |
Status |
| Apache 2.0 |
DiffusionGemma, Olmo Hybrid, Qwen 4, Gemma 4.5 12B, Command A+, Mellum2, Snowflake Arctic |
Approved |
| MIT |
GLM-5.2, InclusionAI Ling/Ring 2.6, DeepSeek V4 Pro/Flash, MiMo-V2.5-Pro |
Approved |
| Modified MIT |
Kimi K2.7 Code, Kimi K2.6, Devstral 2, Mistral Small 4 |
Conditional |
| MiniMax Community License |
MiniMax M3 |
Conditional |
| xAI Custom License |
Grok 4 Open 100B-A20B |
Conditional |
| OpenMDW-1.1 |
Nemotron 3 Ultra, Cosmos 3 |
Conditional |
| License unconfirmed |
Nex-N2 Pro, Nex-N2-mini (4th cycle, no LICENSE file) |
Pending |
| License TBD |
Qwen-Robot Suite, FastContext-1.0-4B-SFT, Llama 5 70B |
Pending |
| Proprietary / API-only |
Claude Fable 5 [OFFLINE], Opus 4.8, GPT-5.5, Gemini 3.5 Flash, Muse Spark |
Restricted |
|
|
|
Closed-Source Context
Proprietary models set the reference ceiling, but the open gap has narrowed sharply this week. Key references:
| Claude Fable 5 / Mythos 5 |
Anthropic · Jun 9 → Jun 12 SHUTDOWN · AA 64.9 (v4.0) / 60 (v4.1) |
| Claude Opus 4.8 |
Anthropic · May 28 · AA v4.1: 56 (effective available leader) |
| GPT-5.5 |
OpenAI · Apr · AA v4.1: 55 · Cost-per-Task $0.99 |
| Qwen3.7-Plus |
Alibaba · ~Jun 14 · Proprietary · 1M context · BenchLM Coding 71.1 |
| Qwen3.7 Max |
Alibaba · May 19 · Proprietary · 1T+ params · AA v4.1: 46 |
| Gemini 3.1 Pro Preview |
Google · May · AA v4.1: 46 · Time-per-Task 1.6 min (2nd fastest) |
| Gemini 3.5 Flash |
Google · Jun · AA v4.0 ~55 · WhatLLM Agentic 70.3 |
| Muse Spark |
Meta · Jun 4 · API launch delayed · First Meta flagship without open weights |
| MAI-Thinking-1 |
Microsoft · Jun 2 · 35B active · 256K context · private preview |
| Unisound U2 |
Unisound · Jun 8 · Proprietary · GPQA 87.9 · SWE-Verified 75 |
| Magistral Medium |
Mistral · ~Jun 3 · Proprietary · AIME-24 73.6% |
|
|
|
What to Watch
|
Fable 5 Restoration Timeline (CRITICAL)
Anthropic engineers went to Washington on June 16. No restoration date announced. If it returns, the closed ceiling jumps back to AA v4.1: 60.
|
|
GLM-5.2 Independent Benchmark Confirmation
SWE-bench / Terminal-Bench / GPQA scores on independent leaderboards are still TBD. Confirmation would solidify GLM-5.2 as the definitive open leader.
|
|
Kimi K2.7 Code Independent Benchmarks
Every Kimi K2.7 Code benchmark is vendor-reported. Zero independent SWE-bench / LiveCodeBench / GPQA scores exist yet.
|
|
MiniMax M3 Independent Convergence
M3 vendor scores (SWE-Bench Pro 59%, GPQA 93%) still sit far above its BenchLM independent rank (#23/124). The starkest vendor-vs-independent gap.
|
|
Nex-N2 Pro License (5th Cycle)
Four consecutive sweeps found no LICENSE file. If confirmed Apache 2.0, it becomes a permissive agentic coding leader.
|
|
Gemini CLI → Antigravity Closure (Jun 18)
Free Gemini CLI access ends June 18. Watch for community forks and whether other labs follow "open code, closed infra."
|
|
Meta's Two-Tier Open/Closed Strategy
Wang said the old playbook "didn't work." Watch for Muse Spark's launch date and whether Llama 6 keeps open weights.
|
|
Microsoft FastContext License & Evaluation
FastContext-1.0-4B-SFT claims up to 60% token reduction for coding agents. License terms and independent benchmarks are pending.
|
|
AA v4.1 Methodology Settling
First cycle under the v4.1 agentic reweighting. Watch whether scores stabilize and whether Cost-per-Task gains adoption.
|
|
DeepSeek V4 Parameter Clarification
HF "862B"/"158B" labels are on-disk GB sizes (FP4+FP8), not parameter counts. A direct config.json read would settle it definitively.
|
|
|
|
Explore the full dashboard
Interactive rankings, 30+ open-weight models, filters, comparison tool, and use-case guide.
Open Dashboard →
|
|
|
LLMWatch
Weekly open-weight intelligence tracker
You're receiving this because you subscribed to LLM Watch.
Unsubscribe · View in browser
|
|
|