LLMWatch
Issue #1 · June 4, 2026
|
|
Weekly Briefing
Agentic Coding Leads the Shift
The strongest open-weight releases now optimize for coding agents, tool-use, and terminal workflows rather than pure benchmark breadth. Sparse MoE architecture remains the dominant route to frontier capabilities, while 1M-token context has normalized as the new baseline.
|
|
|
By the Numbers
|
87
BenchLM open leader
|
80.6%
SWE-bench Verified
|
1M
Baseline Context
|
$0.18
Cheapest per M tok
|
|
|
|
Top Releases
|
MiniMax M3
Jun 1 · Weights Pending · 1M context
|
BREAKOUT
|
Frontier coding/agentic performance with native multimodal capabilities (image and video inputs). Native 1M-token context window.
|
|
DeepSeek V4 Pro (Max)
Apr · MIT · 1.6T params / 49B active · MoE
|
The open-weight leaderboard champion. Delivers a BenchLM score of 87 and SWE-bench Verified score of 80.6% with native 1M-token context.
|
|
Kimi K2.6
Apr/May · Modified MIT · ~1T params / ~32B active · MoE
|
High-throughput open-weight model with a BenchLM score of 85. Designed for RAG and coding tasks with excellent speed ergonomics.
|
|
GLM-5.1 (Reasoning)
May · MIT · 744B params / 40B active · MoE
|
Strong reasoning-focused open model scoring 83 on BenchLM. Integrates a 203K-token context window optimized for complex math and logic.
|
|
Gemma 4 12B
Jun 3 · Apache 2.0 · 12B dense
|
Laptop-friendly local multimodal flagship from Google. Integrates native audio and image inputs directly into the LLM backbone without external encoders.
|
|
Qwen3.7 Max / Plus
Jun · Open-weight · 1M context
|
Alibaba's new generation. Max represents the current open-source quality leader on WhatLLM, paired with a solid mid-tier Plus variant.
|
|
|
|
Key Trends
|
Agentic coding wins
The strongest releases now optimize for coding agents, browser-use, tool-use, and terminal workflows rather than pure benchmark breadth.
|
|
MoE stays dominant
Sparse and hybrid MoE designs remain the best route to frontier capability without explosive serving costs (e.g. DeepSeek V4 Pro, Kimi K2.6).
|
|
Chinese labs lead pace
DeepSeek, Qwen, GLM/Z.AI, MiniMax, and Xiaomi continue to set much of the open-weight release tempo.
|
|
Licensing drives adoption
Apache 2.0 and MIT remain the most deployment-friendly licenses. Buyers are increasingly selecting models by legal fit as much as score.
|
|
Speed is a weapon
High-throughput models like MiniMax-M2.7 and Kimi K2.6 show that latency and tokens/sec are now first-class product metrics.
|
|
|
|
Benchmark Snapshot
| SWE-bench Verified (closed) |
Claude Opus 4.8 — ~82% |
| SWE-bench Verified (open) |
DeepSeek V4 Pro — 80.6% |
| SWE-bench Pro |
MiniMax M3 — 59.0% |
| LiveCodeBench (open) |
DeepSeek V4 Pro — 93.5 |
| WideSearch |
Kimi K2.6 — 80.8% |
| BenchLM Open-Weight Overall |
DeepSeek V4 Pro (Max) — 87 |
| Cheapest Input (per 1M tokens) |
MiMo-V2.5 — $0.18 |
| Largest Context (closed) |
Grok 4 Fast — 2.0M |
|
|
|
Best Picks This Week
|
Frontier Coding (self-hosted)
DeepSeek V4 Pro → Kimi K2.6
DeepSeek is the safest benchmark anchor for self-hosting; Kimi K2.6 brings strong throughput and a very competitive composite score.
|
|
|
Cost-Efficient Production
MiMo-V2.5 → DeepSeek V4 Flash
MiMo-V2.5 is the cheapest high-ranking open option ($0.18/M); DeepSeek V4 Flash offers a stronger benchmark anchor for slightly higher budgets.
|
|
|
Single-GPU Coding
Devstral Small 2 → Mellum2
Devstral Small 2 (68.0% SWE-Verified) targets local consumer hardware; Mellum2 is optimized for low-latency coding workflows.
|
|
|
Computer Use Agents
MiniMax M3 → Step 3.7 Flash
MiniMax M3 is the standout open-weight candidate for agentic tasks; Step 3.7 Flash is the practical fallback with strong tool benchmarks.
|
|
|
|
Closed-Source Context
Proprietary models still maintain an edge in general agent reasoning. Recommended routing:
| Claude Opus 4.8 |
Frontier coding — SWE-bench Verified ~82% |
| Gemini 3.5 Flash |
Fast coding/multimodal — 78% SWE-bench Verified |
| GPT-5.5 |
Broad proprietary frontier baseline |
|
|
|
What to Watch
|
MiniMax M3 Weights
Announced on June 1. Check HuggingFace for weights and final license terms.
|
|
Qwen Open-Weight Cadence
Qwen3.7 Max & Plus are live; look out for new open-weight coding variants.
|
|
DeepSeek V4 Family Stabilization
V4 Pro and Flash aim to establish a new price/performance ceiling for open-weight models.
|
|
Mistral's Coding Cadence
Devstral 2 and Small 2 continue Mistral's strong orientation toward developer-first self-hosting.
|
|
|
|
Explore the full dashboard
Interactive rankings, 31 open-weight models, filters, comparison tool, and use-case guide.
Open Dashboard →
|
|
|
LLMWatch
Weekly open-weight intelligence tracker
You're receiving this because you subscribed to LLM Watch.
Unsubscribe · View in browser
|
|