Agentic coding and MoE models lead the frontier. DeepSeek V4 Pro, MiniMax M3, and the shifting open-weight landscape.

LLMWatch

Issue #1 · June 4, 2026

Weekly Briefing

Agentic Coding Leads the Shift

The strongest open-weight releases now optimize for coding agents, tool-use, and terminal workflows rather than pure benchmark breadth. Sparse MoE architecture remains the dominant route to frontier capabilities, while 1M-token context has normalized as the new baseline.

By the Numbers

87

BenchLM
open leader

80.6%

SWE-bench
Verified

1M

Baseline
Context

$0.18

Cheapest
per M tok

Top Releases

MiniMax M3

Jun 1 · Weights Pending · 1M context

BREAKOUT

Frontier coding/agentic performance with native multimodal capabilities (image and video inputs). Native 1M-token context window.

DeepSeek V4 Pro (Max)

Apr · MIT · 1.6T params / 49B active · MoE

The open-weight leaderboard champion. Delivers a BenchLM score of 87 and SWE-bench Verified score of 80.6% with native 1M-token context.

Kimi K2.6

Apr/May · Modified MIT · ~1T params / ~32B active · MoE

High-throughput open-weight model with a BenchLM score of 85. Designed for RAG and coding tasks with excellent speed ergonomics.

GLM-5.1 (Reasoning)

May · MIT · 744B params / 40B active · MoE

Strong reasoning-focused open model scoring 83 on BenchLM. Integrates a 203K-token context window optimized for complex math and logic.

Gemma 4 12B

Jun 3 · Apache 2.0 · 12B dense

Laptop-friendly local multimodal flagship from Google. Integrates native audio and image inputs directly into the LLM backbone without external encoders.

Qwen3.7 Max / Plus

Jun · Open-weight · 1M context

Alibaba's new generation. Max represents the current open-source quality leader on WhatLLM, paired with a solid mid-tier Plus variant.

Key Trends

Agentic coding wins

The strongest releases now optimize for coding agents, browser-use, tool-use, and terminal workflows rather than pure benchmark breadth.

MoE stays dominant

Sparse and hybrid MoE designs remain the best route to frontier capability without explosive serving costs (e.g. DeepSeek V4 Pro, Kimi K2.6).

Chinese labs lead pace

DeepSeek, Qwen, GLM/Z.AI, MiniMax, and Xiaomi continue to set much of the open-weight release tempo.

Licensing drives adoption

Apache 2.0 and MIT remain the most deployment-friendly licenses. Buyers are increasingly selecting models by legal fit as much as score.

Speed is a weapon

High-throughput models like MiniMax-M2.7 and Kimi K2.6 show that latency and tokens/sec are now first-class product metrics.

Benchmark Snapshot

SWE-bench Verified (closed) Claude Opus 4.8 — ~82%
SWE-bench Verified (open) DeepSeek V4 Pro — 80.6%
SWE-bench Pro MiniMax M3 — 59.0%
LiveCodeBench (open) DeepSeek V4 Pro — 93.5
WideSearch Kimi K2.6 — 80.8%
BenchLM Open-Weight Overall DeepSeek V4 Pro (Max) — 87
Cheapest Input (per 1M tokens) MiMo-V2.5 — $0.18
Largest Context (closed) Grok 4 Fast — 2.0M

Best Picks This Week

Frontier Coding (self-hosted)

DeepSeek V4 Pro → Kimi K2.6

DeepSeek is the safest benchmark anchor for self-hosting; Kimi K2.6 brings strong throughput and a very competitive composite score.

Cost-Efficient Production

MiMo-V2.5 → DeepSeek V4 Flash

MiMo-V2.5 is the cheapest high-ranking open option ($0.18/M); DeepSeek V4 Flash offers a stronger benchmark anchor for slightly higher budgets.

Single-GPU Coding

Devstral Small 2 → Mellum2

Devstral Small 2 (68.0% SWE-Verified) targets local consumer hardware; Mellum2 is optimized for low-latency coding workflows.

Computer Use Agents

MiniMax M3 → Step 3.7 Flash

MiniMax M3 is the standout open-weight candidate for agentic tasks; Step 3.7 Flash is the practical fallback with strong tool benchmarks.

Closed-Source Context

Proprietary models still maintain an edge in general agent reasoning. Recommended routing:

Claude Opus 4.8 Frontier coding — SWE-bench Verified ~82%
Gemini 3.5 Flash Fast coding/multimodal — 78% SWE-bench Verified
GPT-5.5 Broad proprietary frontier baseline

What to Watch

MiniMax M3 Weights

Announced on June 1. Check HuggingFace for weights and final license terms.

Qwen Open-Weight Cadence

Qwen3.7 Max & Plus are live; look out for new open-weight coding variants.

DeepSeek V4 Family Stabilization

V4 Pro and Flash aim to establish a new price/performance ceiling for open-weight models.

Mistral's Coding Cadence

Devstral 2 and Small 2 continue Mistral's strong orientation toward developer-first self-hosting.

Explore the full dashboard

Interactive rankings, 31 open-weight models, filters, comparison tool, and use-case guide.

Open Dashboard →

LLMWatch

Weekly open-weight intelligence tracker

You're receiving this because you subscribed to LLM Watch.
Unsubscribe · View in browser