Browse past issues of LLM Watch. Open-weight releases, benchmark shifts, licensing changes, and deployment picks — delivered weekly.
The strongest open-weight releases now optimize for coding agents, tool-use, and terminal workflows rather than pure benchmark breadth. MoE remains dominant, 1M context is standard baseline.
GLM-5.2 (Z.ai) claims the open-weight frontier at AA v4.1: 51 as the US government forces Anthropic's Fable 5 offline — narrowing the open-to-closed gap to ~5 points. Artificial Analysis reweights toward agentic workloads.
Subscribe to get the next issue delivered to your inbox.