GLM-5.2: 1M Context and IndexShare Sparse Attention for Coding
Technical guide to GLM-5.2 — Z.ai's open-weight 744B MoE model with 1M token context, IndexShare sparse attention, and local execution via llama.cpp.
Category
Technical deep dives into model architectures, training approaches, and inference optimizations.
Technical guide to GLM-5.2 — Z.ai's open-weight 744B MoE model with 1M token context, IndexShare sparse attention, and local execution via llama.cpp.
Hands-on analysis of Moonshot AI's K2.7-Code — 384-expert MoE, $0.95/M tokens, beats Opus 4.8 on MCP tool calling, with API walkthrough and deployment guide.
Anthropic's Sonnet 5 delivers near-Opus 4.8 agentic performance at mid-tier pricing. Benchmark analysis, API changes, tokenizer shock, and migration guidance.
Tencent Hy3: 295B open MoE with Apache 2.0, 21B active, 256K context, $0.20/M tokens. Hallucinations halved, 47% fewer tokens than GLM-5.2 in production.
Grok 4.5: 1.5T MoE trained with Cursor, $2/$6 per M tokens, beats GPT-5.5 on SWE-bench, 4.2x efficient vs Opus 4.8. CursorBench contamination clouds claims.
Google targets July 17 GA for Gemini 3.5 Pro: a reported 2M-token context, Deep Think reasoning, and autonomous coding — with no government access restriction.
Moonshot AI's Kimi K3: 2.8T parameters, 1M context, open weights July 27. Beats Claude Fable 5 on coding, scores 57 on AI Index at $3/M input.
Alibaba's Qwen 3.8 packs 2.4T parameters, open-weight release imminent. Max Preview is live on QwenCloud's Token Plan. Stacks up against Kimi K3 and Fable 5.
Google launches Gemini 3.6 Flash (17% cheaper, smarter coding), 3.5 Flash-Lite (350 tok/s at $0.30/M), and Flash Cyber. Gemini 4 pre-training begins.
ETH Zurich and EPFL release Apertus 1.5 — fully open weights, data, and training code. 8B/70B models with image understanding, thinking mode, 262K context.
Meta ships Muse 2, a frontier-class open-weight model rivaling Fable 5 and GPT-5.6 Sol on coding and reasoning. Weights open at $0.60/$2.40 per million.