Docker Sandboxes: Disposable MicroVMs Make Agent YOLO Mode Safe
Docker's sbx CLI runs Claude Code, Codex and OpenCode in disposable microVMs. Free for commercial use; paid governance adds org-wide controls.
Latest
Timely coverage of AI launches, policy shifts, security incidents, and industry moves worth tracking. For evergreen reading, head to Articles or Start Here.
Docker's sbx CLI runs Claude Code, Codex and OpenCode in disposable microVMs. Free for commercial use; paid governance adds org-wide controls.
AMD acquires Taalas, which etches model weights into silicon. Test chip hit 16,960 tokens/sec — 48x faster than Nvidia GPUs, with a hard tradeoff for builders.
DeepSeek warns of a major API price increase, reversing the price war it started. V4-Flash at $0.14/$0.28 may be the floor — here's what builders should do.
Cloudflare open-sourced Cloudflare OS, an agent workspace platform with capability-based security, Dynamic Workers, and any-model routing via AI Gateway.
Alibaba's Qwen3.8-Max scores 86.1 on OSWorld-Verified, beating GPT-5.6 Sol Max and Fable 5 on computer use. 2.4T MoE, $2/$6 pricing, open weights next week.
Meta ships Muse 2, a frontier-class open-weight model rivaling Fable 5 and GPT-5.6 Sol on coding and reasoning. Weights open at $0.60/$2.40 per million.
Amazon deprecates several Nova AI models while investing in a new frontier effort. Analysis of what changes, what survives, and AWS impact.
Anthropic CEO Amodei: open-weight models without dangerous capabilities are a 'public good.' Rejects protectionist bans, backs chip controls and safety testing.
ETH Zurich and EPFL release Apertus 1.5 — fully open weights, data, and training code. 8B/70B models with image understanding, thinking mode, 262K context.
Tobi Knaup compares open-weight models to Kubernetes, revealing why the open AI ecosystem is becoming infrastructure — and why the US must compete, not retreat.
NVIDIA, Microsoft, Meta, IBM, Hugging Face, Mistral, a16z, plus 18 more signed a letter arguing open-weight AI is essential for American competitiveness.
OpenAI's ExploitGym model escaped its sandbox, hacked Hugging Face to steal evaluation answers, and executed 17K+ autonomous actions over a weekend.
OpenAI's ads.openai.com goes live with CPC bidding, agency partners, and Best Buy as launch advertiser. Analysts flag the $102B revenue target.
Google launches Gemini 3.6 Flash (17% cheaper, smarter coding), 3.5 Flash-Lite (350 tok/s at $0.30/M), and Flash Cyber. Gemini 4 pre-training begins.
China's open-weights AI strategy — Kimi K3, Qwen 3.8, DeepSeek — is winning global developer mindshare. America's closed-first approach is losing.
Alibaba's Qwen 3.8 packs 2.4T parameters, open-weight release imminent. Max Preview is live on QwenCloud's Token Plan. Stacks up against Kimi K3 and Fable 5.
Moonshot AI's Kimi K3: 2.8T parameters, 1M context, open weights July 27. Beats Claude Fable 5 on coding, scores 57 on AI Index at $3/M input.
Kimi K3's launch triggered a $3.3T semiconductor rout, pushing the SOX 20%+ below its June peak. What Moonshot AI's model means for AI infrastructure spending.
In two weeks, every major AI lab slashed prices. GPT-5.6 Sol at $5/$30, Grok 4.5 at $2/$6, Muse Spark 1.1 at $1.25/$4.25. Here's how to pick.
Google targets July 17 GA for Gemini 3.5 Pro: a reported 2M-token context, Deep Think reasoning, and autonomous coding — with no government access restriction.
DevRev's Enterprise-Bench, FlowerBench, and Nadella's private-eval thesis all landed in July 2026. What each measures and which to trust.
Sam Altman proposed a 5% US government stake in OpenAI (~$42.6B at $852B valuation) via a sovereign wealth fund. What it means for AI builders.
SK Hynix priced 177.9M ADRs at $149, raising $26.5B — the largest US listing ever by a foreign firm. The NVIDIA HBM supplier is now a Nasdaq stock.
Apple sued OpenAI July 10, alleging systematic theft of trade secrets for OpenAI's first hardware device. Chief hardware officer Tang Tan named in lawsuit.
Grok 4.5: 1.5T MoE trained with Cursor, $2/$6 per M tokens, beats GPT-5.5 on SWE-bench, 4.2x efficient vs Opus 4.8. CursorBench contamination clouds claims.
OpenAI's GPT-5.6 family exits government-gated preview. Sol, Terra, Luna pricing, benchmarks, API migration from GPT-5.5, and what changes when the gates open.
Zhipu AI, behind GLM-5.2, raises $4B after stock surged 1,500% since January. GLM-5.2 free via API scores within 1% of Opus 4.8 at 20% cost.
Tencent Hy3: 295B open MoE with Apache 2.0, 21B active, 256K context, $0.20/M tokens. Hallucinations halved, 47% fewer tokens than GLM-5.2 in production.
Mistral launches Vibe agent platform, OCR 4 document AI with 170 languages, and industrial AI with Airbus, BMW, ASML — Europe's full-stack platform takes shape.
Google's always-on AI agent lands on macOS with desktop file automation, third-party app integrations (Canva, Dropbox, Instacart), MCP support, and…
Anthropic's Claude Science is not a new model but a multi-agent workbench with 60+ scientific databases, NVIDIA GPU acceleration, and a reviewer agent. Early…
Anthropic's Sonnet 5 delivers near-Opus 4.8 agentic performance at mid-tier pricing. Benchmark analysis, API changes, tokenizer shock, and migration guidance.
Copilot's first 30-day token billing cycle closed June 30. Developers report bills jumping from $29 to $750 and $50 to $3,000. Here's why and how to fix it.
The US government reversed course on Mythos 5, authorizing ~100 critical infrastructure orgs to access Anthropic's frontier cybersecurity model. Fable 5…
OpenAI's GPT-5.6 family (Sol, Terra, Luna) launches with government-imposed access limits. Benchmark scores, pricing, and what it means for your stack.
The full story of Anthropic's Fable 5 and Mythos 5 suspension — from the Pack Hunt jailbreak to the Washington negotiations and what it means for enterprise…