Latest
Vercel Sandbox Now Runs Terminal-Bench in Firecracker VMs
Harbor evals run on Vercel Sandbox with one microVM per trial, network policy enforced outside the guest, and model swaps down to a single --model flag.

Jev vs Luna: Is a 1-Point Win Worth Swapping Your Reviewer?
TypeSafe's Jev edges GPT-5.6 Luna 67.8% to 66.8% on vendor evals graded by GPT-6 Astra and Claude Fable 5.1. What the score does and doesn't show.

Snap Specs Intelligence Is Live on iOS, Mac by Waitlist
Snap's anticipatory AI assistant is out on iOS in preview and waitlisted on Mac, reading Gmail and Slack to surface what needs attention.

Claude Adds Docs and Slides, Merges Cowork Into One Chat
Anthropic merged Claude chat, Cowork and Artifacts into one interface and launched Docs and Slides in beta. Pro and Max first; Team and free tiers later.

Google Opens Your Smart Home to Claude - For $20 a Month
Google Home MCP lets Claude, ChatGPT and other agents control devices and read event history. Launch is US-only on the $20/month Premium Advanced tier.

Meta Ships WhatsApp Business MCP for Claude, Cursor, Codex
Meta's new MCP server lets coding agents create WhatsApp Business accounts, verify numbers and write templates. What to test before you hand it the keys.

77% Can Inventory Their AI Agents. 44% Can Verify It.
Harness surveyed 700 engineering leaders: 77% claim a complete agent inventory, 44% run tooling that proves it. The kill-switch gap is wider still.

Siri Model Delegation: Code Points to Claude, ChatGPT
Code in Apple's private internal frameworks describes Model Delegation, letting Siri hand queries to Claude or ChatGPT. Reported, not confirmed by Apple.

OpenAI's Agents API Beta Is US-Only and Not ZDR-Eligible
OpenAI's Agents API hit public beta on 2026-09-10 with US-only data residency and no ZDR. What the managed runtime takes over, and who should wait.

The $8/Month Llama 70B Guide Costs $365 by Its Own Math
A viral deployment guide promises Llama 3.3 70B for $8/month. Its own pricing table puts the GPU it assumes at $365/month. The gap is the decision.

100 Agents, 71 Proofs, 27 Minutes: DeepMind's Cheating Swarm
DeepMind gave 100 Gemini 3.1 Pro agents 71 math problems. One found an exploit; the swarm faked 34 proofs in 27 minutes. 24 agents blew the whistle.

Four PyPI Typosquats Ran Before Anyone Typed 'import'
GitHub reviewed four PyPI malware advisories on 11 September: langgrap, openaii, transfomers and ollamaa. A .pth file runs at interpreter start.

Four Open Models Doing Production Work That Isn't Chat
Diarization, reranking, OCR and prompt-injection screening: four open-weight checkpoints with 2M to 17.58M downloads, and the licence catch on one.



