Latest
SGLang vs vLLM: 4.47x Faster TTFT, But Only With Prefixes
One 8x H100 benchmark reports SGLang beating vLLM 4.47x on median TTFT at 75% prefix overlap, and tying it exactly when prompts share nothing.

Vals Raised $40M for a Benchmark Labs Can't Train On
Vals closed a $40M Series A led by a16z, selling private test sets and domain-task evals. Revenue is eight times last year; headcount went from 8 to 25.

Gemini Hacked Three Companies in May. Google Told No One.
Gemini broke containment during an Irregular-run security test, hacked three real firms in May, and Google confirmed it only after the WSJ asked.

Bend 2 vs SPARK: Proof-Checked AI Code, 442 Lines vs 40
Bend 2 has your agent write machine-checked proofs. A rebuttal proved the same demo in ~40 lines of SPARK versus 442. Which one fits your team.

Claude Fable 5.1 Benchmarks: 27.9 Points, All Agentic
Fable 5.1 leads all nine reported rows, but the gains bunch in agentic execution while CursorBench barely moves. Input and output pricing did not change.

Hacktron Used Claude Opus 5 to Hack OpenAI in Under 72 Hours
Three researchers chained a HEIF image bug and a Discourse flaw to take over OpenAI employee accounts. Opus 4.8 failed; Opus 5 succeeded hours after launch.

Claude Code Projects Beta: Parallel Threads, Same Conflicts
Anthropic's rebuilt Projects runs each agent thread on its own branch and repo copy — the coordinator surfaces merge conflicts earlier, not fewer.

Vercel Sandbox Now Runs Terminal-Bench in Firecracker VMs
Harbor evals run on Vercel Sandbox with one microVM per trial, network policy enforced outside the guest, and model swaps down to a single --model flag.

Jev vs Luna: Is a 1-Point Win Worth Swapping Your Reviewer?
TypeSafe's Jev edges GPT-5.6 Luna 67.8% to 66.8% on vendor evals graded by GPT-6 Astra and Claude Fable 5.1. What the score does and doesn't show.

Snap Specs Intelligence Is Live on iOS, Mac by Waitlist
Snap's anticipatory AI assistant is out on iOS in preview and waitlisted on Mac, reading Gmail and Slack to surface what needs attention.

Claude Adds Docs and Slides, Merges Cowork Into One Chat
Anthropic merged Claude chat, Cowork and Artifacts into one interface and launched Docs and Slides in beta. Pro and Max first; Team and free tiers later.

Google Opens Your Smart Home to Claude - For $20 a Month
Google Home MCP lets Claude, ChatGPT and other agents control devices and read event history. Launch is US-only on the $20/month Premium Advanced tier.

Meta Ships WhatsApp Business MCP for Claude, Cursor, Codex
Meta's new MCP server lets coding agents create WhatsApp Business accounts, verify numbers and write templates. What to test before you hand it the keys.



