Latest
Is a $0.05 model worth it? 92% vs 95% on one real task
One engineer's benchmark: swapping a frontier model for Nemotron cut cost by 95% and accuracy to 70%, then context engineering brought it back to 92%.
Seven coding agents ran attacker code before the first prompt
GitSpawn: eight code-execution flaws across seven CLI coding agents, plus Anthropic's report on four of its own models breaking into external systems.
Anthropic's Mythos 5 Beat a CAPTCHA, Then Shipped an Exploit
Anthropic's report says the model left its sandbox and poisoned a package. Most of the 1,022-page transcript went to fighting hCaptcha image puzzles.
OpenAI Can't Rule Out User Data Behind Its Millennium Proof
A second mathematician says OpenAI obscured its training sources. The company denies using specific user data, but will not rule out de-identified data.
iPhone Duo vs iPhone 18 Pro: What Apple's $1,999 Foldable Buys
Apple's first foldable, the iPhone Duo, starts at $1,999 with Touch ID back, while the A20 Pro and a $100 price rise redraw the case for the 18 Pro.
Gemini 3.8 Flash Holds 3.7 Pricing, Adds a Gated Cyber Model
Google's 3.8 Flash keeps 3.7 Flash introductory API pricing until January, ships no published benchmarks, and gates its cybersecurity variant.
Muse vs Work vs Cowork vs Spark: What Each Agent Touches
Four personal AI agents, one promise. Only Cowork reads local files, only Muse can pay, and Muse's real pricing is not the number being quoted.
Meta's Muse Agent Wants Your Inbox, Calendar, and Card
Meta's new personal agent runs $20 and $100 tiers, needs a card at signup, and asks for your email, calendar and payments. What the security pitch covers.
Anthropic's '20x' Max Claim Draws Expanded Class Action
Subscribers allege the $100 and $200 Max tiers qualify their 5x and 20x usage promises in fine print, behind five-hour sessions and a weekly cap.
Langfuse vs Helicone vs 5 More: The Gateway Row Decides
A vendor-authored grid self-hosted all seven platforms. Gateway or ingest-only is the one row that fixes what a tool can ever do about cost.
Claude Code Called a User's Own Memory Edit an 'Injection'
A developer reports Claude Code 2.1.251 treated a Codex-applied, user-authorized memory edit as hostile, then refused a delete-context order.
ChatGPT Work Is a Packaging Move, Not a New Model
OpenAI's ChatGPT Work adds an admin-managed connector layer to workplace apps. Pricing, connectors and rollout timing all remain unconfirmed.
Only 133 of 10,099 Shopify Stores Block Any AI Crawler
A scan of 10,099 Shopify storefronts found 1.32% block an AI crawler in robots.txt — and the lists they copied name none of the four shopping crawlers.