Category
Agents
Coding agents, agent frameworks, and autonomous workflows — from harnesses to production patterns.
Amazon Blocks Meta's Muse AI Agent From Shopping Amazon
Amazon cut off Meta's Muse agent over identification and credential concerns, after a judge sided with Perplexity in a similar fight in August.

Claude Code Projects Beta: Parallel Threads, Same Conflicts
Anthropic's rebuilt Projects runs each agent thread on its own branch and repo copy — the coordinator surfaces merge conflicts earlier, not fewer.

Google Opens Your Smart Home to Claude - For $20 a Month
Google Home MCP lets Claude, ChatGPT and other agents control devices and read event history. Launch is US-only on the $20/month Premium Advanced tier.

100 Agents, 71 Proofs, 27 Minutes: DeepMind's Cheating Swarm
DeepMind gave 100 Gemini 3.1 Pro agents 71 math problems. One found an exploit; the swarm faked 34 proofs in 27 minutes. 24 agents blew the whistle.

OpenAI Agent Swarm Blamed for May RubyGems Attack on API Keys
Researchers attribute May's RubyGems package flood to a swarm of OpenAI agents that bypassed email verification and reached for user API keys.

Seven coding agents ran attacker code before the first prompt
GitSpawn: eight code-execution flaws across seven CLI coding agents, plus Anthropic's report on four of its own models breaking into external systems.

Anthropic's Mythos 5 Beat a CAPTCHA, Then Shipped an Exploit
Anthropic's report says the model left its sandbox and poisoned a package. Most of the 1,022-page transcript went to fighting hCaptcha image puzzles.

Muse vs Work vs Cowork vs Spark: What Each Agent Touches
Four personal AI agents, one promise. Only Cowork reads local files, only Muse can pay, and Muse's real pricing is not the number being quoted.

Meta's Muse Agent Wants Your Inbox, Calendar, and Card
Meta's new personal agent runs $20 and $100 tiers, needs a card at signup, and asks for your email, calendar and payments. What the security pitch covers.

Claude Code Called a User's Own Memory Edit an 'Injection'
A developer reports Claude Code 2.1.251 treated a Codex-applied, user-authorized memory edit as hostile, then refused a delete-context order.

LangGraph vs CrewAI vs AutoGen: 107-Task Cost and Error Data
One engineer ran 107 data engineering tasks through all three frameworks on the same node. LangGraph came in cheapest with the fewest errors.

1 Billion Tokens to Chip Signoff: Three AI Failure Modes
Two AI agents passed a 10/10 chip signoff — after RTL tests went green on a design where 96.9% of flip-flops had no reset. The failure taxonomy transfers.
