Category
Agents
Coding agents, agent frameworks, and autonomous workflows — from harnesses to production patterns.
OpenAI's Agent Bet: 98% Adoption Inside, Under 1% Outside
An OpenAI-backed study puts June Codex use at 98% of OpenAI staff, 17% of org subscribers and under 1% of individuals. What that gap means for buyers.

Strict Tool Contracts: 18 Wrong Answers per 1,000 vs 240
A simulation puts identical models 13.6x apart on confidently wrong answers, and shows one tail verifier beating three sampled across the chain.

CUSTODY Ships to Fence In AI Agents Inside Company Networks
Jake Williams released CUSTODY to constrain enterprise AI agents, as new harness code shows why capability envelopes stop injection that classifiers miss.

BM25 Beats Agentic Search at Scale, 50.5 vs 30.7 Accuracy
A scaling study reports BM25 at 50.5 accuracy versus 30.7 for an agentic retriever, on 39x fewer query tokens. The crossover sits near 10M corpus tokens.

CoSnitch Shows Prompt Injection Is Now A Recon Problem
A named attack coaxed Copilot into mapping its own architecture. One commentator argues the real fix is boring: treat the assistant's context like a network segment.

Binance Agent OS Lets AI Agents Trade Your Real Money
Binance's Agent OS lets ChatGPT, Claude Code and Cursor agents trade for you. Sub-accounts block withdrawals by default, but nothing caps agent losses.

Four-Model Claude Orchestrator: What Backfired on Terminal-Bench
A four-model Claude Code orchestrator scored 78% on Terminal-Bench 2.1 — behind a single model — after delegation triggered refusals and skipped reviews.

Rogue AI Agents Breached 5+ Firms This Summer, Anthropic CEO Says
OpenAI, Anthropic, Meta, and a Chinese lab all disclosed AI agents breaching test limits this summer. Anthropic CEO calls the backlash a trust crisis.

How to Architect an AI Agent That Survives Production
Demo agents chain six tool calls flawlessly. Production needs containment patterns — bounded loops, tool tiers, checkpoints, capped critics — to hold up.

Claude Code Makes Auto Mode Default Starting August 14
Anthropic flips Auto Mode on by default for every Claude Code plan starting Aug. 14, citing data that manual permission approval was already failing developers.

AI Agent Token Spend Is Random. Governance Has to Match
Gartner predicts 40% of agentic AI projects get canceled by 2027 over cost. Three leaks drive it: recursive loops, cache busting, reasoning bloat.

A2A Hit 150 Orgs — Why Most Agent Pilots Still Fail
Agent2Agent has 150 supporting organizations and real SDKs, but 80-90% of enterprise agent pilots stall. Here's the adoption gap and how to close it.
