Tag
#claude
Inherent Says Its 27B Agent Beat Claude Opus and GPT-5.5
Inherent's Faraday, on a 27B Qwen 3.6, reportedly beat frontier agents at paper replication. Plus Ora's harness benchmark pointing the same way.

Claude Code vs eve: Same Models, eve Takes 7% Fewer Steps
Ora benchmarked Claude Code against Vercel's eve on live sites using the same models: 7% fewer steps, 2x native success, 9% more valid endpoints.

Binance Agent OS Lets AI Agents Trade Your Real Money
Binance's Agent OS lets ChatGPT, Claude Code and Cursor agents trade for you. Sub-accounts block withdrawals by default, but nothing caps agent losses.

Claude Code Sandbox Escape Rated CVSS 7.7 on macOS
A glob-parsing flaw let a folder name widen Claude Code's macOS sandbox write rules, reaching hook config and running commands before authentication.

Four-Model Claude Orchestrator: What Backfired on Terminal-Bench
A four-model Claude Code orchestrator scored 78% on Terminal-Bench 2.1 — behind a single model — after delegation triggered refusals and skipped reviews.

Anthropic Watermarks Claude Text and Images for EU Rules
Anthropic will embed invisible watermarks and signed provenance metadata in Claude output to meet EU AI Act rules — and some users are already pushing back.

Vercel Sandbox Runtimes Are Dead, Managed Images Take Over
Vercel replaced Sandbox runtimes with versioned Managed Images. Sandbox SDK v3 defaults to a Ubuntu image bundling claude-code, codex, and opencode.

Claude Code Makes Auto Mode Default Starting August 14
Anthropic flips Auto Mode on by default for every Claude Code plan starting Aug. 14, citing data that manual permission approval was already failing developers.

Claude Code vs Cursor for EU Teams: Pricing and GDPR Compared
Cursor costs $40/seat monthly; Claude Code bills per token, 400-2,000 EUR/month for a team of 15. Here is the GDPR, EU AI Act, and real-world tradeoff.

DeepSeek's Price Hike Signals the End of Cheap AI APIs
DeepSeek plans to raise API prices across the board, ending a year of cuts. Claude Code costs and a $12/month self-hosted Llama setup show what it means.

Kimi K3: The Biggest Open-Weight Model You Still Can't Run
Moonshot AI's Kimi K3 hits 2.8T parameters, the largest open-weight model yet, rivaling Claude Opus — but it needs an enterprise GPU cluster, not a homelab.

GhostApproval Symlink Flaw Hit Six AI Coding Assistants
A July 8 Wiz Research disclosure showed how a symlink disguised as a settings file could trick agents into writing to your SSH keys behind an approval dialog.
