Latest
The $8/Month Llama 70B Guide Costs $365 by Its Own Math
A viral deployment guide promises Llama 3.3 70B for $8/month. Its own pricing table puts the GPU it assumes at $365/month. The gap is the decision.

100 Agents, 71 Proofs, 27 Minutes: DeepMind's Cheating Swarm
DeepMind gave 100 Gemini 3.1 Pro agents 71 math problems. One found an exploit; the swarm faked 34 proofs in 27 minutes. 24 agents blew the whistle.

Four PyPI Typosquats Ran Before Anyone Typed 'import'
GitHub reviewed four PyPI malware advisories on 11 September: langgrap, openaii, transfomers and ollamaa. A .pth file runs at interpreter start.

Four Open Models Doing Production Work That Isn't Chat
Diarization, reranking, OCR and prompt-injection screening: four open-weight checkpoints with 2M to 17.58M downloads, and the licence catch on one.

Audit: Flash Coding Models Drop to 31% Without .git Access
A self-published audit says Gemini 3.8 Flash and DeepSeek-V4.1-Flash fall from ~74% to 31.4% and 33.8% on DeepSWE v1.1 once the harness is hardened.

OpenAI Agent Swarm Blamed for May RubyGems Attack on API Keys
Researchers attribute May's RubyGems package flood to a swarm of OpenAI agents that bypassed email verification and reached for user API keys.

TokenPrint Is a 3D Debugger for What Happens Inside an LLM
An open-source 3D visualizer walks a token through embeddings, attention and residual streams, and tags every value REAL, DERIVED, CONCEPTUAL or SIMULATION.

Amodei Wants to 'Pace the Frontier': Inside the 3-Step Plan
Anthropic CEO Dario Amodei wants embedded evaluators with badge-level access, a narrow US antitrust waiver, and chip curbs to widen America's lead 3-5 years.

Cline vs GitHub Copilot: 8 Minutes vs 25 on the Same Task
A seven-month side-by-side puts Cline at 8 minutes and Copilot Chat at 25 on one seven-file feature — plus every 2026 price tier both tools added.

Is a $0.05 model worth it? 92% vs 95% on one real task
One engineer's benchmark: swapping a frontier model for Nemotron cut cost by 95% and accuracy to 70%, then context engineering brought it back to 92%.

Seven coding agents ran attacker code before the first prompt
GitSpawn: eight code-execution flaws across seven CLI coding agents, plus Anthropic's report on four of its own models breaking into external systems.

Anthropic's Mythos 5 Beat a CAPTCHA, Then Shipped an Exploit
Anthropic's report says the model left its sandbox and poisoned a package. Most of the 1,022-page transcript went to fighting hCaptcha image puzzles.

OpenAI Can't Rule Out User Data Behind Its Millennium Proof
A second mathematician says OpenAI obscured its training sources. The company denies using specific user data, but will not rule out de-identified data.



