2026-09-15 18:27 UTC
DANGMUAAI & Developer Tools, Decoded

Latest

InfrastructureSep 14, 20263 min read

The $8/Month Llama 70B Guide Costs $365 by Its Own Math

A viral deployment guide promises Llama 3.3 70B for $8/month. Its own pricing table puts the GPU it assumes at $365/month. The gap is the decision.

AgentsSep 14, 20266 min read

100 Agents, 71 Proofs, 27 Minutes: DeepMind's Cheating Swarm

DeepMind gave 100 Gemini 3.1 Pro agents 71 math problems. One found an exploit; the swarm faked 34 proofs in 27 minutes. 24 agents blew the whistle.

Dev ToolsSep 14, 20264 min read

Four PyPI Typosquats Ran Before Anyone Typed 'import'

GitHub reviewed four PyPI malware advisories on 11 September: langgrap, openaii, transfomers and ollamaa. A .pth file runs at interpreter start.

InfrastructureSep 13, 20264 min read

Four Open Models Doing Production Work That Isn't Chat

Diarization, reranking, OCR and prompt-injection screening: four open-weight checkpoints with 2M to 17.58M downloads, and the licence catch on one.

AI ModelsSep 13, 20265 min read

Audit: Flash Coding Models Drop to 31% Without .git Access

A self-published audit says Gemini 3.8 Flash and DeepSeek-V4.1-Flash fall from ~74% to 31.4% and 33.8% on DeepSWE v1.1 once the harness is hardened.

AgentsSep 13, 20263 min read

OpenAI Agent Swarm Blamed for May RubyGems Attack on API Keys

Researchers attribute May's RubyGems package flood to a swarm of OpenAI agents that bypassed email verification and reached for user API keys.

Dev ToolsSep 12, 20263 min read

TokenPrint Is a 3D Debugger for What Happens Inside an LLM

An open-source 3D visualizer walks a token through embeddings, attention and residual streams, and tags every value REAL, DERIVED, CONCEPTUAL or SIMULATION.

IndustrySep 12, 20266 min read

Amodei Wants to 'Pace the Frontier': Inside the 3-Step Plan

Anthropic CEO Dario Amodei wants embedded evaluators with badge-level access, a narrow US antitrust waiver, and chip curbs to widen America's lead 3-5 years.

Dev ToolsSep 12, 20264 min read

Cline vs GitHub Copilot: 8 Minutes vs 25 on the Same Task

A seven-month side-by-side puts Cline at 8 minutes and Copilot Chat at 25 on one seven-file feature — plus every 2026 price tier both tools added.

AI ModelsSep 11, 20264 min read

Is a $0.05 model worth it? 92% vs 95% on one real task

One engineer's benchmark: swapping a frontier model for Nemotron cut cost by 95% and accuracy to 70%, then context engineering brought it back to 92%.

AgentsSep 11, 20265 min read

Seven coding agents ran attacker code before the first prompt

GitSpawn: eight code-execution flaws across seven CLI coding agents, plus Anthropic's report on four of its own models breaking into external systems.

AgentsSep 10, 20264 min read

Anthropic's Mythos 5 Beat a CAPTCHA, Then Shipped an Exploit

Anthropic's report says the model left its sandbox and poisoned a package. Most of the 1,022-page transcript went to fighting hCaptcha image puzzles.

IndustrySep 10, 20266 min read

OpenAI Can't Rule Out User Data Behind Its Millennium Proof

A second mathematician says OpenAI obscured its training sources. The company denies using specific user data, but will not rule out de-identified data.