2026-10-10 11:35 UTC
DANGMUAAI & Developer Tools, Decoded

Category

Infrastructure

Inference, serving, evals, and the infrastructure layer underneath AI-powered products.

InfrastructureOct 10, 20263 min read

Vercel Starts Charging Pro Teams for Deployment Storage

Vercel will bill all Pro teams $0.10 per GB-month for deployment and functions storage, and delete deployments older than 30 days starting October 23.

InfrastructureOct 01, 20263 min read

Cloudflare Opens a 402 Beta: Agents Pay Per Call in USDC

Monetization Gateway prices websites, APIs, MCP tools and datasets per request over HTTP 402, settling in USDC on Base. Four customers are live, US only.

InfrastructureSep 30, 20264 min read

Restate Raises $20M to Take On Temporal in Agent Durability

Berlin's Restate raised a $20M Series A led by Singular. It was built for durable workflows in 2022, not agents - and that turned out to be the fit.

InfrastructureSep 28, 20264 min read

Nvidia Ships OpenShell to GA and Adds Sentry for Agents

Nvidia's kernel-level agent sandbox hits general release alongside Sentry, a BlueField-resident monitor built to quarantine agents that stray.

InfrastructureSep 27, 20266 min read

Docker's Kit Spec Puts Agent Permissions in the OCI Image

Docker shipped Sandbox Kit Spec v3 on September 24, moving an agent's network and credential grants out of shell history into an OCI image annotation.

InfrastructureSep 26, 20263 min read

Anthropic Commits $11.6B to Akamai in a Bet on CPUs

Anthropic will pay Akamai $11.6 billion over seven years for cloud capacity, six times its May deal, with a warrant that could push it to $20 billion.

InfrastructureSep 20, 20265 min read

Speculative Decoding Pays 4x Until Concurrency Hits 128

A production write-up reports EAGLE-3 speculative decoding at 4x-5.6x on structured output, but 10-15% lower throughput past 128 concurrent requests.

InfrastructureSep 20, 20263 min read

SGLang vs vLLM: 4.47x Faster TTFT, But Only With Prefixes

One 8x H100 benchmark reports SGLang beating vLLM 4.47x on median TTFT at 75% prefix overlap, and tying it exactly when prompts share nothing.

InfrastructureSep 15, 20265 min read

77% Can Inventory Their AI Agents. 44% Can Verify It.

Harness surveyed 700 engineering leaders: 77% claim a complete agent inventory, 44% run tooling that proves it. The kill-switch gap is wider still.

InfrastructureSep 14, 20263 min read

The $8/Month Llama 70B Guide Costs $365 by Its Own Math

A viral deployment guide promises Llama 3.3 70B for $8/month. Its own pricing table puts the GPU it assumes at $365/month. The gap is the decision.

InfrastructureSep 13, 20264 min read

Four Open Models Doing Production Work That Isn't Chat

Diarization, reranking, OCR and prompt-injection screening: four open-weight checkpoints with 2M to 17.58M downloads, and the licence catch on one.

InfrastructureSep 04, 20263 min read

Can You Run a 70B Model on a 4GB GPU? AirLLM Says Yes

AirLLM claims 70B inference on a 4GB GPU by streaming one layer at a time. The VRAM table, the 3x quantization option, and the numbers it omits.