Category
Infrastructure
Inference, serving, evals, and the infrastructure layer underneath AI-powered products.
ChatGPT, Claude and Grok All Broke the Same Morning
OpenAI, Anthropic and xAI all had outages inside three hours on September 3. The timeline, what is still unknown, and the dependency check worth running.

Local-First LLMs: Measure Four Numbers Before You Commit
Local-first LLM setups leak context through telemetry, and slow to tens of seconds on weak laptops. Measure four numbers per tier before you pick one.

NVIDIA Posted $96B the Quarter OpenAI's Own Chip Arrived
NVIDIA booked $96B at a 75% margin while its biggest customer benchmarked a 700W inference ASIC. What the numbers say about inference economics.

M5 Ultra vs DGX Spark: Which Desktop Runs a 125B Model?
Apple's M5 Ultra brings 1.2 TB/s bandwidth and 512GB unified memory. Here is how it stacks up against DGX Spark and Strix Halo for local inference.

OpenAI says Jalapeño beats Blackwell on inference benchmarks
OpenAI's Jalapeño ASIC posted 1.5-1.9x more AI work per watt and 1.7-3.6x lower latency than Nvidia GB200/GB300 on InferenceX. Volume ships in 2027.

Vercel Sandbox Goes Global: Four Regions, Failover on Pro
Vercel Sandbox now runs in iad1, sfo1, cle1 and cdg1, with failover for Pro and Enterprise. Snapshots can't move regions, so plan a rebuild.

Vercel Bills Deployment Storage at $0.10 per GB per Month
Vercel's rollback history now has a price tag: $0.10 per GB per month, 10 GB included on Hobby. How to trim the bill without losing instant rollback.

Vercel Offers $1M to Anyone Who Escapes Its AI Sandbox
Vercel opened a two-week HackerOne program paying up to $1M, with $50,000 per cross-tenant break, and published the microVM architecture it wants attacked.

Nvidia Puts $1.5B Into SoftBank's OpenAI Data Center Site
Nvidia invested $1.5B in SB Energy, the SoftBank/OpenAI data center developer, becoming the site's sole compute supplier with a $105B credit line.

SQLite-vec Vs. Pinecone: Is Local Vector Search Worth It?
A self-run benchmark shows sqlite-vec beating a Pinecone pod on latency and cost for AI agent memory. What the numbers show, and where they do not apply.

TPU Raiden: Google's Unconfirmed Answer to NVIDIA's NIXL
A single tweet claims Google open-sourced TPU Raiden, a KV-cache library rivaling NVIDIA NIXL. No official confirmation yet — here is what to watch.

pgvector vs LanceDB: What the Benchmark Numbers Mean
A 100k-vector benchmark shows LanceDB ingests 22x faster while pgvector wins at 8 concurrent clients — plus what embedding model choice costs at scale.
