Category
Infrastructure
Inference, serving, evals, and the infrastructure layer underneath AI-powered products.
M5 Ultra vs DGX Spark: Which Desktop Runs a 125B Model?
Apple's M5 Ultra brings 1.2 TB/s bandwidth and 512GB unified memory. Here is how it stacks up against DGX Spark and Strix Halo for local inference.

OpenAI says Jalapeño beats Blackwell on inference benchmarks
OpenAI's Jalapeño ASIC posted 1.5-1.9x more AI work per watt and 1.7-3.6x lower latency than Nvidia GB200/GB300 on InferenceX. Volume ships in 2027.

Vercel Sandbox Goes Global: Four Regions, Failover on Pro
Vercel Sandbox now runs in iad1, sfo1, cle1 and cdg1, with failover for Pro and Enterprise. Snapshots can't move regions, so plan a rebuild.

Vercel Bills Deployment Storage at $0.10 per GB per Month
Vercel's rollback history now has a price tag: $0.10 per GB per month, 10 GB included on Hobby. How to trim the bill without losing instant rollback.

Vercel Offers $1M to Anyone Who Escapes Its AI Sandbox
Vercel opened a two-week HackerOne program paying up to $1M, with $50,000 per cross-tenant break, and published the microVM architecture it wants attacked.

Nvidia Puts $1.5B Into SoftBank's OpenAI Data Center Site
Nvidia invested $1.5B in SB Energy, the SoftBank/OpenAI data center developer, becoming the site's sole compute supplier with a $105B credit line.

SQLite-vec Vs. Pinecone: Is Local Vector Search Worth It?
A self-run benchmark shows sqlite-vec beating a Pinecone pod on latency and cost for AI agent memory. What the numbers show, and where they do not apply.

TPU Raiden: Google's Unconfirmed Answer to NVIDIA's NIXL
A single tweet claims Google open-sourced TPU Raiden, a KV-cache library rivaling NVIDIA NIXL. No official confirmation yet — here is what to watch.

pgvector vs LanceDB: What the Benchmark Numbers Mean
A 100k-vector benchmark shows LanceDB ingests 22x faster while pgvector wins at 8 concurrent clients — plus what embedding model choice costs at scale.

Production RAG Pipelines: Fix Ingestion and Chunking First
Most RAG failures trace back to ingestion and chunking, not the LLM. Here's the pattern for structuring documents and metadata before you touch embeddings.

AI Infra Reality Check: $1B Compute Deals and Uptime Gaps
Reflection AI signs a $1B Nebius compute deal, an indie builder ditches GPT-4o for cheaper models, and a new tool tracks uptime across 77 AI APIs.

How OpenAI Scaled PostgreSQL to 800 Million ChatGPT Users
OpenAI's engineering team details how one PostgreSQL primary and nearly 50 read replicas handle millions of queries per second for 800 million users.
