Category
Infrastructure
Inference, serving, evals, and the infrastructure layer underneath AI-powered products.
Production RAG Pipelines: Fix Ingestion and Chunking First
Most RAG failures trace back to ingestion and chunking, not the LLM. Here's the pattern for structuring documents and metadata before you touch embeddings.

AI Infra Reality Check: $1B Compute Deals and Uptime Gaps
Reflection AI signs a $1B Nebius compute deal, an indie builder ditches GPT-4o for cheaper models, and a new tool tracks uptime across 77 AI APIs.

How OpenAI Scaled PostgreSQL to 800 Million ChatGPT Users
OpenAI's engineering team details how one PostgreSQL primary and nearly 50 read replicas handle millions of queries per second for 800 million users.

Vercel AI Gateway's 8-Model Week: Sonnet 4.6 to Opus 4.7
In one week, Vercel's AI Gateway added Claude Sonnet 4.6, Opus 4.7, Kimi K2.6, GPT-5.5, GPT-5.4 Mini/Nano, GPT Image 2, Responses API, and team ZDR.

AMD ROCm Skills Were Zero on skills.sh. NVIDIA Had 428
One developer counted 428+ agent skills for NVIDIA on skills.sh and zero for AMD ROCm, then built the first open-source pack of 10 — plus a multi-GPU trick.

How to Monetize an MCP Server With x402 Micropayments
A developer turned the useless HTTP 402 status code into a working pay-per-call system for an MCP tool server, using USDC on Base instead of subscriptions.

How to Give AI Agents Live GPU Metrics on OKE, No Prometheus
Your AI coding agent can read logs and pod status on Oracle Kubernetes Engine, but GPU utilization stays invisible. Here's how to fix that without Prometheus.
