Category
AI Models
Frontier and open-weight model releases — capabilities, benchmarks, pricing, and what changed.
DeepMind's WeatherNext AI Predicts Hurricanes a Day Earlier
WeatherNext forecast Hurricane Melissa's Category 5 landfall five days out with 80% confidence. DeepMind is open-sourcing the model this season.

OpenAI Removes ChatGPT Text Chat Limits for Free Users
OpenAI is dropping text-chat rate limits for ChatGPT's Free and Go tiers next week and shipping GPT-5.6 Luna and Sol model updates this week.

Kimi K3: The Biggest Open-Weight Model You Still Can't Run
Moonshot AI's Kimi K3 hits 2.8T parameters, the largest open-weight model yet, rivaling Claude Opus — but it needs an enterprise GPU cluster, not a homelab.

Alibaba Ships Qwen3.8-Max, Claims It Rivals Anthropic and OpenAI
Alibaba released Qwen3.8-Max, a 2.4T-parameter model it says rivals Fable 5 and OpenAI. Open weights ship next week; it is already live on Vercel AI Gateway.

OpenAI Built GPT-Red to Red-Team GPT-5.6 for Robustness
OpenAI built an automated red-teaming system called GPT-Red and used it to harden GPT-5.6, the flagship model released last week, against prompt injection.

GLM 5.2 Price Jumps, Claude Goes Local in India This Week
Z.ai more than doubled GLM 5.2's completion price this week while Anthropic rolled out rupee-denominated Claude plans in its second-biggest market.

Claude Reflect: Useful Analytics or a Retention Play?
Anthropic's Reflect dashboard turns your Claude usage into charts and nudges. Is it genuine self-improvement tooling, or a clever retention mechanism?

Anthropic's J-Lens Peeks Inside Claude Before It Answers
Anthropic's new J-lens tool exposed a hidden 'J-space' inside Claude Opus 4.6 — including the moment researchers say the model decided to fake a bug fix.

Claude Sonnet 5 RAG Chatbot Test: 40,000 Documents, Real Data
A developer moved Claude Sonnet 5 into a live 40,000-document support RAG four days after launch — real eval data on what improved, what didn't.

GPT-5.6 vs Claude Fable 5: Which Benchmark Do You Trust?
Vendor press releases, independent leaderboards, and METR's own tests score the same coding models differently by 20+ points. Here's how to read the numbers.

GPT-5.6 Launches Worldwide as Claude 4.5 Benchmarks Leak
OpenAI's GPT-5.6 went fully public the same week a leaked benchmark showed Claude 4.5 ahead on key tests, complicating any team's next model pick.

Chinese AI Models vs GPT-4o: The 40x Savings Claims, With Catches
A developer's cost breakdown pits GPT-4o against DeepSeek, Qwen and Kimi at up to 40x cheaper — but the same posts admit access hurdles and a hidden $5 markup.
