Category
AI Models
Frontier and open-weight model releases — capabilities, benchmarks, pricing, and what changed.
Jev vs Luna: Is a 1-Point Win Worth Swapping Your Reviewer?
TypeSafe's Jev edges GPT-5.6 Luna 67.8% to 66.8% on vendor evals graded by GPT-6 Astra and Claude Fable 5.1. What the score does and doesn't show.

Audit: Flash Coding Models Drop to 31% Without .git Access
A self-published audit says Gemini 3.8 Flash and DeepSeek-V4.1-Flash fall from ~74% to 31.4% and 33.8% on DeepSWE v1.1 once the harness is hardened.

Is a $0.05 model worth it? 92% vs 95% on one real task
One engineer's benchmark: swapping a frontier model for Nemotron cut cost by 95% and accuracy to 70%, then context engineering brought it back to 92%.

Gemini 3.8 Flash Holds 3.7 Pricing, Adds a Gated Cyber Model
Google's 3.8 Flash keeps 3.7 Flash introductory API pricing until January, ships no published benchmarks, and gates its cybersecurity variant.

VIDRAFT Says 2M Hugging Face Downloads, No Pretraining
Korean startup VIDRAFT merges and fine-tunes open LLMs instead of pretraining them. What AETHER and POCKET offer, and which rankings are its own.

Two $10/$50 Flagships in 72 Hours: LLM Price Index Up 29%
GPT-6 Astra took OpenAI's index slot at $10/$50 and moved a ten-model price index 29% in a day. DeepSeek went time-of-day; Sol's cut expires Nov 21.

Claude Fable 5.1 Pricing: Only the Cache Read Got Cheaper
Fable 5.1 kept four of its five published rates and cut the cache read from $1.00 to $0.25 per million tokens. Whether that pays is your hit ratio.

OpenAI Rates GPT-6 Astra 'Critical' for Cybersecurity
OpenAI rated GPT-6 Astra Critical for cybersecurity and shipped it off by default. What the evals measured, how access works, and the audit gap left behind.

Gemini 3.8 Flash Ships at $0.75 per Million Input Tokens
Google's third Flash model in six weeks holds 3.7 Flash pricing until January, ships a restricted Cyber variant, and warns it may burn more tokens.

Gemini 3.5 Transcribe: 4.0% WER in streaming preview
Google's new speech-to-text model reports 4.0% streaming WER, 85+ languages and a three-speaker diarization ceiling. Here's what shipped and what didn't.

Inherent Says Its 27B Agent Beat Claude Opus and GPT-5.5
Inherent's Faraday, on a 27B Qwen 3.6, reportedly beat frontier agents at paper replication. Plus Ora's harness benchmark pointing the same way.

GPT-5.6 Sol Is 50% Off Until Sept 18, And o3 Retires Aug 26
Vercel is discounting gpt-5.6-sol by 50% through September 18 via AI Gateway, with no code change. OpenAI retires o3 on August 26. Two clocks, one config.
