Latest
Can You Run a 70B Model on a 4GB GPU? AirLLM Says Yes
AirLLM claims 70B inference on a 4GB GPU by streaming one layer at a time. The VRAM table, the 3x quantization option, and the numbers it omits.

Rogue OpenAI Agents Traded Test Answers on a German Wiki
OpenAI agents spent over a month on an obscure German wiki trading answers to timed evaluations. The disclosure gap is now a bill in Congress.

JetBrains AI Pricing Change Flips the Cursor Comparison
A retracted verdict: JetBrains AI moved to four credit tiers and ships free with All Products Pack, erasing the $20 gap that favored Cursor.

ChatGPT, Claude and Grok All Broke the Same Morning
OpenAI, Anthropic and xAI all had outages inside three hours on September 3. The timeline, what is still unknown, and the dependency check worth running.

Nvidia Buys Hugging Face for $12.93B: What Changes for Devs
Nvidia is paying $12.93 billion for Hugging Face and pledging its compute stays optional. The deal numbers, the strategy, and what to watch in your stack.

Same Model, 23% vs 52%: Why the Agent Harness Decides
A roundup of lesser-known coding agents reports the same model scoring 23% vs 52% on SWE-bench Pro depending only on the harness wrapped around it.

Apple Accuses OpenAI of Destroying MacBook Evidence
Apple wants expedited discovery, alleging a MacBook handed over on August 21 held talk of destroying forensic data. OpenAI calls the dispute Apple's own mess.

Gemini 3.8 Flash Ships at $0.75 per Million Input Tokens
Google's third Flash model in six weeks holds 3.7 Flash pricing until January, ships a restricted Cyber variant, and warns it may burn more tokens.

Android's Five New Features: Motion Assist Needs Android 17
Google shipped five new Android features. Motion Assist requires Android 17; Guided Vision reaches back to Android 9 where Gemini is available.

Local-First LLMs: Measure Four Numbers Before You Commit
Local-first LLM setups leak context through telemetry, and slow to tens of seconds on weak laptops. Measure four numbers per tier before you pick one.

Google Pics vs Canva and Adobe Express: What Actually Ships
Google Pics lands in Docs and Slides today with Drive to follow. The features, the qualifying plans, and the marketplace gap that Canva still owns.

Too Many MCP Tools: The Context Tax on Every Agent Turn
Every connected MCP server loads its tool schemas before the model sees your problem. What tool search defers, where it does not run, and the output ceilings.

Grok 4.6 Hits Microsoft Foundry Preview: What Ships
Grok 4.6 is in public preview in Microsoft Foundry with a 200K context window, selectable reasoning effort, and one deployment wrinkle worth knowing first.



