An Open 180B Model Tops 10 Hugging Face Boards, Runs on CPU
Darwin-180B-RSI leads 10 of 48 official Hugging Face leaderboards, and a 111 GB 4-bit build of it runs GPU-free at 18-21 tokens per second on CPU.

An open-weight 180B model now tops 10 of Hugging Face's 48 official leaderboards, and a 4-bit build runs GPU-free at 18-21 tokens per second.
The model is Darwin-180B-RSI, built by VIDRAFT on top of Alibaba's Qwen3.8-Flash-Next. According to the team's own write-up, that tally of 10 is the most of any organization competing on those 48 curated boards, where 95 orgs are in the field and the runner-up holds 4. The stated per-org count is FINAL-Bench (the VIDRAFT research org) 10, Z.ai 4, Xiaomi 3, DeepSeek 3, Moonshot 3.
Which boards, and which ones matter
The posted scores span science, math, law, vision and document work: GPQA Diamond 94.44, MMLU-Pro 88.12, MMMU-Pro 79.48, AIME 2026 and HMMT 2026 at 100, LEXam 68.94, LEXam-hard 45.72, ExtractBench 90.29, IFStruct 98.95, MDPBench 83.65.
Two of those are the ones to read if you ship software rather than take exams. IFStruct, built by Liquid AI, checks whether a model emits data in the exact requested shape and fails a response if a single field is missing, mistyped, out of range or invented; the post reports 1,979 of 2,000 items passed for 98.95, ahead of Agents-A1 (93.25), gpt-oss-20b (91.95) and Nemotron-3-Nano-30B (86.80). ExtractBench, built by LlamaIndex, measures field extraction across 370 PDFs running from a few pages to roughly 190. The 90.29 there came from Darwin-180B-RSI-R3, a further self-improved revision that edged past its own base model, Qwen3.8-Flash-Next, at 89.88.
Schema adherence is a gating requirement for a pipeline that feeds model output into the next program, so a board that fails on one wrong character is closer to a procurement test than a benchmark.
The part that changes your hardware math
A separate post covers POCKET-Darwin-180B-GGUF, a 4-bit build packaged to run without a GPU. The figures given, all measured in October 2026: 111 GB across 4 GGUF files, down from 360 GB in BF16 across 131 files; a single 16-thread server CPU generating 18.4 to 21.0 tokens/s at 78.8 GB peak memory; an RTX 5060 Laptop (8 GB VRAM) with 32 GB RAM running it at 4.17 tokens/s; a 128 GB mini PC holding the whole model in memory with no GPU at all.
The reason it fits is architectural, not a trick of compression. Darwin-180B is a mixture-of-experts model with 512 expert sub-networks, of which 10 are selected per token, so roughly 3B of the 180B parameters are active for any given token. The quantization is selective on top of that: only the portion the self-improvement training actually modified, about 3% of total volume, is kept at higher precision.
On accuracy, the post reports a paired comparison on MMLU-Pro across 2,000 questions in which the original and the 4-bit build both scored 87.65%. On SuperGPQA, 1,000 graduate-level science questions the team says were never used in training, the quantized POCKET build scored 61.55% against 59.10% for the same quantization of its base model, a 2.45-point gain while using about 13% fewer tokens per question.
Is it worth your weekend?
Running it needs a llama.cpp build b11048 or later, the four GGUF files totalling 111 GB, and enough RAM to hold or page the model. The 128 GB mini PC path is the only one of the three that looks comfortable; the 32 GB laptop figure of 4.17 tokens/s is a demonstration that it loads, not a usable coding loop.
Weigh the leaderboard claims accordingly. They are self-reported, and the write-up says so plainly: the numbers are registered scores, each measured by the respective submitter using the board's official harness, and different boards normalize differently, so a cross-board ranking is a portfolio view rather than a single global ranking. The independently interesting claim is the narrower one — that a paired, item-by-item comparison showed no measurable loss from 4-bit on one benchmark. That is the result worth waiting for a third party to reproduce before you buy 128 GB of RAM.
More from DangMua