2026-09-25 11:38 UTC
DANGMUAAI & Developer Tools, Decoded
BackAI Models

PrismML's 1-Bit LLM Hits Snapdragon Glasses, None Shipping Yet

Qualcomm showed PrismML's 1-bit Bonsai running on Snapdragon AR1 Gen 1 glasses: a 2B vision model, 4x smaller. No glasses have been announced yet.

DangMua EditorialSep 25, 20264 min read
PrismML's 1-Bit LLM Hits Snapdragon Glasses, None Shipping Yet

Qualcomm showcased PrismML's 1-bit Bonsai LLM running locally on Snapdragon AR1 Gen 1 smart glasses this week — a 2-billion-parameter vision-language model.

The demo came at Qualcomm's Snapdragon Summit on Wednesday. TechCrunch reports the glasses build shrinks a larger model by 4x while "retaining almost all of their performance on standard benchmarks," and is tuned for vision and language so wearers can ask what they are looking at in real time.

Who PrismML is

PrismML was founded by a group of Caltech researchers and is led by Babak Hassibi, a Caltech professor and an expert in compression technologies. Ion Stoica — a Databricks co-founder and director of Berkeley's Sky Computing Lab — advises the company, which is backed by Khosla Ventures, Cerberus Capital and Caltech on a $22.25 million seed round.

The pitch is open-weight AI that runs on hardware people already own, positioned as an alternative to depending on the privacy promises of proprietary AI labs. "You are going to have intelligence at your fingertips, and it's going to be free because it's going to run on the device you already bought," Stoica told TechCrunch last week. "It's also going to be private, because you're not going to send it to the cloud."

The compression numbers

PrismML's flagship release is Bonsai 2 27B, out September 17, which compresses Alibaba's widely used open-source Qwen3.8 27B down to 5.9 GB — a 9x to 10x memory reduction that TechCrunch reports is small enough for a PC and possibly a high-end smartphone. Hassibi's claim is that almost nothing is lost: Bonsai 2 matches 98% of Qwen's aggregate benchmark scores, up from 95% for the first Bonsai, released in March. PrismML says that original model has been downloaded over 11 million times and its smaller models another 2.6 million.

The technique is ternary weights. Where a weight normally requires 16 bits, PrismML's approach simplifies it to three values: +1, −1, or 0. The labels do not line up cleanly across the two announcements — the September 17 coverage describes ternary weights, while the glasses model is billed as "1-bit Bonsai." Neither report reconciles the two.

What is not shipping

No smart glasses running PrismML have been announced yet. The Snapdragon Summit appearance is a platform showcase, and no OEM, price or date is attached to it. Anyone planning around on-glasses inference should treat AR1 Gen 1 support as a capability demonstrated on a reference platform, not a shipping product.

TechCrunch also reported on September 17 that PrismML is rumored to be in talks with Apple, and that Hassibi declined to comment on it. That is an unconfirmed report, not a disclosed deal.

Why it matters if you ship on-device

The number worth watching is not 4x — it is 98%. Aggressive quantization has never been hard to find; quantization that holds benchmark scores has been. If a small vision-language model keeps its accuracy at this ratio on wearable-class silicon, the design question moves from whether local inference is possible to which turns still justify a round trip to the cloud.

Hassibi told TechCrunch the next releases, hoped for within a couple of months, will be in the several-hundred-billion-parameter range, and that he expects larger models to be easier to compress without losing intelligence. PrismML is not alone here: TechCrunch names Multiverse Computing, founded by a professor from Spain's Donostia International Physics Center, as another company working on LLM compression.

What to watch

  • A named glasses OEM. Until one ships, AR1 Gen 1 support is a reference build.
  • Per-task benchmarks for the 2B vision model. The 98% figure belongs to Bonsai 2 27B, not to the glasses build.
  • The several-hundred-billion-parameter release, and whether the retention claim holds at that size.

More from DangMua