A 400MB Model Ran a Browser Agent on a 2017 Galaxy Note 8
Qwen3-0.6B scored 10/10 on a live browser task from a 2017 Note 8 and 0/3 without a perception layer — all figures self-reported by the layer's team.

A 400 MB model running on a 2017 Samsung Galaxy Note 8 passed a live browser-navigation task 10 out of 10 times — and 0 out of 3 without a perception layer in front of it.
The figures come from the team behind that layer, who state it up front: "we build this layer (e2llm, the SiFR format)." Treat every number below as a vendor self-report published alongside open logs, not independent verification.
The task that separated the two runs
The task starts on an unrelated site, navigates to the Wikipedia page for the Galaxy Note series, picks exactly "Note 8" among decoys — "Note 8.0", "Galaxy Note 8.0" and "Note FE" — then reads the release date from the infobox. Expected answer: 15 September 2017.
Fed raw HTML, the model drowns. The page runs to 466,744 characters; a 16k context fits roughly 40,000 of them, about 9%. The reported result is 0 out of 3, with every run returning the address of the page it was already on, at roughly 12,300 tokens and 35 to 44 minutes per attempt.
With the layer, the team reports the same model on the same phone hitting 10 out of 10 at a median 49 seconds per run — and picking "Note 8" every time even as node ids changed between runs, which they read as evidence it selects by text rather than a memorized path.
On an easier target the gap narrows but does not close. On the books.toscrape.com sandbox, raw HTML passes 3 out of 4 at about 12,200 tokens and 1,360 seconds per run; with the layer, 25 seconds and 10 out of 10 — framed as roughly 25 times fewer tokens and 50 times less time.
Old hardware is not the bottleneck
| Device | Year | Chip | RAM | Sec/run |
|---|---|---|---|---|
| Galaxy Note 8 | 2017 | Exynos 8895 | 5.2 GB | 25 |
| Galaxy S21 | 2021 | Exynos 2100 | 7 GB | 29 |
| Galaxy A04e | 2022 | Helio P35 | 2.7 GB | 62 |
All three run the same llama.cpp commit, the same model files and the same prompt. Qwen3-0.6B is reported as the only model scoring 10/10 across all three, including the A04e — a handset the write-up puts at about $100 with 2.7 GB of RAM.
The most useful number for anyone sizing edge hardware: four years of flagship silicon bought nothing, at 19 seconds of model time on both the Note 8 and the S21. Scale shows up with parameters instead — Ministral 3 3B takes 193 seconds on the Note 8 against 47 on the S21.
Parameter count does not predict the score
Fourteen models were run on the Note 8. Six, from five vendors, reportedly scored 10 out of 10: Qwen3-0.6B, Qwen2.5-1.5B, GLM-Edge-1.5B, Gemma-2-2B, Llama-3.2-3B and Ministral 3 3B. Qwen2.5-0.5B managed 6 out of 10 at nearly the same size as the 0.6B that scored perfectly, while Llama-3.2-1B, Gemma-3-1B and LFM2.5-1.2B scored zero — answering with a placeholder "ID" or clicking at random.
That pattern fits the architecture. The model never sees the page; it gets a short candidate list and returns one decision, such as {"target": "div036"}. Capture, click, paste and check are plain Python. Format compliance, not reasoning capacity, is what the scores measure.
What to check before believing it
The team's own caveat is the right starting point: "This isn't a benchmark. These are measurements on specific tasks." Results are per-device, and they do not claim the tasks are impossible without the layer — only cheaper, faster and repeatable with it.
A longer relay run — Ministral 3 3B on an S21 shuttling messages between Gemini in Chrome and Z.ai in Firefox — reports 10/10 runs and 40/40 hops, but also the costs: run times from 7:59 to 11:55, a second-half median 11.6% above the first half, battery from 79% to 63%, and temperature climbing 32.5 to 35.4 °C. That thermal drift is the number to watch if you plan sustained on-device agent work.
Code and logs sit at github.com/e2llm/edge-browser-agent, and the write-up points to a replay.py that reproduces the series from open JSONL with no phone or browser attached. Replaying those logs is the cheapest way to decide whether these numbers survive contact with your own hardware.
More from DangMua