2026-10-09 11:46 UTC
DANGMUAAI & Developer Tools, Decoded
BackDev Tools

Argent vs agent-device: Same Fixes, One-Third the Cost

Both React Native agent tools scored 12/12 on the same planted bugs. The gap was $1.27 versus $0.41 per run, and which one needed setup first.

DangMua EditorialOct 09, 20264 min read
Argent vs agent-device: Same Fixes, One-Third the Cost

Two React Native agent tools scored 12 out of 12 on the same planted bugs. One burned $1.27 a run; the other $0.41.

A developer planted three bugs in a React Native 0.87 app and handed the same bug reports to Claude Code twice — once with only Software Mansion's Argent, once with only Callstack's agent-device. Three tasks, two runs per tool per task, 12 scored runs. The app, the harness, the rubric, every transcript and every score sit in a public repo.

The bench: the store-audit app on React Native 0.87.1 with FlashList, in the iPhone 17 simulator on an Apple M1 with 8 GB of memory. Argent 0.26.0 against agent-device 0.21.19, driving Claude Code 2.1.42 in headless mode with claude-sonnet-5-5, one session per run.

Both tools solved everything

TaskArgentagent-device
Reproduce6 / 66 / 6
Fix and verify6 / 66 / 6
Performance6 / 66 / 6

Every run reproduced or verified the bug in the running app, named the right cause, and produced fixes that type-checked. Two runs went past the report: one Argent run noticed a recycled-row bug while fixing a layout bug, and one agent-device run rewrote a slow keyword check after confirming the new version gave identical results on all 48 questions. On work at this level, the tool is not what decides whether the agent succeeds.

The gap is context, not capability

Median per runArgentagent-device
Time138 s108 s
Tool calls3416.5
Screenshots in conversation110.5
Text returned by tools59 KB17 KB
Tokens processed1.84 M0.34 M
Cost per run (average)$1.27$0.41

The split comes from what each tool hands back. Every Argent interaction — tap, swipe, launch — returns a screenshot plus the full accessibility tree, and every image stays in the conversation. agent-device returns a compact text snapshot or a diff with short handles like @e8 to tap, and screenshots only when the agent asks. Across a session that is roughly 1.8 million tokens against 0.34 million — about a fifth of the tokens and a third of the money, by the author's measurement.

Text-first paid off in one concrete way here: agent-device's snapshot printed a pill's state in the line itself, as @e29 [button] "yes for Entrance mats clean" [selected], so the agent read the selection without spending an image on it.

Setup cuts the other way

Argent's React profiler worked with no setup on React Native 0.87. agent-device's profiler needed a package installed and two project files patched first. If you are picking for a one-off investigation rather than a standing workflow, that difference may outweigh the token bill.

Two operational notes from the same run book: you cannot start Claude Code from inside Claude Code — claude -p from an agent session fails with “Claude Code cannot be launched inside another Claude Code session” — and the command-line tool needs its own claude auth login; being signed in to the desktop app is not enough.

Why “it compiles” is not the bar

The rubric was committed before the first scored run, and it withheld points that need evidence from the running app whenever the agent had only read the code. A separate experiment shows why that guard earns its place. Given Claude Code on Haiku 5.5 a task whose test suite needed a database the VM could not run, the agent's own reasoning read: “Running the test suite likely won't work since the DB on localhost:5433 probably isn't up … worth flagging, but let me make the edit first.” It tried two substitutes and raised the problem only at the end. That author puts the cost of a checker model watching a whole session at under two cents.

What to watch

Price the screenshot. If your agent runs are measured in dollars rather than cents, the question is not which tool is smarter but how many images your transcript is carrying. Benchmark your own app before committing: this is one developer, one 0.87 app, 12 runs, and a rubric its author wrote themselves.

More from DangMua