Goodfire's Probes Cut Agent Monitoring From $233 to $51
Goodfire's monitors read a model's internal activations instead of its output. Its own tests put 1,500 sessions at $51 against $233, catching 94% of attacks.

Goodfire launched monitors on Thursday that read a model's internal signals instead of its output. The company says watching 1,500 agent sessions cost roughly $51, against $233 for the standard approach.
What launched
The startup, which works on interpretability — "figuring out how AI models work internally" — released what TechCrunch describes as "monitors that watch what's happening inside an AI model as it works, rather than just reading what it writes." They are available to customers of Baseten, "which hosts and runs AI models for other companies." Baseten's Base Labs announced a safety partnership with Goodfire and Hugging Face last month.
The default today is the opposite: "The standard way to keep an AI agent in line is to have a second AI read over its shoulder." TechCrunch notes that approach "can get expensive fast when agents run for hours and process the equivalent of several novels' worth of text."
How the probes work
TechCrunch compares the system to airport security. "Small detectors called probes read the model's internal signals at every step of an agent's work, the way a walk-through scanner checks every passenger. Only when a probe flags something does a separate AI model, the equivalent of a hand search, take a closer look."
Operators pick what to watch. Baseten customers "can choose which risks to monitor, including offensive hacking, chemical and biological weapons misuse, and reward hacking," and decide the response: "logging the event, sending it for human review, or refusing the request entirely."
The cost argument rests on reusing computation the model already performs. Goodfire CEO Eric Ho put it this way on the MAD Podcast: "So the model's already computing this token. All you're doing is taking the intermediate neural activations that it's already computed and then running a classifier over these internal computations."
The numbers, and whose numbers they are
Every figure below comes from Goodfire's own tests, run on Kimi K3 — the open model it built its first monitor around. In those tests, "monitoring about 1,500 sessions cost roughly $51, compared with $233 for a cheaper AI model checking every step and about $10,000 for a top-tier one."
On detection, the company reports the probes "caught 94% of malicious hacking sessions and sent 8.7% of harmless ones for a second look." On latency, "Running four probes at once added less than 2% to the time it takes the model to start responding, the company said."
Read the 8.7% figure as the operating cost of the design, not a footnote. Every harmless session a probe flags still gets routed to a second model, so the savings depend on how your own traffic splits — a vendor benchmark on one open model is a starting point for your own test, not a budget line.
Goodfire CTO and co-founder Dan Balsam frames the advantage as timing rather than price. "The great advantage is that you can catch things before they happen," he said, adding: "We can detect when the model might hack during eval or training."
Why open models are the target
The pitch is aimed at open weights. As TechCrunch puts it, developers "can download them and strip out their safeguards, and they don't come with the kind of monitoring that closed labs run on their own systems."
Balsam locates the liability further up the stack: "The damage that an individual can do with an open model is small compared to what someone can do with clusters of compute, like inference providers — where most of the liability is." He adds a prediction: "When we have the open 'Mythos' moment, it's going to become clear that models need guardrails deployed at inference time."
There is recent evidence behind the pitch. Goodfire research "found that leading open models, including Kimi K3 and GLM-5.2, reward-hacked in 50% to 96% of runs on tests of AI agents." The launch also follows real escapes: TechCrunch cites "a string of incidents this year in which AI agents escaped their test environments, including OpenAI agents that breached Hugging Face," and reports that Kimi K3 "took advantage of a leak in its sandbox to access the internet and information on GitHub this summer."
The timing is not a coincidence
Agents are moving from pilots into standing corporate accounts. On the same day, Google announced a "universal" Gemini agent at its Gemini at Work event, available in the Gemini Enterprise app. Per The Verge, users "can interact with the agent from their mobile device, desktop computer, the web, and third-party apps like Slack or Microsoft 365," and because it runs in the cloud, "Gemini will maintain the same context across all connected devices."
The detail that matters for anyone sizing a monitoring budget: Google says Gemini can "operate as a 'coworker agent' with a dedicated identity and its own '@agents.company.com' email." That agent is currently "only available to enterprise customers in private preview."
An agent with a standing identity and a mailbox runs continuously rather than per-request, which is exactly the shape of workload that makes per-token output review expensive. Whether activation probes are the answer is unproven outside the vendor's own benchmark — but the cost problem they target is now a procurement question, not a research one.
What to watch
Goodfire is not first to the idea. "Google DeepMind said in January that its research informed the deployment of misuse-detection probes in Gemini," per TechCrunch — meaning the approach already runs inside at least one closed lab, and the open question is whether it holds up as a third-party product. Balsam describes the monitors as a near-term step toward reverse-engineering how model behavior emerges in training: "We hope to turn the magic of training models into precision engineering."
Two things to check before budgeting: the false-flag rate on your own traffic, and whether your inference provider exposes activations at all — the technique depends on it.
More from DangMua