2026-10-04 18:28 UTC
DANGMUAAI & Developer Tools, Decoded
BackAgents

Four Controls That Turn an AI Agent From Demo to System

Two sponsored 2026 agent guides, product names stripped out, leave the same four controls — and four questions you can make any vendor demonstrate live.

DangMua EditorialOct 04, 20264 min read
Four Controls That Turn an AI Agent From Demo to System

A chatbot answers. An agent acts — it searches a document, fills in a form, drafts the reply and sends it. That difference is why an agent needs rules before it needs features, and it is the organizing idea behind a dev.to post that strips the product names out of two 2026 Sifted guides on AI agents, one sponsored by Box and one by Salesforce, to see what survives.

Four controls appear in both. Each is framed as a question you can put to a vendor — including, the author notes, to their own agency.

1. Scoped access, time-bounded, fully logged

The sharpest version comes from Box's chief information security officer, Heather Ceylan: agents should never get broad standing access just because they are useful. Permissions should be scoped to a task, time-bounded where possible, and fully logged.

The question to ask: what exactly can the agent read and change, and who can see what it touched? The post is blunt about the failure case — if the answer is "everything", that is not an answer.

2. A second check before anything ships

The guides say plainly that agents make mistakes; one founder quoted in them describes agents as "well-intentioned, slightly forgetful children". Every team that got past the pilot stage added a second step: either a person approves before anything leaves the building, or a separate evaluator checks the first agent's work independently and scores its confidence.

The post cites one London company in the Box report, Deliverance, that builds this as a pair of agents — an executor that does the work and an evaluator that critiques it. Its founder's line is the one worth borrowing in a procurement meeting: "It's not magic. It's a governed system."

The question to ask: what checks the output before a customer sees it?

3. Written rules for when it does not know

A reliable agent knows when not to act, and the post argues that matters more than being right most of the time — an agent that recognises its limits does far less damage than one that is merely usually correct.

The rules get written in advance: if it cannot find the source, it says so; if confidence is low, it asks a human; if money or an upset customer is involved, it escalates. Then you test exactly those cases, because the tidy requests almost always work and the oddly phrased ones are where the failures live.

The question to ask: show me what it does with a request it cannot answer.

4. A log of every step

For each run: what the agent received, what it used, what it produced. When something goes wrong, that log tells you whether the input was bad, the instructions were unclear or a tool failed — instead of leaving you to guess. It is also what lets you answer, months later, where a particular answer came from.

The question to ask: if I pick any answer from last week, can you show me its sources?

What to do with this list

Two caveats are worth carrying. Both underlying guides are sponsored content — one by Box, one by Salesforce — and the dev.to author runs an agency selling exactly the kind of assistants the four controls describe, which they disclose in the piece. That does not make the controls wrong; it does mean the list is shaped by vendors who already sell the governance layer.

Our read: the four questions are more useful than the four controls. A control is something a vendor can claim in a deck. A question like "show me what it does with a request it cannot answer" has to be demonstrated live, on your data, and the demo either happens or it does not. Ask all four in the same meeting and the gap between a product and a pilot usually closes itself.

More from DangMua