Don't Ship Autonomy Without Evals
Back to BlogAI Governance

Don't Ship Autonomy Without Evals

Nuno Lopes
Nuno Lopes
CEO
September 18, 2026•9 min read

Enterprises are not waiting for better models. They are waiting for enough trust to let an agent act. Deloitte's 2026 State of AI survey found that only one in five companies has a mature governance model for autonomous agents — while agentic usage is set to rise sharply. Gartner forecasts that by 2029, at least 70% of organisations running production agents in infrastructure and operations will face a material incident linked to weak runtime controls. The bottleneck moved. Capability is abundant. Control is scarce.

Why written policy is not enough

Traditional software governance assumes deterministic code and human decision-makers. Agents break both assumptions. They propose multi-step plans, call tools at machine speed, and reason under uncertainty. A PDF of corporate policy cannot stop a destructive refund, a runaway purchase, or a token burn that empties the budget overnight. Runtime controls have to live beneath the model — in the harness that evaluates, authorises, and sometimes refuses what the model wants to do.

Evals as release gates

The cheapest place to catch failure is before deploy. An eval suite is a fixed set of cases that every model, prompt, or harness change must pass: task success, groundedness against your data, policy adherence, jailbreak resistance, and escalation rate when the agent should ask a human. Report cost and latency beside accuracy — otherwise teams will buy high scores with expensive models and never notice.

Treat the suite like a test suite. Wire it into CI. Block the release if the gate fails. When a change improves one metric and tanks another, that is a product decision, not a surprise. Memory systems deserve the same discipline: new benchmarks such as DolphinBench and AgentMemBench measure whether the past actually changes the next action — not whether a retriever can find a related string.

Runtime: least privilege and circuit breakers

At runtime, separate reasoning from execution. The agent may propose; a control layer decides. Give tools least privilege, budget caps, and human approval for irreversible or high-stakes actions. Keep a circuit breaker that trips on anomaly — spend, error rate, or policy violations — and fails closed. Trace every request with OpenTelemetry so you can answer "what happened?" without reading chat logs by hand.

Identity matters here. Agents that act need principals, not shared API keys. OAuth discovery, protected resource metadata, and signed agent cards turn "who did this?" into an auditable fact. Without that, every incident becomes a post-mortem of shared credentials and guesswork.

Human-in-the-loop is a design problem

Escalation is not a checkbox. It is an interface: show the plan, the evidence, the risk, and a clear Approve / Edit / Reject. Progressive autonomy — start narrow, earn scope — works only if the UI makes the trust gradient visible. That is why AG-UI and careful approval patterns belong in the same conversation as evals. Governance without legibility is theatre; legibility without gates is a pretty dashboard over an uncontrolled agent.

A stack you can ship

Start with twenty golden cases that represent the jobs you care about and the failures you fear. Add policy cases and adversarial prompts. Put them in CI. Wrap tool execution with budgets and approvals. Instrument traces. Red-team before you widen autonomy. Map the same controls to EU AI Act and ISO/IEC 42001 language so legal and engineering share a vocabulary.

Autonomy is only as good as the controls beneath it. If you are ready to turn demos into systems that can say yes safely, we design and ship that stack — evals, guardrails, and the interfaces that make them usable. Book a consultation at calendly.com/lopezi/.

Share:

Ready to ship agents you can trust?

Let's map one workflow worth handing to an agent — and what it would take to govern it.

Schedule a consultation