AI agent safety · technical architecture

Beyond Code Generation: Building a Deterministic Control Plane for Autonomous Agents

Autonomous agents can propose code, commands and infrastructure changes faster than a human can review them. The missing layer is not another critic model. It is a deterministic gate that binds the exact proposal to explicit policy, bounded execution and a reproducible decision receipt.

EvidenceBound Research · canonical edition

A capable model is not an execution control

Prompt instructions can guide an agent, but they do not create an enforceable runtime boundary. A model may misunderstand a rule, choose a semantically different branch, emit an unsafe import, exceed a resource budget or produce an explanation that sounds compliant while the generated action is not.

Self-critique has the same structural weakness: the system asking whether an action is safe is still probabilistic. A second model can improve review, but it cannot prove that the exact candidate satisfied an exact execution contract.

The verifier sits between proposal and authority

EvidenceBound treats the agent as a proposal generator and deterministic verification as a separate authority for execution eligibility. The model may generate a candidate, explain it and revise it after rejection. It cannot manufacture the final verdict or authorize promotion.

AI agent proposal
      ↓
canonical source + declared evidence
      ↓
versioned policy contract
      ↓
bounded deterministic verification
      ↓
VERIFIED / REFUTED / BLOCKED / UNKNOWN
      ↓
content-addressed receipt
      ↓
human review
Public examples are controlled evidence. They do not accept arbitrary confidential code and they do not prove a private deployment boundary that has not been independently accepted on the target host.

Why canonical structure matters

Source text alone is a weak policy boundary. Formatting, aliases and equivalent syntax can obscure what a candidate actually does. A verifier can normalize allowed representations and bind them to the request, evidence set and execution artifact, making drift inspectable.

Bounded execution closes a second gap

Static checks cannot prove every runtime property. A candidate that passes structural policy may still consume excessive time, memory or output. Boundary violations therefore become explicit negative results rather than ambiguous success.

A receipt is more useful than a screenshot

A screenshot records what a UI displayed. A bounded decision receipt records the relationships between exact request, evidence, policy, candidate, result and verdict. If a retained artifact changes later, content identity exposes the mismatch instead of silently reusing an old conclusion.

What the system proves — and what it does not

A VERIFIED result means the exact declared obligations completed for the exact retained inputs and policy. It does not mean the model is generally trustworthy, that future inputs are safe, that a deployment is authorized, or that an organization is compliant.