Deterministic Control Plane for AI Agents
Why agent proposals need policy and verification outside probabilistic self-review.
EvidenceBound is an early-stage AI Safety infrastructure project for binding agent actions to evidence, provenance, executable policy, deterministic verification, bounded blast radius and recoverable human control.
The core pattern is deliberately narrow: an AI system may propose or reason, but authority is bounded by verifiable state outside the model process.
Bind an exact claim or action to exact source identity, provenance and retained artifacts.
Version the operator-owned obligations that decide what is allowed, blocked or requires human approval.
Keep deterministic checks and typed outcomes outside probabilistic model self-evaluation.
Scope what a verified task may affect instead of treating verification as universal deployment authority.
Interrupt, correct and resume with lineage rather than silently mutating history after a failure.
Preserve explicit review, approval and kill boundaries when the operator remains accountable.
EvidenceBound is being developed to make control evidence inspectable: which policy applied, what evidence was available, which deterministic checks ran, what failed, what was approved, and how recovery changed state.
Evidence artifacts can support selected GOVERN, MAP, MEASURE and MANAGE activities, especially traceability, TEVV evidence, decision records, incident response and recovery. The NIST AI RMF is voluntary and broader than this technical runtime.
Policy lineage, verification receipts, human approvals and corrective-action evidence can support an organization's AI management-system records. EvidenceBound is not ISO certification and cannot certify an organization.
Where legally relevant, retained logs, traceability, human-control evidence and bounded risk controls may support organizational work around high-risk AI obligations. Applicability depends on role, system classification and legal context.
EvidenceBound does not guarantee safety, legal compliance, model truthfulness, regulatory approval or fitness for a domain. It verifies bounded technical statements against declared evidence and controls.
Prompt rules and model self-critique can assist review, but they do not create a deterministic runtime boundary. EvidenceBound keeps executable policy, typed failure states and approval outside the model's ability to rewrite its own verdict.
Governance becomes inspectable when request identity, policy version, candidate identity, execution evidence, verdict and approval are bound at decision time instead of reconstructed from disconnected logs later.
The public OSS core contains executable conformance tests, signing and persistence paths, adversarial cases, graph recovery and correction evidence. Public demonstrations are evidence of implemented properties, not evidence of customer adoption.
Unsupported, blocked, stale or incomplete evidence must not be promoted into a successful result.
Cryptographic digests make changed artifacts and stale receipts detectable instead of silently reusable.
Recovery produces explicit lineage so operators can distinguish the original state from a corrected successor.
A frozen reproducible agent-safety benchmark tested EvidenceBound recovery-authority semantics on stock NVIDIA OpenShell v0.1.2 as an independent runtime enforcement substrate. The same runtime and non-idempotent target were exercised first as a negative control and then under fail-closed authority enforcement.
The negative control reproduced two consequences after an uncertain outcome and retry. Under frozen controlled conditions the benchmark returned FULL_PASS: 12/12 hard gates, 8/8 frozen scenarios, and no middleware-denied request ID appeared in target-contact evidence.
EvidenceBound demonstrated outcome-aware recovery authority and trusted operation-lineage enforcement at the OpenShell pre-effect HTTP boundary under the tested benchmark conditions.
Why agent proposals need policy and verification outside probabilistic self-review.
Binding policy, exact artifacts, execution and human approval into a reproducible evidence chain.
A sports reference case for separating algorithm conformance from real-world outcome claims.
SignalReview is a separate sports-intelligence product that exercises provenance visibility, deterministic evidence, explicit missing-data states and bounded AI reasoning. It is useful as an applied EvidenceBound case study without making EvidenceBound part of the sports product's commercial identity.
Creator and maintainer of EvidenceBound. Product and systems work focuses on verifiable agent behavior, evidence-bound decisions, deterministic control boundaries, human corrigibility, recovery semantics and production AI infrastructure.
EvidenceBound Core, human-control research, open conformance evidence, and applied reference implementations. Claims remain evidence-scoped: working code is not the same as a security certification or a production customer deployment.
Python, TypeScript, agent orchestration, policy-as-code, provenance, cryptographic receipts, cloud delivery, production QA and commercialization of trustworthy AI systems.