Certanum  /  Research note

Runtime control is not numerical evidence.

Agent infrastructure is moving enforcement outside the model. On 28 September 2026 NVIDIA announced its Open Agent Safety Platform, with OpenShell, an open-source secure runtime that NVIDIA says it has been building over the past year, at its core. OpenShell controls what an AI agent can access and do. This note sets out where runtime control ends, and where the evidence for a consequential number begins.

Research note  ·  1 October 2026  ·  Certanum Technologies Inc.

01What OpenShell does

A secure runtime boundary, enforced outside the agent.

In NVIDIA’s own words:

“an open source secure runtime for executing autonomous AI agents in sandboxed environments with kernel-level isolation”

NVIDIA Technical Blog, 28 Sep 2026 ↗

“Operators define which files, networks, tools, processes, and credentials an agent can access. OpenShell checks those limits before the agent runs and enforces them as it works.”

NVIDIA Technical Blog ↗

“a prover shows that its policy cannot escape the intent of the operator”

NVIDIA Technical Blog ↗

NVIDIA Sentry, its hardware counterpart, “correlates agent interactions, policy decisions, and tool and data access to create a contextual record of agent activity.”

NVIDIA Technical Blog ↗
02The shared principle

Do not ask a probabilistic model to guarantee what deterministic infrastructure can guarantee from outside it.

OpenShell applies that principle to access and behaviour: which files, networks, tools, processes and credentials an agent may touch. Certanum applies it to consequential numbers: whether a figure was computed from authoritative source data, by a declared operation, and can be reproduced.

The two are complementary. Neither replaces the other, and neither depends on the other: the numerical question exists whichever runtime an agent runs in.

03Two questions

Permission answers whether an agent may perform an action. Numerical admissibility asks whether a consequential figure has the evidence to be released.

Runtime controle.g. NVIDIA OpenShell
Numerical evidenceAP-1 · Certanum
Question it answers
May the agent take this action?
Does this figure have the evidence to be released?
Operates on
The execution environment: files, networks, tools, processes and credentials, under policy the operator defines
The operands of a consequential figure, the declared operation, and the result
Where it sits
Outside the agent, before and during execution
Outside the model, between computation and release
What it records
Agent activity: actions, policy decisions, tool and data access
An evidence record: each operand’s source, the declared operation, and the re-execution result
When the check fails
Policy enforcement denies or constrains the action, according to the configured controls
The figure is withheld

The difference is the property each layer establishes, not how much a runtime can inspect: runtimes can be extended to examine execution in more detail. Runtime-control descriptions summarise NVIDIA’s published materials linked below; they are not an assessment of OpenShell.

04Architecture

Two layers of external control.

The runtime layer governs the action. The numerical layer governs the figure the action produces.

Two layers of external control for AI agents An AI agent sits above a runtime and policy layer, for example NVIDIA OpenShell, which controls files, networks, tools, processes and credentials and asks whether the agent may take an action. A permitted tool invocation passes operands 82.4 and 1.78 into a numerical evidence layer, which checks operand sources, the declared operation and re-execution, and either releases the figure 26.0 with its evidence record or withholds it. AI AGENT · ANY MODEL RUNTIME AND POLICY LAYER e.g. NVIDIA OpenShell · enforced outside the agent Files Networks Tools Processes Credentials May the agent take this action? permitted call tool invocation · BMI.v1 operands 82.4 · 1.78 NUMERICAL EVIDENCE LAYER AP-1 · Certanum (in development) · outside the model Operand source Declared operation Re-execution Release / withhold Does this figure have the evidence to be released? withheld 26.0 + evidence record CONSEQUENTIAL USE
Illustrative architecture with synthetic values. The runtime layer is shown generically; OpenShell is one example. Certanum is in development and is not integrated with OpenShell.
05A worked case

A permitted action can still carry an unsupported number.

An agent is permitted to call a calculation tool. The call is inside policy and recorded. The tool executes correctly. But one operand it passes is 84.2, while the source record holds 82.4.

The runtime layer may have permitted, bounded and recorded the action. That does not by itself establish that the value supplied to the calculation matches an authoritative source value. Establishing that relationship requires additional evidence about the operand and its source: a question about the number, not the action.

The numerical layer resolves each operand against its source, finds the mismatch, and withholds the figure. See the evidence record →

06An open question

Could numerical evidence attach to a runtime-governed execution without changing the runtime’s security model?

NVIDIA notes that, as open-source software, OpenShell “can also be extended”. A natural attachment point for numerical evidence is the record of what actually executed: an evidence record keyed to the execution it describes, referencing the decision that permitted it, rather than a change to the policy layer itself.

This is an open engineering question, not an existing integration. The same composition, provenance bound on the execution side, is under discussion in the FINOS AI Governance Framework (#359) ↗.

What this note does not claim

Certanum is not affiliated with, endorsed by, or partnered with NVIDIA. OpenShell, Sentry and NVIDIA are referred to only to describe their published capabilities. This note does not suggest that OpenShell lacks anything it was designed to do; it describes a different question. Certanum’s commercial infrastructure is in development, and no independent evaluation of AP-1 has been completed.