Agent runtime enclosed by a policy boundary, with an approved service path and a blocked route.

AI Agent Sandbox Security Testing: What Must Be Proven

A CTO is choosing a runtime for an agent that reads project files, calls a model, and updates an internal service. The vendor demonstrates an isolated workspace and blocked internet access. Engineering wants approval to mount a larger document directory and connect the agent to production. Procurement asks whether the sandbox has been security tested. The demonstration cannot yet answer that question for this deployment.

AI agent sandbox security testing should establish which files, processes, destinations, and credentials the deployed runtime can actually reach, including when the agent attempts something outside its assigned task. The business decision concerns the authority you are granting and the evidence that contains it.

Recent containment-platform announcements make this a timely purchasing question. However, a supplier capability, a configuration file, and a passing test are different kinds of evidence. Before approving broader access, require proof of the boundary your organization will operate, together with a clear account of what remains outside it.


Why sandbox selection now needs deployment evidence

On September 28, 2026, NVIDIA announced its Open Agent Safety Platform, describing OpenShell runtime policy enforcement and Sentry out-of-band monitoring on BlueField-4 DPUs. These are vendor-described capabilities. Their presence in an announcement does not establish the controls available in your selected deployment, their configured behavior, or their effectiveness against your business risks.

Separately, OpenAI published its model-misalignment reporting framework on September 16. Its initial reports concern individual observations during training or evaluation; OpenAI explicitly cautions against interpreting them as a measure of how often misalignment occurs. They are not evidence that OpenShell, another sandbox, or your agent has failed.

The practical inference for buyers is narrower: model behavior and oversight assumptions deserve observable tests. Containment should continue enforcing the approved boundary even when a model attempts an inappropriate operation. A cooperative demonstration alone does not exercise that condition.

Useful purchasing triggers include a new agent runtime, additional writable mounts, an expanded destination list, a new credential broker, or migration to shared infrastructure. Each can change effective authority without changing the product’s visible workflow. Commission focused validation when an existing assessment no longer represents those permissions.

Ask vendors to distinguish available features, required dependencies, supported deployment modes, and features your team must enable. This prevents a purchasing decision based on the strongest possible architecture when the actual rollout uses a smaller subset of its protections.


What AI agent sandbox security testing must prove

A useful assessment starts with a written containment claim. For example: this agent may read the current project, write only its task workspace, contact two approved services, and use a task-scoped identity. It may not read neighboring workspaces, change runtime policy, reach administrative services, or obtain reusable production secrets. Replace those examples with your actual workflow and risk tolerance.

The test must distinguish three outcomes: an operation succeeds within policy, an operation is independently denied, or the result remains unverified. A model refusing a request can be helpful behavioral evidence, but it does not prove that the runtime would block the same operation if attempted. An absent error message is also insufficient evidence of success or denial.

Put those claims into an independent AI penetration testing scope. Name the enforcement point for each restriction and the records needed to corroborate it. Agree whether testing begins from the agent’s ordinary tool interface, an authorized process inside the sandbox, or both. These positions answer different questions about workflow controls and runtime containment.

Successful validation needs permitted-workflow checks alongside prohibited-action checks. A policy that blocks everything can appear secure while making the product unusable. The assessor should show that legitimate tasks still complete under the same identity and configuration used for the denied tests.

Keep the assurance statement bounded. Record the runtime release, image, host or compute driver, effective policy revision, credential configuration, and tested integration paths. State material exclusions, including inaccessible supplier internals. The result supports a decision about that deployment; it does not prove that every future agent, operating system, or tool combination is contained.

Finally, separate runtime isolation from application correctness. The sandbox may enforce every local restriction while the approved service accepts an unauthorized customer-record change. That residual risk belongs in the decision record and, where material, in a connected application assessment.


Choose an architecture by its enforcement boundary

Architecture names are useful starting points, but buyers should compare the actual separation and authority they provide. Containers, virtual machines, policy proxies, and application permissions can coexist. None automatically answers every containment question.

Control approachBoundary to examineQuestion still requiring evidencePlanning implication
Agent instructions and tool permissionsWhat the application offers and authorizesCan a process or alternate integration route perform the same prohibited action?Include application and runtime owners; prompt changes alone may leave the authority intact.
Container with hardened runtime policyNamespaces, host integration, mounts, process identity, and additional enforcementDo shared-kernel exposure and mounted resources fit the threat model?Review host dependencies and deployment exceptions before treating a standard image as a baseline.
VM or microVM-backed executionGuest-to-host separation and guest configurationCan guest credentials, approved APIs, or attached storage still enable an unacceptable action?Include guest maintenance, image lifecycle, capacity, and operational visibility.
Brokered egress and credentialsDestination, request, and identity mediationAre alternate routes excluded, and are allowed requests limited to the intended operation?Integration work grows with protocols, service identities, and endpoint exceptions.
Managed agent runtimeSupplier controls plus customer-controlled integrationsWhich claims can be tested directly, and which depend on supplier evidence?Agree audit access, configuration visibility, and shared responsibilities before contracting.

This comparison is a scoping aid, not a universal ranking. A carefully restricted container may suit a bounded internal task. A stronger guest boundary may be appropriate where the organization expects untrusted code execution and higher isolation requirements. Both can still expose sensitive data through legitimately allowed services.

The buyer should ask where policy decisions happen, who can change them, and which components are trusted to enforce them. Include administrators, gateways, image builders, and credential services. A runtime boundary is less useful if the agent can modify its own enforcement configuration or use a mounted management interface to obtain broader authority.

For agents operating browsers and business applications, pair runtime validation with computer-use agent launch controls. That guide addresses the separate questions of action authorization and meaningful approval. This article’s scope is the runtime policy beneath those workflows.

Runtime, policy mediator, credential broker and downstream authorization shown as separate security boundaries.

Build an AI agent sandbox security testing control matrix

Convert the chosen architecture into acceptance criteria before testing begins. The following matrix proposes buyer requirements; it does not report findings from a Pentest Testing Corp engagement. Specify the actual data classes, services, identities, and enforcement mechanisms for your deployment.

BoundaryWhat the assessment should establishEvidence to retainApproval concern
FilesystemOnly approved task data is readable; writes stay within authorized locations; neighboring workspaces remain separate.Effective mounts and permissions, synthetic-file results, identity, and policy revision.Access to unrelated customer data, host secrets, or writable enforcement configuration.
Process and host integrationPermitted programs operate under the intended identity without unintended management or neighboring-process access.Resolved identity, runtime restrictions, host interfaces, and controlled denied-operation results.Authority to alter the host, inspect protected processes, or control other workloads.
Network egressRequired destinations work; prohibited external, internal, and management destinations are blocked across relevant routes.Effective destination rules, resolver/proxy coverage, denial records, and controlled destination receipts.A route around the approved mediator or unnecessarily broad destination access.
Request policyWhere claimed, allowed endpoints accept only approved methods and operations, with blocking enabled.Inspection mode, allowed/denied request pairs, and service-side outcomes.Rules only log violations, or a broad allowed host permits unintended operations.
CredentialsSecrets stay outside untrusted task content where designed; brokered access is limited to approved identities and destinations.Credential flow, synthetic-secret exposure checks, binding decisions, and downstream permissions.Reusable privileged credentials or a broker capable of exercising unnecessary authority.
Policy lifecycleOnly authorized operators broaden access; revised policy reaches the intended runtime; unsupported enforcement is visible.Change history, approval identity, effective revision, and startup or degraded-state observations.Silent exceptions, stale restrictions, or a workload able to authorize its own expansion.

Filesystem and process restrictions protect different assets

Read-only access protects integrity, but it does not protect confidentiality. If an agent can read an entire document archive, it may process information unrelated to its task without modifying anything. Decide which records the workflow needs, then test separation using synthetic documents representing both permitted and prohibited areas.

Process scope should include tools the agent can invoke and the interfaces exposed by deployment choices. A non-root identity is useful, but its mounted-volume permissions and service authority still matter. Ask whether a child process remains subject to the intended boundary and whether management interfaces are reachable from the workload.

Egress permission is not permission for every business action

Allowing a service hostname can expose more operations than a task needs. A collaboration service may support reading documents, uploading files, and sending communications. Test the control level actually claimed: destination restriction, inspected request restriction, and downstream business authorization are different boundaries.

Include legitimate intermediaries and all relevant supported traffic paths in the scope. DNS resolution, proxies, package services, and observability destinations deserve review because they can carry data or expand reach. A denied public connection says little about a separately permitted internal service.

What OpenShell buyers should verify specifically

NVIDIA’s OpenShell best-practices documentation describes filesystem, process, network, and provider-credential controls. It distinguishes static restrictions from controls that can change during operation. Verify the deployed version and driver rather than assuming every documented capability applies unchanged.

The policy schema distinguishes request audit from enforcement: audit records violations while allowing the request. It also separates destination admission from protocol inspection. A network policy therefore needs evidence of the effective inspection and blocking mode, not simply a permitted-host list.

Filesystem compatibility needs similar care. The documented mandatory baseline differs from additional configured restrictions that may not fully apply under best-effort compatibility. Require the startup outcome, applied restrictions, skipped paths, and any degraded-enforcement findings. Do not describe this as loss of every isolation layer, or accept it silently where the missing restriction protects sensitive data.

Credential mediation adds another boundary. Verify whether the workload sees an opaque placeholder or a reusable secret, which destinations may resolve it, and what the downstream identity can do. For the broader authority question, use cloud agent identity and delegated authority as companion scope.

Configured runtime policy compared with permitted and denied test results tied to an effective revision.

Hypothetical scenario: isolation works, but the allowed path is too broad

Consider a fictional professional-services platform evaluating an agent that prepares project summaries. It receives a task workspace and access to a collaboration API. The intended business policy allows it to read current-project documents and save a draft for internal review.

In this illustrative staging assessment, filesystem separation holds: synthetic files in another workspace remain inaccessible. Prohibited external destinations are also denied. However, the collaboration identity includes an upload operation that the workflow does not need. A harmless synthetic document can be transferred through that permitted service without the intended review step.

This does not establish a sandbox escape. It establishes a mismatch between the business policy and the authority available through an allowed path. The consequence in a comparable production configuration could be unintended disclosure to users or workspaces reachable through that service, depending on its permissions. No real client, incident, or measured loss is implied.

Remediation should address the narrowest demonstrated cause: remove unnecessary operations, restrict the downstream identity, and enforce any required review at the action boundary. If sensitive material should never enter the agent’s context, reduce its readable data as well. A new model instruction would leave the unnecessary API authority available.

The retest should prove that the unintended transfer fails and the approved summary workflow still succeeds. Leadership can then distinguish a working local isolation layer from a corrected application permission gap, rather than paying for a runtime replacement that does not address the cause.


Define the assessment scope, evidence and effort

Provide the assessor with the architecture diagram, representative workflow, runtime and image versions, deployment driver, effective policy, mounts, destination list, and credential flow. Include the people who own the agent, infrastructure, application permissions, and supplier relationship. These inputs reduce time spent resolving what the control was intended to achieve.

Agree on testing positions and access explicitly. A configuration review can identify broad permissions, but cannot establish every operating result. Testing solely through the product interface may miss runtime paths that the model never attempts. An authorized in-sandbox starting position can examine those paths, while clearly recording the assumed access and avoiding claims about how an attacker would obtain it.

Use representative staging, synthetic records, harmless operations, and controlled destinations wherever possible. Document differences from production before making an approval recommendation. Any production validation should target the specific dependency that staging cannot reproduce, with authorized limits and an owner able to stop the exercise.

For internal reachability beyond the sandbox’s egress boundary, commission segmentation testing from the AI workload. This extension answers whether permitted network paths expose management systems or other sensitive zones. It should be separately named when a runtime-only assessment cannot cover the wider infrastructure.

Evidence should support the decision and the retest

Request an allowed-and-denied operation matrix tied to the tested policy revision. Each material result should identify the initiating workflow or process, effective identity, attempted operation, relevant control, observed outcome, and corroborating record. Separate demonstrated failures from configuration concerns, unavailable evidence, and accepted exceptions.

Report limitations in ordinary business terms. If the supplier does not expose its host configuration, the assessment cannot independently verify that host boundary. If downstream records are unavailable, a client-side denial may support a narrower conclusion than an end-to-end outcome. Procurement can then seek additional evidence or reduce the authority granted.

Effort follows boundaries and integration complexity

A single disposable workspace with limited destinations is simpler to assess than multiple runtime types, shared volumes, several credential brokers, and custom network protocols. Policy review, controlled validation, reporting, remediation, and retest each need time. Confirm a schedule after checking environment readiness and evidence access.

Remediation effort also varies. Narrowing an unnecessary mount may be straightforward. Separating a shared identity or introducing operation-level enforcement can require application changes and integration testing. Budget for those dependencies rather than treating every containment issue as a configuration toggle. Request a scope-based quote; this article does not supply a fixed runtime-testing price or delivery promise.


What leadership should decide before approval

Approve a specific operating envelope: readable data, writable locations, callable services, permitted operations, credentials, and policy-change authority. Name the owner who maintains it and the changes that trigger focused retesting. A broader mount, new provider binding, runtime migration, or access exception should prompt review of the affected boundary.

Hold expansion when testing demonstrates an unacceptable path to sensitive data, stronger authority, or a prohibited business action. Missing evidence for a material claim should remain unresolved. A restricted pilot may be reasonable when compensating controls demonstrably reduce exposure, but give each exception an owner, expiry date, and measurable condition for removal.

Acceptance should include usability. Record which normal workflows passed under the restricted configuration and which require redesign. This helps leadership compare the cost of a narrower workflow with the cost and exposure of broader authority.

Map the evidence to risk management

OWASP LLM01:2025 Prompt Injection is relevant when untrusted content influences an agent’s attempted action. Runtime containment can limit consequences without proving that the model always rejects such instructions.

OWASP LLM06:2025 Excessive Agency addresses excessive functionality, permissions, or autonomy. Apply that mapping where unnecessary tools or authority enable the demonstrated business risk. A broad operating-system permission alone should not automatically be presented as a model-behavior finding.

The NIST AI RMF supplies a voluntary governance structure through GOVERN, MAP, MEASURE, and MANAGE. An application here is to assign ownership, describe intended authority, measure boundary performance, and decide treatment or residual-risk acceptance. NIST states that revision of AI RMF 1.0 is in progress; retain the edition used in the assessment.

NIST SP 800-115 provides established guidance for planning security tests, examining results, and developing mitigation strategies. It is not an AI-specific certification. These mappings organize evidence and responsibilities; they do not establish universal compliance or replace the organization’s own acceptance criteria.

Keep residual application and consumption risks visible. Even a correctly contained agent can use approved services too frequently or make an incorrect permitted change. The companion guide to agentic API abuse and cost controls addresses spending and capacity boundaries that runtime isolation alone does not establish.

Leadership decision path from tested claims to obtaining evidence, remediation and retest, or approval of defined access.

Before buying containment controls or expanding production access, scope a runtime-containment assessment around your deployed architecture. Bring the policy, integrations, identities, and approval criteria. The useful outcome is a defensible decision about which boundaries hold, which require remediation, and what authority the organization can safely approve.


Frequently asked questions

Can we assess a runtime before committing to a vendor?

Yes. A proof-of-concept assessment can compare candidate architectures using the same synthetic workflow and acceptance criteria. Record differences in dependencies and visibility. A later production configuration still needs validation if its identity, mounts, policy, or integrations differ materially.

What if a supplier will not permit active security testing?

Review its testing policy and request architecture-specific evidence, configuration access, and supported assurance options. Test customer-controlled integrations within authorization. Document the supplier boundary as unverified where evidence is insufficient, then decide whether narrower access or another supplier is necessary.

Does a sandbox test cover malicious packages or agent skills?

It can evaluate the authority available to components running inside the boundary if they are named in scope. Package provenance, integrity, update controls, and dependency review are separate supply-chain questions. A contained component may still misuse allowed capabilities.

Can existing pentest results reduce the assessment effort?

They can, when their tested architecture, versions, identities, and evidence remain applicable. Ask the assessor to identify reusable results and remaining gaps. A web or API report should not be assumed to include filesystem, process, or runtime-policy validation.

Should we give every tenant a separate sandbox?

That depends on data sensitivity, workload behavior, isolation requirements, and operating cost. Evaluate dedicated and shared designs against cross-tenant access and lifecycle criteria. Separate sandboxes do not remove the need for tenant authorization in shared downstream services.

How should procurement word an assurance requirement?

Ask for evidence against named containment claims for the proposed deployment, including permitted and denied results, configuration versions, limitations, remediation status, and retest outcomes. Avoid requiring an undefined “AI-safe” certificate that neither party can explain or verify.

Can test evidence itself expose sensitive information?

Yes. Prompts, traces, files, and network records may contain business data or credentials. Prefer synthetic data, redact evidence appropriately, restrict access, and agree on retention before testing. Preserve enough context to substantiate results without collecting unnecessary production content.


Leave a Comment

Scroll to Top
Pentest_Testing_Corp_Logo
Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.