Chapter 05 · Execution Safety

Sandbox & Execution

The first three layers decide whether a tool may run. This is layer 4: when an allowed tool actually runs code, it runs inside a box with no capabilities it wasn't granted — and its result is cryptographically attested. The earlier layers can be tricked by a clever prompt; this one can't, because it isn't asking the model anything. It's an operating-system-level cage with a signed receipt.

Two boxes: process and container

"Sandbox" means a confined execution environment: the code runs, but the walls of the box decide what it can see and do. The harness ships two boxes, picked by how strong the isolation needs to be.

Process sandbox
ProcessSandboxExecutor runs the tool as an isolated Windows subprocess, caged by a kernel-level Job Object (covered next). Lighter weight; no container runtime required. File: src/Content/Infrastructure/Infrastructure.AI/Sandbox/ProcessSandboxExecutor.cs.
Docker sandbox
DockerSandboxExecutor runs the tool inside a Docker container (driven through the Docker.DotNet client). Stronger isolation — a whole container boundary between the untrusted code and the host. File: src/Content/Infrastructure/Infrastructure.AI/Sandbox/DockerSandboxExecutor.cs.

The Docker executor enforces a minimum-isolation invariant. When a request's permission profile requires MinimumIsolation = Container and Docker is unavailable at runtime, the execution fails rather than quietly falling back to the weaker process sandbox.

No silent downgrade

A common failure pattern: a control "degrades gracefully" when its strong path is unavailable, and nobody notices the box got weaker. The Docker executor refuses to do this. If you asked for container isolation and the container runtime isn't there, you get an error, not a silently downgraded process sandbox. The security posture you configured is the security posture you get — or you get told it can't be met.

Windows Job Objects: the resource cage

The process sandbox is held shut by a Job Object. Defining the term plainly: a Job Object is a Windows kernel object that caps and contains a group of processes — you attach a process (and everything it spawns) to the job, set limits on the job, and the kernel enforces those limits on every process in it. The harness wires this up via P/Invoke (calling the native Windows API directly) in two files:

  • src/Content/Infrastructure/Infrastructure.AI/Sandbox/WindowsProcessResourceLimiter.cs
  • src/Content/Infrastructure/Infrastructure.AI/Sandbox/WindowsJobObjectManager.cs

The default limits — all configurable through ResourceLimits — are:

Limit Default Enforced by What it stops
Memory 256 MB ResourceLimits.MemoryLimitBytes Memory-exhaustion / OOM attacks on the host
CPU time 30 s JOB_OBJECT_LIMIT_PROCESS_TIME CPU-burning loops, cryptomining
Child processes max 5 JOB_OBJECT_LIMIT_ACTIVE_PROCESS Fork bombs, self-replication
Disk quota 100 MB ResourceLimits.DiskQuotaBytes Disk-filling denial of service
Wall-clock timeout 30 s CancellationTokenSource.CancelAfter Hangs, infinite waits, slow-loris-style stalls
Kill on close always JOB_OBJECT_LIMIT_KILL_ON_JOB_CLOSE Orphaned / escaped processes outliving the job

The last row is the one that makes the cage trustworthy. JOB_OBJECT_LIMIT_KILL_ON_JOB_CLOSE guarantees that when the job handle closes, every process in the job dies with it — there is no way for a spawned child to survive the teardown and keep running on the host. Combined with the active-process cap, this closes the most direct denial-of-service against the sandbox itself.

✕
Attack scenario: the fork bomb

A tool is coerced (by a prompt-injection payload, say) into running code that spawns children in a loop, each of which spawns more — classic self-replication designed to exhaust the host's process table and bring the whole machine down.

Inside the Job Object it gets nowhere. JOB_OBJECT_LIMIT_ACTIVE_PROCESS caps the job at 5 processes, so the 6th spawn fails immediately. The bomb can't replicate. When the tool's turn ends, JOB_OBJECT_LIMIT_KILL_ON_JOB_CLOSE reaps whatever is left. The host never sees the pressure.

The closed-by-default capability model

Resource limits stop a tool from overwhelming the host. Capabilities stop it from reaching things it has no business reaching. A capability is a named permission for a category of effect. The flags live in src/Content/Domain/Domain.AI/Sandbox/ToolCapability.cs:

Capability Grants
FileReadRead files
FileWriteWrite files
NetworkAccessMake network connections
SubprocessSpawn child processes
EnvReadRead environment variables
DatabaseReadRead from a database
DatabaseWriteWrite to a database
LlmInvocationCall an LLM

The model has three parts, and the order matters:

1. Declare
A tool states the capabilities it needs by overriding ITool.RequiredCapabilities — a default interface member, so a tool that declares nothing inherits None automatically. This is a static, reviewable contract — you can read a tool's source and know exactly what it claims to need, and a build-time test (DefaultPolicyCapabilityAlignmentTests) checks the shipped governance policy's risk tiers against these declarations so the two cannot silently drift apart.
2. Enforce
At run time, src/Content/Application/Application.AI.Common/Services/Sandbox/CapabilityEnforcer.cs checks each requested effect against what was granted. A capability that wasn't granted yields Result.Forbidden() — the effect simply does not happen.
3. Default to nothing
The default grant is NONE. A tool with no declared capability can read no files, touch no network, and spawn no process. You don't lock a tool down — it arrives locked, and you open exactly the doors it proves it needs.
✓
Closed-by-default is the whole point

An allowlist that defaults to "deny everything" cannot be defeated by forgetting to add a rule — forgetting just keeps the door shut. A blocklist that defaults to "allow everything" fails open the moment you miss a case. The capability model is the former: a new tool, a refactored tool, or a tool whose author forgot to declare a capability all fail safe, not silently broad.

An operator can narrow what a specific tool is allowed to do below its own declaration via a per-tool DeniedCapabilities override in ToolOverrideConfig. The harness keeps the tool's own declaration (ToolPermissionProfile.RequiredCapabilities) and the operator's deny list separate rather than folding one into the other — a tool whose declared requirement overlaps the deny is refused outright, not silently granted a smaller set of capabilities than it actually needs. EffectiveCapabilities (RequiredCapabilities minus DeniedCapabilities) is the value sandbox provisioning and the attestation reads. There is no ambiguity to exploit and no ordering trick: deny always wins, and a denial that the tool actually needs is a hard refusal, never a quiet downgrade. (This control also scopes allow/deny to specific filesystem paths and network hosts via AllowedPaths/DeniedPaths/AllowedHosts/DeniedHosts. CapabilityEnforcer checks a call's requested paths/hosts against these lists whenever the tool declares which of its parameters carry them, and this is enforced identically across all three tool-invocation entry points: the agent's conversational turn, the direct-invoke HTTP surface, and plan/DAG steps. One narrow gap remains, tracked for follow-up rather than fixed: whole-shell-command tools that bypass the capability enforcer entirely. (A plan step chaining a path/host value from an upstream step's output used to bypass scoping the same way; that's now fixed — such a value is normalized identically to a directly-declared one before the check runs.)

!
Capabilities are an enforced boundary only at the Container tier

EffectiveCapabilities is always signed into the run's attestation, but it is only a genuine confinement when MinimumIsolation resolves to Container — that's the tier DockerContainerLaunchPreparer reads it to set network mode and mount read/write-ness. At the Process tier, neither ProcessSandboxExecutor nor ProcessSandboxLaunchPreparer reads Isolation or any capability bit — the only real constraint on a Process-tier run is AllowedPrograms. The attestation records which case applied via a capabilitiesEnforcedBy field ("container" or "declaration-only") so an auditor never mistakes a signed capability set for an enforced one on a run that never left the Process tier.

i
Two similarly-named config types — don't mix them up

SandboxOptions is the Domain-layer configuration for the sandbox model itself; SandboxExecutionOptions is the Application-layer configuration for a specific container execution. They are different types with different jobs. If you find yourself reaching for one and the field you want isn't there, you probably want the other.

Long-lived sessions: a bundle's own MCP server, sandboxed

Everything above assumes a tool runs once: input in, output out, box torn down. A stdio-style MCP (Model Context Protocol — the standard the harness uses to expose tools to the model) server doesn't fit that shape. It's a local process that stays running for the whole conversation, answering many tool calls over one open connection — closer to a phone call than a letter. The one-shot sandbox has no way to hold that call open, so a bundle (an uploaded, untrusted package of instructions and files) that declared its own local MCP server used to be rejected outright at upload time. This section covers the sandbox's second shape: one built to stay open.

ISandboxSession / ISandboxSessionFactory
A duplex, long-lived counterpart to the one-shot executors above — a session stays open across many exchanges instead of closing after one. File: src/Content/Application/Application.AI.Common/Interfaces/Sandbox/ISandboxSession.cs (and its factory interface in the same folder).
DockerSandboxSessionFactory
The real containment boundary for this feature — unprivileged user, dropped capabilities, read-only root filesystem, no network unless explicitly granted. This is the only tier a bundle-owned server is allowed to run on. File: src/Content/Infrastructure/Infrastructure.AI/Sandbox/DockerSandboxSessionFactory.cs.
ProcessSandboxSessionFactory
Exists for the same reason the one-shot process sandbox does — lighter weight, no container runtime — but it is not a containment boundary, and the harness treats it that way: it fails closed and refuses any session request that carries a workspace seed, which is exactly what a bundle-owned server's files need. A bundle's stdio server can only ever land on the Docker tier.

The bridge between this new session shape and the MCP protocol is SandboxedStdioClientTransport (src/Content/Infrastructure/Infrastructure.AI.MCP/Services/SandboxedStdioClientTransport.cs) — it makes an ISandboxSession look, from the MCP SDK's point of view, like an ordinary local process it's talking to over standard input/output. The model's side of the conversation doesn't change at all; only where the other end of the pipe actually executes does.

The registration gate that makes any of this reachable lives in BundleStagingService (src/Content/Infrastructure/Infrastructure.AI/Bundles/BundleStagingService.StdioMcp.cs). Before this capability existed — and still today, unless an operator turns it on — a bundle declaring a stdio-type MCP server in its manifest was rejected outright. Turning it on is a separate, explicit opt-in from the one that already lets a bundle register a remote (http/sse) server:

Setting Default What it controls
StdioMcpServers.Enabled false Master switch. Off: a bundle's stdio declaration is parsed (so it can be logged) but never registered — no container is ever created.
StdioMcpServers.ContainerImage "" (empty) The runtime image every bundle-owned stdio session runs in on this host. Left empty, staging refuses to register a server even with Enabled = true — an empty image would fall back to the harness's own .NET image, which can't run most MCP servers. Whatever an operator sets is still checked against ContainerSandboxOptions.AllowedImagePrefixes, the same allowlist every other sandboxed image goes through.
StdioMcpServers.MaxServersPerBundle 2 How many distinct stdio servers one bundle may declare. Not a container count on its own: each concurrent run against the bundle's staged handle gets its own container per declared server (a stdio session cannot safely be shared across callers), so this multiplies by however many runs are concurrent — see MaxConcurrentSessions below for the actual host-wide cap.
StdioMcpServers.MaxConcurrentSessions 8 A separate, host-wide cap: how many bundle-owned sessions may be running at once, across every staged bundle. Without this, the per-bundle cap alone doesn't bound anything — enough concurrently-staged bundles could still pin an unbounded number of containers between them. Enforced in McpConnectionManager.StartSandboxedStdioSessionAsync.

All four live under AppConfig:AI:BundleExecution:StdioMcpServers — deliberately separate from BundleExecutionConfig.AllowBundleDeclaredMcpServers, the flag for remote servers. Opting into one does not opt into the other; each capability needs its own decision.

The bundle's own files reach the container by copy, not mount — the sandbox workspace is seeded from the bundle's staged directory at session start (src/Content/Infrastructure/Infrastructure.AI/Sandbox/SandboxWorkspace.cs), never bind-mounted. A live session's lifetime and the staging directory's lifetime aren't the same thing; a mount would tie them together and let a running container see later changes to — or write changes back into — files outside the exact snapshot it was granted.

✕
Attack scenario: a bundle names its own command to run

A bundle is just an uploaded package — its manifest is attacker-controllable by definition. A bundle that could name any local command in its own mcp.json and have the host launch it directly would be remote code execution gated by nothing stronger than upload permission.

It never reaches the host that way. It reaches the Docker sandbox instead: unprivileged user, dropped capabilities, read-only root filesystem, an operator-chosen and allowlist-checked image, and — see below — no network. The same containment this chapter already describes for every other sandboxed tool run, applied to a server that happens to stay open for longer than one call.

i
Known limits, stated plainly

No network. A bundle-owned stdio server's capability grant resolves to ToolCapability.None the same way any other bundle-owned tool name does, so its container runs with NetworkMode = "none". It must be fully self-contained in the bundle's own staged files — bridging it to the network would sidestep the egress attribution the remote (http/sse) bundle path already enforces; see Egress & SSRF. An operator must supply the runtime image — the capability stays inert until ContainerImage is set. A failed session start isn't cached — each tool call retries the container start; acceptable today, a candidate for a follow-up if it proves noisy in practice. Container lifetime follows the bundle's own handle — the default HandleTtl is 30 minutes, sliding — so operators sizing a host should account for that, separately from the one-shot executors' 30-second timeout above.

Argument-injection prevention

Even a correctly-caged subprocess can be subverted if you build its command line by gluing strings together. The harness's execution request, src/Content/Domain/Domain.AI/Sandbox/SandboxExecutionRequest.cs, removes the chance entirely: subprocess arguments are passed as an ArgumentList — an array where each element is handed to the process directly, with no shell interpretation in between.

The contrast is the whole defense. Consider a tool that runs grep over user-supplied input:

// SAFE — each array element is a literal argument, never parsed by a shell.
var request = new SandboxExecutionRequest
{
    ArgumentList = ["grep", userInput]   // userInput is ONE argument, whatever it contains
};

// UNSAFE — string concatenation hands the input to a shell to re-parse.
var cmd = $"grep {userInput}";           // a shell now interprets metacharacters in userInput
With ArgumentList, if userInput is the string "; rm -rf /" it is passed to grep as a single literal argument — a pattern to search for, which matches nothing useful and harms nothing. With the concatenated string, a shell sees the ; as a command separator and the rm -rf / as a brand-new command to run.

To force this safe path everywhere, the older string-valued Arguments property is deprecated with error: true — code that still uses it won't compile, so there is no quiet legacy path left that re-parses a shell string.

Shell metacharacters

Characters a command shell treats specially instead of literally: ; and && chain commands, | pipes output, $(...) and backticks substitute command output, > redirects to a file, * expands file names. Argument injection is the trick of smuggling these into an argument so the shell runs your command instead of treating the text as data. Passing an argument array with no shell in the middle means none of these characters are ever special — there is no shell to interpret them.

HMAC attestation: a signed receipt for every run

The cage stops a tool from doing damage. Attestation answers a different question: can you later prove this output really came from a sandboxed run of this input, and wasn't tampered with afterward? The harness signs an attestation after every sandboxed execution in src/Content/Infrastructure/Infrastructure.AI/Attestation/HmacAttestationService.cs, using HMAC-SHA256.

HMAC

A keyed cryptographic signature over some data. Unlike a plain hash — which anyone can recompute — an HMAC can only be produced or verified by someone holding the secret key. So an HMAC over a tool's result proves two things at once: the result hasn't changed, and it was signed by the harness (the keyholder), not forged by whoever stored it.

The signature covers a fixed payload. On a successful run:

// Success payload that gets HMAC-signed:
{toolName}|{inputHash}|{outputHash}|{timestamp:O}|egress:{egressDigest}

// Failure payload (no output to hash, so the failure reason is hashed instead):
{toolName}|{inputHash}|null|{failureHash}|{timestamp:O}
inputHash and outputHash are SHA-256 hashes of the input and the output. On failure there is no output, so a hash of the failure reason takes its place. The timestamp binds when it ran; the egressDigest binds which network destinations were permitted before the untrusted tool ran — see Egress & SSRF for that side of containment.

Why this matters: the attestation proves a given output was produced by a real sandboxed execution of a given input, and was not altered after the fact. The egress digest in the payload ties the run to the exact set of network destinations that were allowed when it started, so a result and its network policy can't be silently decoupled later.

✕
Attack scenario: tampering with a cached result

An attacker who can reach the result cache swaps a stored tool output for a poisoned one — a doctored summary, a fake "all clear," an injected instruction the agent will later read as truth.

The swap changes the bytes, so the outputHash no longer matches the signed attestation. Verification fails. Because the attacker doesn't hold the HMAC key, they can't re-sign the forged payload to make it pass either. A tampered result is detectable, not silently trusted.

Where the signing key lives

An HMAC is only as trustworthy as the secrecy of its key, so the key handling is deliberate. Keys come from IOptionsMonitor<AttestationKeyOptions> — src/Content/Infrastructure/Infrastructure.AI/Attestation/AttestationKeyOptions.cs — sourced from User Secrets in development and Azure Key Vault in production. They are never read from appsettings.json.

  • CurrentKeyVersion supports rotation: the signed payload carries the key version, so keys can be rolled forward without invalidating the ability to identify which key signed an older attestation.
  • Key bytes are zeroed after use with CryptographicOperations.ZeroMemory, so the secret doesn't linger in process memory longer than the signing operation needs it.
i
Secrets handling has its own page

The full story on User Secrets, Key Vault, rotation, and why nothing sensitive ever lands in appsettings.json is on Data Protection & Privacy. This page only covers what the attestation service needs from it.

What a skill can read: two file sandboxes, not one

File access on the host runs through two separate sandboxes with deliberately different permissions. They share one implementation of the rules — allow/deny geometry, symlink resolution, hard-link identity checks — so they cannot drift on what "inside" means, but they permit different directories and offer different operations.

Sandbox Reachable by Permits Operations
IFileSystemService The model, via the file_system tool AppConfig:Infrastructure:FileSystem:AllowedBasePaths, plus the configured logs directory Read, write, list, search
ISkillFileReader The harness's own skill loader only The configured skill content roots: AppConfig:AI:Skills, AppConfig:AI:Agents, and the bundle staging root Read only — the interface exposes no write operation

The guarantee, stated plainly: skill loading is confined to the configured skill content roots, read-only — and to nothing in the model's own file allowlist. Everything a skill can load — its SKILL.md, and any reference or template file it discloses on demand through read_skill_resource — must live inside one of those roots. A path outside them is refused with an error, never silently treated as missing: a refusal that read as "this directory holds no skills" would be indistinguishable from a directory that genuinely holds none, so a misconfigured root would boot an agent quietly missing its skills instead of failing.

The two allowlists are disjoint by design, so the guarantee is not the union of the two — the model's file tool gains nothing from the skill roots, and skill loading gains nothing from the model's configured paths.

Why these are not one sandbox

Merging them is the obvious simplification and it is unsafe. Skill content sits outside the model's sandbox by default — the shipped configuration allows workspace while skills live in skills — so unifying them means adding the skill roots to the allowlist the file_system tool uses. That tool exposes an ungated write. The model would then be able to rewrite its own SKILL.md files, including the allowed-tools list that constrains which tools it may call. Two narrow sandboxes over one shared rulebook avoids that; a regression test (ModelFileSandbox_DoesNotCoverSkillRoots_SoSkillsCannotBeRewritten) fails if the two allowlists are ever merged.

The skill sandbox resolves its permitted roots from live configuration rather than snapshotting them at startup. Plugin-supplied skill directories are registered during host start, after the dependency container is built; a snapshot taken at registration time would refuse exactly the plugin skills the harness advertises support for. A plugin cannot widen the sandbox arbitrarily — its skill directory is accepted only after being verified as contained within the plugin's own directory, and that directory comes from operator configuration under AppConfig:AI:Plugins:Packages.

Where the network side lives

This layer contains what a tool can do on the host: how much it can consume, what it can touch, and proof of what it produced. The other half of containment is where it can reach on the network — blocking SSRF, cloud-metadata credential theft, and exfiltration. That's the egress allowlist whose digest you saw signed into the attestation payload above. It gets the full treatment on the next page.