Sandbox & Execution
The first three layers decide whether a tool may run. This is layer 4: when an allowed tool actually runs code, it runs inside a box with no capabilities it wasn't granted — and its result is cryptographically attested. The earlier layers can be tricked by a clever prompt; this one can't, because it isn't asking the model anything. It's an operating-system-level cage with a signed receipt.
Two boxes: process and container
"Sandbox" means a confined execution environment: the code runs, but the walls of the box decide what it can see and do. The harness ships two boxes, picked by how strong the isolation needs to be.
ProcessSandboxExecutor runs the tool as an isolated Windows subprocess, caged by
a kernel-level Job Object (covered next). Lighter weight; no container runtime required.
File:
src/Content/Infrastructure/Infrastructure.AI/Sandbox/ProcessSandboxExecutor.cs.
DockerSandboxExecutor runs the tool inside a Docker container (driven through
the Docker.DotNet client). Stronger isolation — a whole container boundary
between the untrusted code and the host. File:
src/Content/Infrastructure/Infrastructure.AI/Sandbox/DockerSandboxExecutor.cs.
The Docker executor enforces a minimum-isolation invariant. When a request's
permission profile requires MinimumIsolation = Container and Docker is unavailable
at runtime, the execution fails rather than quietly falling back to the weaker
process sandbox.
A common failure pattern: a control "degrades gracefully" when its strong path is unavailable, and nobody notices the box got weaker. The Docker executor refuses to do this. If you asked for container isolation and the container runtime isn't there, you get an error, not a silently downgraded process sandbox. The security posture you configured is the security posture you get — or you get told it can't be met.
Windows Job Objects: the resource cage
The process sandbox is held shut by a Job Object. Defining the term plainly: a Job Object is a Windows kernel object that caps and contains a group of processes — you attach a process (and everything it spawns) to the job, set limits on the job, and the kernel enforces those limits on every process in it. The harness wires this up via P/Invoke (calling the native Windows API directly) in two files:
src/Content/Infrastructure/Infrastructure.AI/Sandbox/WindowsProcessResourceLimiter.cssrc/Content/Infrastructure/Infrastructure.AI/Sandbox/WindowsJobObjectManager.cs
The default limits — all configurable through ResourceLimits — are:
| Limit | Default | Enforced by | What it stops |
|---|---|---|---|
| Memory | 256 MB | ResourceLimits.MemoryLimitBytes |
Memory-exhaustion / OOM attacks on the host |
| CPU time | 30 s | JOB_OBJECT_LIMIT_PROCESS_TIME |
CPU-burning loops, cryptomining |
| Child processes | max 5 | JOB_OBJECT_LIMIT_ACTIVE_PROCESS |
Fork bombs, self-replication |
| Disk quota | 100 MB | ResourceLimits.DiskQuotaBytes |
Disk-filling denial of service |
| Wall-clock timeout | 30 s | CancellationTokenSource.CancelAfter |
Hangs, infinite waits, slow-loris-style stalls |
| Kill on close | always | JOB_OBJECT_LIMIT_KILL_ON_JOB_CLOSE |
Orphaned / escaped processes outliving the job |
The last row is the one that makes the cage trustworthy.
JOB_OBJECT_LIMIT_KILL_ON_JOB_CLOSE guarantees that when the job handle closes,
every process in the job dies with it — there is no way for a spawned child to survive
the teardown and keep running on the host. Combined with the active-process cap, this closes the
most direct denial-of-service against the sandbox itself.
A tool is coerced (by a prompt-injection payload, say) into running code that spawns children in a loop, each of which spawns more — classic self-replication designed to exhaust the host's process table and bring the whole machine down.
Inside the Job Object it gets nowhere. JOB_OBJECT_LIMIT_ACTIVE_PROCESS caps
the job at 5 processes, so the 6th spawn fails immediately. The bomb can't replicate.
When the tool's turn ends, JOB_OBJECT_LIMIT_KILL_ON_JOB_CLOSE reaps whatever
is left. The host never sees the pressure.
The closed-by-default capability model
Resource limits stop a tool from overwhelming the host. Capabilities stop it from
reaching things it has no business reaching. A capability is a named permission for a
category of effect. The flags live in
src/Content/Domain/Domain.AI/Sandbox/ToolCapability.cs:
| Capability | Grants |
|---|---|
FileRead | Read files |
FileWrite | Write files |
NetworkAccess | Make network connections |
Subprocess | Spawn child processes |
EnvRead | Read environment variables |
DatabaseRead | Read from a database |
DatabaseWrite | Write to a database |
LlmInvocation | Call an LLM |
The model has three parts, and the order matters:
ITool.RequiredCapabilities — a default interface member, so a tool that
declares nothing inherits None automatically. This is a static, reviewable
contract — you can read a tool's source and know exactly what it claims to need, and a
build-time test (DefaultPolicyCapabilityAlignmentTests) checks the shipped
governance policy's risk tiers against these declarations so the two cannot silently
drift apart.
src/Content/Application/Application.AI.Common/Services/Sandbox/CapabilityEnforcer.cs
checks each requested effect against what was granted. A capability that wasn't granted
yields Result.Forbidden() — the effect simply does not happen.
An allowlist that defaults to "deny everything" cannot be defeated by forgetting to add a rule — forgetting just keeps the door shut. A blocklist that defaults to "allow everything" fails open the moment you miss a case. The capability model is the former: a new tool, a refactored tool, or a tool whose author forgot to declare a capability all fail safe, not silently broad.
An operator can narrow what a specific tool is allowed to do below its own declaration via a
per-tool DeniedCapabilities override in ToolOverrideConfig. The
harness keeps the tool's own declaration (ToolPermissionProfile.RequiredCapabilities)
and the operator's deny list separate rather than folding one into the other — a tool whose
declared requirement overlaps the deny is refused outright, not silently granted a smaller
set of capabilities than it actually needs. EffectiveCapabilities
(RequiredCapabilities minus DeniedCapabilities) is the value sandbox
provisioning and the attestation reads. There is no ambiguity to exploit and no ordering
trick: deny always wins, and a denial that the tool actually needs is a hard refusal,
never a quiet downgrade. (This control also scopes allow/deny to specific filesystem paths
and network hosts via AllowedPaths/DeniedPaths/AllowedHosts/DeniedHosts.
CapabilityEnforcer checks a call's requested paths/hosts against these lists
whenever the tool declares which of its parameters carry them, and this is enforced
identically across all three tool-invocation entry points: the agent's conversational
turn, the direct-invoke HTTP surface, and plan/DAG steps. One narrow gap remains, tracked
for follow-up rather than fixed: whole-shell-command tools that bypass the capability
enforcer entirely. (A plan step chaining a path/host value from an upstream step's output
used to bypass scoping the same way; that's now fixed — such a value is normalized
identically to a directly-declared one before the check runs.)
EffectiveCapabilities is always signed into the run's attestation, but it
is only a genuine confinement when MinimumIsolation resolves to
Container — that's the tier DockerContainerLaunchPreparer
reads it to set network mode and mount read/write-ness. At the
Process tier, neither ProcessSandboxExecutor nor
ProcessSandboxLaunchPreparer reads Isolation or any
capability bit — the only real constraint on a Process-tier run is
AllowedPrograms. The attestation records which case applied via a
capabilitiesEnforcedBy field ("container" or
"declaration-only") so an auditor never mistakes a signed capability set
for an enforced one on a run that never left the Process tier.
SandboxOptions is the Domain-layer configuration for the sandbox model
itself; SandboxExecutionOptions is the Application-layer configuration for a
specific container execution. They are different types with different jobs. If you find
yourself reaching for one and the field you want isn't there, you probably want the
other.
Long-lived sessions: a bundle's own MCP server, sandboxed
Everything above assumes a tool runs once: input in, output out, box torn down. A
stdio-style MCP (Model Context Protocol — the standard the harness uses to expose
tools to the model) server doesn't fit that shape. It's a local process that stays running for
the whole conversation, answering many tool calls over one open connection — closer to a phone
call than a letter. The one-shot sandbox has no way to hold that call open, so a bundle (an
uploaded, untrusted package of instructions and files) that declared its own local MCP server
used to be rejected outright at upload time. This section covers the sandbox's second shape:
one built to stay open.
src/Content/Application/Application.AI.Common/Interfaces/Sandbox/ISandboxSession.cs
(and its factory interface in the same folder).
src/Content/Infrastructure/Infrastructure.AI/Sandbox/DockerSandboxSessionFactory.cs.
stdio server can only ever land on the Docker tier.
The bridge between this new session shape and the MCP protocol is
SandboxedStdioClientTransport
(src/Content/Infrastructure/Infrastructure.AI.MCP/Services/SandboxedStdioClientTransport.cs)
— it makes an ISandboxSession look, from the MCP SDK's point of view, like an
ordinary local process it's talking to over standard input/output. The model's side of the
conversation doesn't change at all; only where the other end of the pipe actually executes
does.
The registration gate that makes any of this reachable lives in BundleStagingService
(src/Content/Infrastructure/Infrastructure.AI/Bundles/BundleStagingService.StdioMcp.cs).
Before this capability existed — and still today, unless an operator turns it on — a bundle
declaring a stdio-type MCP server in its manifest was rejected outright. Turning it
on is a separate, explicit opt-in from the one that already lets a bundle register a
remote (http/sse) server:
| Setting | Default | What it controls |
|---|---|---|
StdioMcpServers.Enabled |
false | Master switch. Off: a bundle's stdio declaration is parsed (so it can
be logged) but never registered — no container is ever created. |
StdioMcpServers.ContainerImage |
"" (empty) | The runtime image every bundle-owned stdio session runs in on this host. Left
empty, staging refuses to register a server even with Enabled = true
— an empty image would fall back to the harness's own .NET image, which can't run
most MCP servers. Whatever an operator sets is still checked against
ContainerSandboxOptions.AllowedImagePrefixes, the same allowlist every
other sandboxed image goes through. |
StdioMcpServers.MaxServersPerBundle |
2 | How many distinct stdio servers one bundle may declare. Not a container count on
its own: each concurrent run against the bundle's staged handle gets its own
container per declared server (a stdio session cannot safely be shared across
callers), so this multiplies by however many runs are concurrent — see
MaxConcurrentSessions below for the actual host-wide cap. |
StdioMcpServers.MaxConcurrentSessions |
8 | A separate, host-wide cap: how many bundle-owned sessions may be running at
once, across every staged bundle. Without this, the per-bundle cap alone
doesn't bound anything — enough concurrently-staged bundles could still pin an
unbounded number of containers between them. Enforced in
McpConnectionManager.StartSandboxedStdioSessionAsync. |
All four live under AppConfig:AI:BundleExecution:StdioMcpServers — deliberately
separate from BundleExecutionConfig.AllowBundleDeclaredMcpServers, the flag for
remote servers. Opting into one does not opt into the other; each capability needs its own
decision.
The bundle's own files reach the container by copy, not mount — the sandbox
workspace is seeded from the bundle's staged directory at session start
(src/Content/Infrastructure/Infrastructure.AI/Sandbox/SandboxWorkspace.cs), never
bind-mounted. A live session's lifetime and the staging directory's lifetime aren't the same
thing; a mount would tie them together and let a running container see later changes to — or
write changes back into — files outside the exact snapshot it was granted.
A bundle is just an uploaded package — its manifest is attacker-controllable by
definition. A bundle that could name any local command in its own mcp.json
and have the host launch it directly would be remote code execution gated by nothing
stronger than upload permission.
It never reaches the host that way. It reaches the Docker sandbox instead: unprivileged user, dropped capabilities, read-only root filesystem, an operator-chosen and allowlist-checked image, and — see below — no network. The same containment this chapter already describes for every other sandboxed tool run, applied to a server that happens to stay open for longer than one call.
No network. A bundle-owned stdio server's capability grant resolves to
ToolCapability.None the same way any other bundle-owned tool name does, so
its container runs with NetworkMode = "none". It must be fully
self-contained in the bundle's own staged files — bridging it to the network would
sidestep the egress attribution the remote (http/sse) bundle path already enforces; see
Egress & SSRF.
An operator must supply the runtime image — the capability stays inert
until ContainerImage is set.
A failed session start isn't cached — each tool call retries the
container start; acceptable today, a candidate for a follow-up if it proves noisy in
practice.
Container lifetime follows the bundle's own handle — the default
HandleTtl is 30 minutes, sliding — so operators sizing a host should
account for that, separately from the one-shot executors' 30-second timeout above.
Argument-injection prevention
Even a correctly-caged subprocess can be subverted if you build its command line by gluing
strings together. The harness's execution request,
src/Content/Domain/Domain.AI/Sandbox/SandboxExecutionRequest.cs, removes the chance
entirely: subprocess arguments are passed as an ArgumentList — an array where
each element is handed to the process directly, with no shell interpretation in between.
The contrast is the whole defense. Consider a tool that runs grep over
user-supplied input:
// SAFE — each array element is a literal argument, never parsed by a shell.
var request = new SandboxExecutionRequest
{
ArgumentList = ["grep", userInput] // userInput is ONE argument, whatever it contains
};
// UNSAFE — string concatenation hands the input to a shell to re-parse.
var cmd = $"grep {userInput}"; // a shell now interprets metacharacters in userInput
ArgumentList, if userInput is the string
"; rm -rf /" it is passed to grep as a single literal argument — a
pattern to search for, which matches nothing useful and harms nothing. With the concatenated
string, a shell sees the ; as a command separator and the
rm -rf / as a brand-new command to run.
To force this safe path everywhere, the older string-valued Arguments property is
deprecated with error: true — code that still uses it won't compile, so there is no
quiet legacy path left that re-parses a shell string.
Characters a command shell treats specially instead of literally:
; and && chain commands,
| pipes output, $(...) and backticks substitute command output,
> redirects to a file, * expands file names. Argument
injection is the trick of smuggling these into an argument so the shell runs your command
instead of treating the text as data. Passing an argument array with no shell in the
middle means none of these characters are ever special — there is no shell to interpret
them.
HMAC attestation: a signed receipt for every run
The cage stops a tool from doing damage. Attestation answers a different question:
can you later prove this output really came from a sandboxed run of this input, and wasn't
tampered with afterward? The harness signs an attestation after every sandboxed execution in
src/Content/Infrastructure/Infrastructure.AI/Attestation/HmacAttestationService.cs,
using HMAC-SHA256.
A keyed cryptographic signature over some data. Unlike a plain hash — which anyone can recompute — an HMAC can only be produced or verified by someone holding the secret key. So an HMAC over a tool's result proves two things at once: the result hasn't changed, and it was signed by the harness (the keyholder), not forged by whoever stored it.
The signature covers a fixed payload. On a successful run:
// Success payload that gets HMAC-signed:
{toolName}|{inputHash}|{outputHash}|{timestamp:O}|egress:{egressDigest}
// Failure payload (no output to hash, so the failure reason is hashed instead):
{toolName}|{inputHash}|null|{failureHash}|{timestamp:O}
inputHash and outputHash are SHA-256 hashes of the input and the
output. On failure there is no output, so a hash of the failure reason takes its place. The
timestamp binds when it ran; the egressDigest binds which network
destinations were permitted before the untrusted tool ran — see
Egress & SSRF for that side of containment.
Why this matters: the attestation proves a given output was produced by a real sandboxed execution of a given input, and was not altered after the fact. The egress digest in the payload ties the run to the exact set of network destinations that were allowed when it started, so a result and its network policy can't be silently decoupled later.
An attacker who can reach the result cache swaps a stored tool output for a poisoned one — a doctored summary, a fake "all clear," an injected instruction the agent will later read as truth.
The swap changes the bytes, so the outputHash no longer matches the signed
attestation. Verification fails. Because the attacker doesn't hold the HMAC key, they
can't re-sign the forged payload to make it pass either. A tampered result is detectable,
not silently trusted.
Where the signing key lives
An HMAC is only as trustworthy as the secrecy of its key, so the key handling is deliberate. Keys
come from IOptionsMonitor<AttestationKeyOptions> —
src/Content/Infrastructure/Infrastructure.AI/Attestation/AttestationKeyOptions.cs —
sourced from User Secrets in development and Azure Key Vault in production. They are
never read from appsettings.json.
-
CurrentKeyVersionsupports rotation: the signed payload carries the key version, so keys can be rolled forward without invalidating the ability to identify which key signed an older attestation. -
Key bytes are zeroed after use with
CryptographicOperations.ZeroMemory, so the secret doesn't linger in process memory longer than the signing operation needs it.
The full story on User Secrets, Key Vault, rotation, and why nothing sensitive ever lands
in appsettings.json is on
Data Protection & Privacy. This page only
covers what the attestation service needs from it.
What a skill can read: two file sandboxes, not one
File access on the host runs through two separate sandboxes with deliberately different permissions. They share one implementation of the rules — allow/deny geometry, symlink resolution, hard-link identity checks — so they cannot drift on what "inside" means, but they permit different directories and offer different operations.
| Sandbox | Reachable by | Permits | Operations |
|---|---|---|---|
IFileSystemService |
The model, via the file_system tool |
AppConfig:Infrastructure:FileSystem:AllowedBasePaths, plus the
configured logs directory |
Read, write, list, search |
ISkillFileReader |
The harness's own skill loader only | The configured skill content roots:
AppConfig:AI:Skills, AppConfig:AI:Agents, and the bundle
staging root |
Read only — the interface exposes no write operation |
The guarantee, stated plainly: skill loading is confined to the
configured skill content roots, read-only — and to nothing in the model's own file
allowlist. Everything a skill can load — its SKILL.md, and any reference or
template file it discloses on demand through read_skill_resource — must live
inside one of those roots. A path outside them is refused with an error, never silently
treated as missing: a refusal that read as "this directory holds no skills" would be
indistinguishable from a directory that genuinely holds none, so a misconfigured root would
boot an agent quietly missing its skills instead of failing.
The two allowlists are disjoint by design, so the guarantee is not the union of the two — the model's file tool gains nothing from the skill roots, and skill loading gains nothing from the model's configured paths.
Merging them is the obvious simplification and it is unsafe. Skill content sits outside the
model's sandbox by default — the shipped configuration allows workspace while
skills live in skills — so unifying them means adding the skill roots to the
allowlist the file_system tool uses. That tool exposes an ungated
write. The model would then be able to rewrite its own SKILL.md
files, including the allowed-tools list that constrains which tools it may
call. Two narrow sandboxes over one shared rulebook avoids that; a regression test
(ModelFileSandbox_DoesNotCoverSkillRoots_SoSkillsCannotBeRewritten) fails if
the two allowlists are ever merged.
The skill sandbox resolves its permitted roots from live configuration rather than snapshotting
them at startup. Plugin-supplied skill directories are registered during host start, after the
dependency container is built; a snapshot taken at registration time would refuse exactly the
plugin skills the harness advertises support for. A plugin cannot widen the sandbox arbitrarily
— its skill directory is accepted only after being verified as contained within the plugin's own
directory, and that directory comes from operator configuration under
AppConfig:AI:Plugins:Packages.
Where the network side lives
This layer contains what a tool can do on the host: how much it can consume, what it can touch, and proof of what it produced. The other half of containment is where it can reach on the network — blocking SSRF, cloud-metadata credential theft, and exfiltration. That's the egress allowlist whose digest you saw signed into the attestation payload above. It gets the full treatment on the next page.