Observability & Safety
Agents are non-deterministic systems that call APIs, write files, and answer with confidence even when they're wrong. You need to see inside them — and stop them when they go astray. This page covers both: the observability pipeline and the safety layers that wrap every turn.
The observability pipeline
Everything in the harness is instrumented with OpenTelemetry. Every MediatR command opens a span. Every agent turn opens a child span. Every tool call opens a child of that. Every LLM request opens a child of that. You can drill from "this conversation failed" all the way down to "the third sub-call to GPT-4o returned a malformed JSON".
An open standard for emitting traces, metrics, and
logs from running code.
A span is a record of one operation — when it started, when it
ended, whether it succeeded, plus arbitrary key/value tags. Think of a span as
"here's one timed thing that happened, with notes attached."
Spans nest. The outer "HandleAgentTurn" span contains a "CallLLM"
span, which contains three "InvokeTool" spans, each of which might contain
"ReadFile" spans. The whole nested tree for one user request is a trace.
Metrics are different — they're aggregated numbers emitted separately
(counters, histograms, gauges) that show patterns over time without recording every
individual event.
We export traces to Jaeger (a free trace viewer you run locally or
in your cluster), metrics to Prometheus, and optionally everything to
Azure Monitor in production.
What gets instrumented
- Every MediatR request — via
RequestTracingBehavior. - Every agent turn — span attributes include agent name, conversation ID, turn number, input/output tokens, cost.
- Every tool call — span attributes include tool name, operation, latency, success/failure.
- Every LLM request — via the Microsoft.Extensions.AI built-in tracing, augmented by our custom
LlmTokenTrackingProcessorthat recognizes spans from Agents.AI / Semantic Kernel and enriches them with agentic context (including cache-read / cache-write token counts). - RAG retrievals — span per query transform, retrieval, rerank, assembly.
- Tool output compression —
ToolOutputCompressionBehaviorlogs compression events, strategy used, and size reduction ratio. - Prompt-cache usage —
PromptCachingPipelinePolicy+TokenUsageMetricstrack cache-read and cache-write token counts, emitted as theagent.tokens.cache_read/agent.tokens.cache_writeinstruments. - Real token streaming — when a turn runs via
RunStreamingAsync, deltas surface throughIAgentTurnStreamSinkas they arrive; the turn span captures the streamed run.
Metric instruments use dotted names like agent.orchestration.turns_total,
agent.tokens.input/output/total,
agent.tokens.cache_read/cache_write,
rag.retrieval.duration, and agent.tool.invocations. Span
operation names follow the GenAI semantic conventions —
invoke_agent, chat, execute_tool,
embeddings — with attributes like gen_ai.request.model and
gen_ai.usage.input_tokens. Never add a harness.* prefix in app
code: the Prometheus agentic_harness_ prefix is applied by the OTel
collector, and prefixing in-app double-prefixes it.
Setting up Jaeger locally
Run Jaeger with one Docker command, then point the OTLP exporter at it:
# Start Jaeger all-in-one
docker run -d --name jaeger \
-p 16686:16686 \
-p 4317:4317 \
jaegertracing/all-in-one:latest
# Browse to http://localhost:16686
Configure the harness to export to Jaeger via the OTLP endpoint in
AppConfig.Observability.Exporters.Otlp.Endpoint. Run an agent example. Refresh
Jaeger and you'll see the conversation timeline span-by-span.
The PostgreSQL audit store
OTel traces are great for debugging but ephemeral. For long-term audit
(compliance, post-incident review, governance), the harness writes a structured record per
turn to PostgreSQL via IObservabilityStore. The connection string lives in
AppConfig.Observability.PostgresConnectionString. Schema includes turn-level
token usage, cost, tool invocations, message previews truncated to 500 chars, plus the
original message body and the args + stdout for every captured tool call. Those last two are
what the deep-link endpoints below serve.
Tamper-evident audit hash-chain
For audit records that must be provably un-altered, the harness writes a
hash-chained JSONL log via HashChainedJsonlWriter: each entry
carries a hash over its own content plus the previous entry's hash, so any retroactive edit or
deletion breaks the chain. AuditChainVerificationService (behind
IVerifiableAuditChain) walks the file and reports the first broken link. This gives
you a cheap, file-based tamper-evidence guarantee without a database — useful for governance
and compliance trails where "nobody quietly changed the record" needs to be demonstrable.
Microsoft Agent 365 — putting agents on the tenant's radar
Everything above tells you what your agents did. None of it tells the organisation that your agents exist. To a tenant's IT and security teams, an agent built here is an unmanaged non-human actor — a "shadow agent" — no matter how well governed it is internally. Microsoft Agent 365 is the tenant-side control plane that closes that gap: a central inventory of every agent, each with a real Microsoft Entra agent identity, and its activity flowing into Microsoft Defender, Microsoft Purview and the Microsoft 365 admin center.
The harness ships this as an opt-in trace exporter. It is off by default and a
host that omits the section registers none of it. Turn it on under
AppConfig.Observability.Exporters.Agent365:
"Observability": {
"WebTelemetryProjects": [ "Presentation.AgentHub" ],
"Exporters": {
"Agent365": {
"Enabled": true,
"AgentAppId": "00000000-0000-0000-0000-000000000000",
"TenantId": "00000000-0000-0000-0000-000000000000",
"BlueprintId": null,
// Only for a host that runs several distinct agents. Keyed by the agent id
// carried on agent-scoped requests; unlisted agents use AgentAppId above.
"Agents": {
"researcher": { "AppId": "00000000-0000-0000-0000-000000000000" }
}
}
}
}
AgentAppId is the appId of the agent identity the host runs
as — not the blueprint's, and not the object id. Both ids must be GUIDs, and a startup
validator refuses to boot if they are not. That strictness is deliberate: Agent 365 does not
return an error for a malformed agent id, it just shows the agent as unidentified, so startup
is the only cheap place to catch it.
What gets reported. Every execution path that establishes an agent context — an ordinary agent turn, an orchestrated task, a plan run, a sub-plan, and a directly invoked tool — reports under the right identity, and you do not have to do anything per path to make that true. Identity is published by the act of establishing the context itself and released when that context's scope ends, so a path cannot establish one and omit the attribution. That used to be a separate call each path made, and three of five paths did not make it: their spans reached the service carrying no agent identity, which Agent 365 discards without reporting an error, so those categories of activity were simply absent from the tenant's records with nothing to indicate it.
What you have to do outside the harness. Registration is a provisioning step, not a runtime one — it creates Entra directory objects, needs a Global Administrator, and needs a named human sponsor accountable for the agent. The harness deliberately automates none of that; a running agent must never hold those privileges. An operator needs to:
- Register an agent identity blueprint in Entra and create an agent identity from it. Microsoft ships a CLI and coding-assistant skills for this. Rule of thumb: one blueprint per credential boundary — credentials live on the blueprint and are shared by every identity minted from it.
-
Grant the blueprint
Agent365.Observability.OtelWriteand have a tenant administrator consent to it once. - Assign a Microsoft 365 E7, Test - Microsoft 365 E7 or Microsoft Agent 365 Frontier licence to at least one user in the tenant.
An unassigned licence discards everything silently. The SKU merely
existing in the tenant is not enough — it must be assigned to a user. Until it is, the
service accepts requests and throws the telemetry away, so the agent simply never
appears. And a host that is not listed in
WebTelemetryProjects cannot export at all, because the exporter
attaches to a pipeline only that host shape builds. The harness refuses to boot in that
case rather than letting you debug a missing dashboard row; today only
Presentation.AgentHub lists itself, so ExecutionApi and FoundryHost must
be added before they can export.
There are two independent baggage stores in .NET/OpenTelemetry, each
with its own propagator — and this harness's own identity attribution (user id,
conversation id, and Agent 365's tenant/agent/blueprint ids) rides the one most
developers don't think to check: System.Diagnostics.Activity.Baggage, not
OpenTelemetry's own Baggage API. A propagator that carries either would
serialise that data onto every outbound HTTP call and accept a caller-supplied version
on every inbound one — so Observability:PropagateBaggage governs
both stores together, installing a trace-context-only propagator on
each whenever it is false (the default), unconditionally, for
every host, not only ones that enable Agent 365. Enabling Agent 365 no
longer changes this behaviour itself; it depends on the same host-wide policy every
other feature does. Set the flag to true only if this host has a
deliberate, reviewed reason to use cross-process baggage — the harness then installs
the standard trace-context-plus-baggage propagator explicitly on both stores, rather
than leaving it to whatever happened to be ambient. A startup check re-asserts both
propagators unconditionally (independent of whether Agent 365 is enabled) and refuses
to boot if something re-registers baggage-carrying propagation afterwards; if Agent 365
is also enabled while this flag is true, the harness logs a warning naming
the egress, since that combination is materially riskier than either setting alone.
Microsoft's exporter persists undeliverable spans to LOCALAPPDATA or
TEMP by default and replays them in the background. Harness spans can
carry prompts, tool arguments and model output, so the harness inverts that default —
where such content lands is a deployment's decision, not a library's. The failure
compounds too: while a tenant is unlicensed or unconsented every export fails,
so everything would be written to disk rather than an occasional retry batch.
Set EnableOfflineStorage plus an
OfflineStorageDirectory you control to opt in — the harness creates that
directory and enforces owner-only permissions on it itself (POSIX; a no-op on Windows,
left to its inherited ACL) and refuses to boot if it cannot confirm the directory is
secured, rather than merely documenting the requirement.
This integration is proven by unit and composition tests and by inspecting spans
locally with the console exporter — span shape, the required root
invoke_agent span, and the agent/tenant attribution that is otherwise the
silent failure. It has not been observed ingesting into a licensed
Agent 365 tenant. Treat first-run ingestion, licensing and consent as unverified and
check the Microsoft 365 admin center after your first turn.
Deep-link endpoints for the dashboard
The harness ships a web dashboard (Foresight, in
Presentation.Dashboard) that lists past sessions. Because the audit store truncates
previews to 500 characters, the dashboard needs a way to fetch the untruncated payload when you
click into a row. Two scoped endpoints on SessionsController do that:
-
GET /api/sessions/{id}/tools/{invocationId}— returnsToolInvocationDetailDtowith the JSON args the LLM passed to the tool, the full stdout returned to the model, the LLM-suppliedCallId, plus the standard metrics (duration, status, error type, result size). Bothargsandstdoutare populated byToolDiagnosticsMiddleware, which interceptsFunctionCallContent+FunctionResultContentand pairs them byCallIdthrough the scopedILlmUsageCapture. Args pass through the optionalISecretRedactorbefore storage. -
GET /api/sessions/{id}/messages/{messageId}— returnsMessageBodyDtowith the fullcontent_fullbody captured before the 500-char preview truncation. The list endpoints still return only the preview to keep payloads cheap; the detail endpoint serves the full body on demand.
Both endpoints scope the lookup to (sessionId, id), so a forged
invocationId or messageId from a different session returns 404
rather than leaking content across session boundaries. The Dashboard wires the deep-links
from ToolsTable (tool name → /sessions/:id/tools/:invocationId)
and SessionTimeline ("view full →" link → /sessions/:id/files/:messageId).
Schema migration: nothing to run by hand. The
content_full, call_id, args and
stdout columns arrive from migration
004_message_and_tool_bodies.sql, which the harness applies itself on the
first database connection it opens. Rows recorded before that migration ran return
contentFull: null; the dashboard renders a banner explaining the preview
fallback in that case.
This page previously told you to apply that file yourself with psql, because
the schema used to be delivered by SQL mounted into the Postgres container — and Postgres
runs those scripts only when it first creates an empty data directory. A database
that already held data could not receive a schema change at all, so hand-applying was the
only route. That is fixed: schema now ships as numbered migrations embedded in the
application, applied under a lock, with a ledger table recording what has already run, so
an installation that has been live for months picks up exactly the changes it is missing.
Two things worth knowing if you are adapting this for your own deployment. To add a schema
change, drop a numbered .sql file into
Infrastructure.Observability/Migrations/ — it is embedded automatically and
applied in numeric order; write it to be safely re-runnable. And
Dashboards/postgres-bootstrap/ is a separate, once-per-cluster step that
creates the read-only role Grafana connects as. That one is not applied by the
application, because creating a role needs a privilege no least-privilege application
account should hold — against a managed Postgres you run it yourself, once.
Safety layers, top to bottom
"Safety" in this codebase isn't one thing. It's a stack of independent layers, each catching a different class of problem.
1 · Prompt injection detection
PromptInjectionBehavior runs deterministic detectors over every user input —
looking for known patterns ("ignore previous instructions", suspicious base64 blobs,
instruction-resembling content in unexpected places). Threats above
InjectionBlockThreshold halt the request before it reaches the LLM.
2 · Content safety middleware
Wrapped around the chat client itself (in AgentFactory). Filters configurable
content categories — PII, profanity, classified information — both on user input and LLM
output. The harness ships with sane defaults; tighten or loosen via
AppConfig.AI.Governance.
3 · Tool-invocation governance — GovernedAIFunction
This is the live tool-gating layer, and it runs on the tool-execution path,
not the MediatR request pipeline. Every converted tool is wrapped by
GovernedAIFunction, whose InvokeCoreAsync runs three ambient gates in
order before the tool executes:
IToolInvocationGovernor— authorization. Internally chainsIToolPermissionService(ThreePhasePermissionResolver) → graded-autonomy risk gate → capability enforcement → the YAML policy engine. Opt-in viaGovernanceConfig.EnforceToolInvocation.IToolClassificationGate— Purview data-classification DLP; opt-in viaAppConfig:AI:Governance:DataClassification:Mode(Off/Audit/Enforce).IProgressEvaluator— spin / no-progress guard.
GovernancePolicyBehavior and ToolPermissionBehavior were deleted
Earlier drafts described these two as MediatR pipeline behaviors. Both were removed in
PR #90 — they keyed on an IToolRequest marker nothing implemented, so they
never fired. The YAML policy engine and the permission resolver still exist, but as
stages inside IToolInvocationGovernor, reached from
GovernedAIFunction.
4 · Permission resolution & autonomy tiers
The first stage of the governor is IToolPermissionService
(ThreePhasePermissionResolver), which evaluates each tool call against the agent's
autonomy tier (Restricted, Supervised, Autonomous).
Rules are resolved across 9 sources including PluginDeclaration,
through a 3-phase resolver (Deny gates → Ask rules → Allow rules). When the decision is
"requires approval" the governor records a PendingApproval and blocks
fail-closed — live mid-call human escalation is deliberately deferred, so it does
not route to IEscalationService in the middle of a tool call. The
tool's RiskTier (a BlastRadius value) feeds the graded-autonomy gate:
higher tiers may auto-approve low-radius tools while still gating high-radius ones.
4b · Plugin-boundary governance
When a skill comes from a local plugin, additional restrictions
apply from the plugin's PluginDeclaration:
AllowedTools— whitelist of tools the plugin's skills can access. Tools not on the list are filtered out during context assembly.DeniedTools— blacklist that is bypass-immune. Even when the agent runs in Autonomous mode, DeniedTools are always blocked. Use this for hard security boundaries.AutonomyLevel— overrides the agent's default autonomy tier for this plugin's tool calls.
This lets harness operators grant plugins access to capabilities while constraining their blast radius — a third-party plugin can read files but never write them, regardless of the agent's own permissions.
An AllowedTools/DeniedTools entry that matches no real tool — first-
party or MCP — is fail-closed, not a silent no-op. If resolvable at startup (no MCP server is
configured anywhere on the host), the host refuses to boot. Otherwise it's checked once the
entry's tool is discovered on some configured server; if still unresolved once every server
has reported, the whole plugin's boundary is marked faulted and every tool from its skills is
denied. This denial isn't limited to the plugin in question, and it starts as soon as any
plugin's boundary is unresolved — not only once it's confirmed faulted: every first-party
tool is denied agent-wide for as long as any plugin's boundary hasn't verified, since
there's no way to know which global tool an unresolved entry was meant to protect.
5 · The sandbox
Tools that touch the filesystem do so through path validators. The agent thinks it has a file system; it actually has a fenced subset. Covered in detail on Tools & Keyed DI.
6 · MCP security scanning
Tools from external MCP servers are scanned at registration time. Suspicious schemas are logged and skipped. See MCP Server & Client.
7 · Response sanitization
ResponseSanitizationBehavior redacts known PII / secret patterns from
tool responses before they reach the LLM. Findings above
ResponseBlockThreshold block the response entirely.
7b · Tool-call replay treatment
Layer 7 guards a tool response on its way into the current turn. When a conversation is later resumed, the harness replays that turn's real tool calls and results back to the model as its own memory — and that is a different trust boundary: the content is durable (persisted in the conversation store indefinitely, not a transient in-flight message) and model-facing (read back as fact on every subsequent turn).
IToolCallReplayTreatment is the layer for that boundary. Before anything is
persisted it runs sanitize → redact → structural secret redaction, then size-tiers the
result: replayed verbatim under MaxVerbatimChars, truncated with a visible marker
above it, and withheld outright above the ceiling where redaction stops being trustworthy. The
structural pass is what distinguishes it from layer 7's — a secret sitting in a JSON field
named "token" is invisible to a plain value-shape scan, so a durable path
protected only by layer 7's patterns would be weaker than the transient one it sits
behind.
Two further settings bound growth the size-tier alone does not reach, because it caps each
payload individually and nothing caps how many of them accumulate.
MaxCallsPerTurn (default 32) limits how many tool calls a single turn persists —
nothing upstream does, since the framework's per-request iteration limit caps tool-calling
rounds, not the calls issued in parallel within one round.
MaxReplayedChars (default 65536, roughly 16k tokens) limits the total treated text
one replayed window sends back to the model, and is the only real ceiling on a resumed
conversation's per-turn prompt cost: the store's dispatch window is capped in rows, and
each row expands into two chat messages per tool call it carries, so a row cap stops bounding
the prompt once turns are tool-heavy. Both trim from the same end — newest kept, oldest
dropped — so they compose into one coherent policy rather than a slice from the middle of a
turn. The two are validated against each other at startup: a window budget smaller than a
single maximum-size call would drop not just that call but every older one behind it.
Operators can disable replay entirely with
AI:Conversations:ToolCallReplay:Enabled. That switch gates both writing new tool
payloads and replaying already-persisted ones — but note it does not
un-persist anything already stored, so containment after a leak needs a store-side
purge as well.
ToolName and CallId are narrowed separately, before either ever
reaches IToolCallReplayTreatment: ToolCallTranscriptExtractor
restricts both to [A-Za-z0-9_-], truncates to 128 characters, and appends a
short hash suffix when sanitization actually changed the value (so two distinct raw ids
can't collapse onto one persisted CallId). This is deliberately not the
sanitize/redact/size-tier pipeline above — that pipeline is built for free text, and an
identifier has a much narrower legitimate shape than a payload does.
8 · Human escalation
When the governance engine decides a decision exceeds the agent's authority, it raises an
escalation. Escalations have a priority level
(Informational, Blocking, Critical) and configurable
timeouts and approval strategies (AnyOf, AllOf). Records persist
under .agent-sessions/escalations/ for audit. See
AppConfig.AI.Governance.Escalation.
The audit trail
AuditTrailBehavior writes a tamper-evident record of every governance decision —
what was requested, what the rule said, who approved or denied, when. Combined with the
PostgreSQL observability store, you have a full reconstruction path for any conversation: who
said what, what the agent decided, what it called, what came back.
Drift detection and learnings
Two longer-term feedback loops:
-
Drift detection — EWMA-based scoring of quality regressions across runs.
When per-task scores trend down, the system flags a drift event. Config:
AppConfig.AI.DriftDetection. -
Learnings — cross-session feedback that biases retrieval, skill loading,
and tool selection. Has decay (older learnings weight less) and pruning (forgotten if no
longer relevant). Config:
AppConfig.AI.Learnings.
The meta-harness loop
Beyond per-run safety, the harness can optimize itself. The meta-harness loop
(RunHarnessOptimizationCommand) does this iteratively:
- Snapshot — capture current skill files as the baseline.
- Propose — a coding agent reads recent execution traces (JSONL files in
.meta-harness/traces/), reasons about why turns failed, and outputs a proposal — which skill files to change, how, and what it observed. - Evaluate — run the proposed skill files against the benchmark task suite in
eval-tasks/and score by regex match. - Regression gate — verify the candidate doesn't regress on tasks prior winners already solved (threshold default 80%).
- Record — promote to new best if score improved and regression gate passed.
This is the system that turns drift detection into improvement: signals from observability feed into proposals, which feed back into the skills the agent uses next time.
Reading a Jaeger trace
Describing a trace is much less useful than seeing one, so here is roughly what a two-turn conversation looks like in the Jaeger UI — one where the second turn went wrong:
RunConversationCommand 2.41s ← root: one whole conversation
├─ ExecuteAgentTurnCommand 0.83s ← turn 1
│ ├─ chat gen_ai.usage.input_tokens=1204 0.79s ← the LLM call itself
│ │ gen_ai.usage.output_tokens=88
│ └─ (no tool calls this turn)
└─ ExecuteAgentTurnCommand 1.57s ← turn 2, the one that failed
├─ chat gen_ai.usage.input_tokens=1533 0.61s
│ gen_ai.usage.output_tokens=142
├─ execute_tool tool=file_system 0.04s
│ otel.status_code=ERROR ← START HERE
│ error=Path escapes sandbox root
└─ chat gen_ai.usage.input_tokens=1791 0.88s ← the model retrying
gen_ai.usage.output_tokens=64 after seeing the error
Four things to read off that, in order:
-
Find the root.
RunConversationCommandis the whole conversation. Its duration is what the user actually waited. (MediatR spans are named after the C# request type, which is why they carry theCommandsuffix.) -
Find the failing turn. Each
ExecuteAgentTurnCommandchild is one turn. Scan for the red one, or the slow one. -
Look at the turn's children. A
chatspan is the model thinking; anexecute_toolspan is the harness doing. In the trace above, the tool span carriesotel.status_code=ERRORand anerrortag naming the sandbox rejection — that is the actual defect, and everything after it is consequence. -
Check the token tags when the model, not a tool, is the problem. The
gen_ai.usage.*attributes tell you whether input grew unexpectedly (context bloat) or output was truncated.
Many identical execute_tool spans in a row means the agent is looping —
that is what the spin guard in
Chapter 04 exists to stop. One enormous
chat span means a context or model problem, not a tool
problem. A turn with no tool spans at all, when you expected some, usually means a
permission denial upstream — check the governance path in the safety layers above,
not the tool.