Skills System
Skills are how this harness teaches an agent what it knows, what it can do, and when to do it. They're plain Markdown files — humans write them, agents read them, the harness loads them in three tiers to keep the token budget under control. This is the most Claude-Code-inspired part of the codebase.
The mental model
Think of a skill as a "job description" the agent receives at runtime. The job description says: here's your role, here are the tools you can use, here's how to think about typical tasks, and here are the deeper references if you need them. The harness can have dozens of skill files on disk. A single-skill agent has one skill "in role" at any given turn; a multi-skill agent merges instructions and tools from several skills simultaneously.
The three tiers — progressive disclosure
An LLM has a finite context window. If you eagerly load every detail of every skill at startup, you blow the budget before the first user message lands. The harness solves this the same way Claude Code does: skills are loaded in three tiers, only as deeply as needed.
| Tier | What's in it | Approx. tokens | When loaded |
|---|---|---|---|
| Tier 1 — Index Card | ID, name, description, category, tags | ~100 | At agent startup. The agent always sees the index for every available skill. |
| Tier 2 — Folder | Full instructions, behavioral guidelines, tool declarations | ~5,000 | When the agent (or the orchestrator) selects this skill for the turn. |
| Tier 3 — Filing Cabinet | Scripts, reference docs, templates, examples, schemas | Unbounded | Only when the skill actively executes and needs that resource. |
At 100 tokens per skill, even 50 skills add only ~5K tokens to the startup budget — enough to let the agent know what it has access to without committing resources. Tier 2 is loaded one-at-a-time per turn. Tier 3 is loaded only when a tool physically opens that file. The token budget is never spent speculatively.
Anatomy of a SKILL.md file
Skills live under skills/ at the repo root (configured via
AppConfig.AI.Skills.BasePath). Each skill is a single Markdown file with two
parts: a block of structured metadata at the top, and human-readable instructions below.
Markdown is the lightweight plain-text formatting language you're
already used to from READMEs (# headings, **bold**, etc.).
YAML is a structured text format that looks like indented
key: value pairs — much friendlier for humans than JSON.
Frontmatter is a block of YAML at the top of a Markdown file, fenced
between two --- lines. It's how a Markdown file carries metadata (id,
tags, dependencies) that tooling can parse without parsing the prose.
Here's what a real SKILL.md looks like:
---
id: research
name: Research Agent
description: Finds and synthesizes information from documents and the web.
category: information-gathering
tags: [research, search, summarize]
allowed-tools:
- file_system
- document_search
version: 1.0.0
model-override: gpt-4o # optional override of AppConfig default
skill_type: research
---
# Research Agent
## Role
You are a thorough research assistant. When given a question, you:
1. Identify what is and isn't known.
2. Use the document_search tool to gather sources.
3. Synthesize findings with citations.
## Behavioral guidelines
- Always cite sources by URL or document ID.
- Prefer primary sources over summaries.
- If sources conflict, surface the conflict rather than picking one.
...
The frontmatter is the Tier-1 metadata: id, name,
description, category, tags are the index card;
allowed-tools declares the tool surface. The parser
(SkillMetadataParser) reads a fixed set of keys — category,
tags, version, model-override, agent-id,
allowed-tools, prerequisites, completion_tool,
skill_type — anything else lands in a generic Metadata bag. The
Markdown body is the Tier-2 content — the agent's instructions for this role. Tier-3 resources
aren't listed in frontmatter at all: they're discovered from the skill folder's
references/, templates/, and scripts/ subdirectories.
The SkillDefinition domain type
When the loader parses a SKILL.md, it produces a SkillDefinition:
public class SkillDefinition
{
// Level 1 — Index Card (always loaded)
public string Id { get; set; } = string.Empty;
public string Name { get; set; } = string.Empty;
public string Description { get; set; } = string.Empty;
// Level 2 — Folder (on demand)
public string? Instructions { get; set; } = string.Empty; // Markdown body
// Categorization + runtime config
public string? Version { get; set; }
public string? Category { get; set; }
public string? SkillType { get; set; } // from skill_type
public IList<string> Tags { get; set; } = new List<string>();
public IList<string>? AllowedTools { get; set; }
public string? ModelOverride { get; set; } // from model-override
public string? AgentId { get; set; } // from agent-id
public string? PluginSource { get; set; } // plugin ID, if Injected
// Prerequisite ordering
public IList<string> Prerequisites { get; set; } = new List<string>();
public string? CompletionTool { get; set; } // from completion_tool
// Level 3 — Filing Cabinet (discovered from subfolders)
public IList<SkillResource> Templates { get; set; } = new List<SkillResource>();
public IList<SkillResource> References { get; set; } = new List<SkillResource>();
public IList<SkillResource> Scripts { get; set; } = new List<SkillResource>();
// Computed — NOT settable
public bool IsPluginSkill => !string.IsNullOrEmpty(PluginSource);
public SkillMode Mode => IsPluginSkill && !HasToolDeclarations && !HasToolRestrictions
? SkillMode.Injected : SkillMode.Managed;
// ... Author, License, StateConfiguration, DecisionFramework, etc.
}
Notice this is a plain mutable class — get; set; properties and
IList<> collections, rather than the immutable records used almost everywhere
else in this codebase. That is deliberate: the parser fills it in field by field as it reads
down the SKILL.md, so it needs somewhere to accumulate.
The one thing you can't set is Mode: it's a computed property
derived from IsPluginSkill plus whether the skill declares any tools. The
PluginSource and Prerequisites properties are covered in the sections
below.
The richer tools: block
allowed-tools: is a bare name list — enough when a skill just needs a tool
available. A skill that needs to say more about how it uses a tool — required
operations, a fallback if the tool is unavailable, or that a tool must be called at most
once per conversation — declares a structured tools: block instead. Each entry
parses into a Domain.AI.Tools.ToolDeclaration:
---
id: diagnostics
name: Diagnostic Agent
tools:
- name: start_diagnostic_session
operations: [open, close]
optional: false
call-once-per-conversation: true # refused durably on a second call — see below
description: Opens the stateful session a diagnostic conversation runs inside.
- name: github_repos
optional: true
fallback: file_system
description: Looks up related issues when available.
---
name— the tool name (keyed-DI key, or an MCP server name in Managed mode). An entry with no name is dropped, not loaded nameless.operations— the specific operations this skill uses; empty means all are allowed.optional— whether the skill can function without this tool. Absent means required:ToolChainBuilderthrows at resolution rather than silently proceeding without a tool the skill depends on.fallback— another tool name to substitute, or the literalmanualfor "a human does it instead."call-once-per-conversation— whentrue, the admission chain'sICallOnceGaterefuses a second call to this tool within the same conversation, durably — surviving across turns, across separate runs continuing the same conversation, and across hosts. Absent (orfalse) is the default: an ordinary, repeatable tool. Enforcement itself is a separate host opt-in — see Autonomy & Governance — so a skill can declare a tool call-once well before a host chooses to enforce it.
call-once-per-conversation for a side effect that can't repeat, not as a rate limitIt's the right fit for a tool that opens a stateful session or triggers a non-idempotent external action — something genuinely unsafe to run twice in one logical conversation. It is a hard guarantee enforced by the harness itself, not a hint the model might forget: even a confused or adversarially-prompted model cannot call the tool a second time once the first call has been recorded.
Enforcement is process-global by tool name, not scoped to this skill. If a different skill resolves a tool that happens to share this exact name, its calls are refused too — safe for the common case, since a first-party keyed-DI tool name is meant to identify one capability across the whole harness, but the wrong choice if two genuinely unrelated tools could ever share a name.
How a skill becomes an agent
The runtime journey of a skill, end-to-end:
-
Discovery at startup
SkillMetadataParserparses eachSKILL.mdfrontmatter into aSkillDefinition, andSkillMetadataRegistry(behindISkillMetadataRegistry) holds them, keyed by id. Only the Tier-1 metadata is eagerly in memory — the Markdown body is held lazily. -
Selection and assembly per turn
When
ExecuteAgentTurnCommandHandlerruns, it looks up the agent's skill references. For a single-skill agent, one skill is selected. For a multi-skill agent,AgentExecutionContextFactorymerges instructions and combines tool lists from all referenced skills into one unifiedAgentExecutionContext. If any skill has prerequisites, the factory enforces ordering — prerequisite skills must be marked complete (via theirCompletionTool) before dependent skills activate. -
Tier-2 promotion
The selected skill's full instructions (Markdown body) are loaded and stitched into the system prompt. Now the agent knows how to play this role.
-
Tool resolution
For each entry in
allowed-tools, the harness resolves a keyed singleton from DI —sp.GetRequiredKeyedService<ITool>("file_system")— then converts it to anAITooland attaches it to the agent. See Tools & Keyed DI. -
Tier-3 access at runtime
A resource under the skill folder's
references/subdirectory isn't read until the agent'sfile_systemtool actually opens it. The path is exposed via the tool, not preloaded into the prompt.
Skill modes — Managed vs Injected
Not all skills originate inside the harness. The SkillMode enum distinguishes two patterns:
| Mode | Tool resolution | When to use |
|---|---|---|
| Managed (default) | Only the tools listed in allowed-tools are resolved from keyed DI or MCP. The harness controls exactly what the agent can call. |
Skills you author inside the harness — full control over tool surface. |
| Injected | All MCP tools from the skill's parent plugin are passed through automatically, bypassing allowed-tools declarations. |
Skills provided by local plugins — the plugin author controls the tool surface. |
Injected skills always set PluginSource to the owning plugin's ID. The harness still applies
plugin-boundary governance (AllowedTools / DeniedTools) even when tools are passed through —
see Observability & Safety.
Prerequisites and completion tracking
In a multi-skill agent, some skills depend on the output of others. A "data gathering" skill should finish before an "analysis" skill starts. The harness supports this with two properties:
Prerequisites— a list of skill IDs that must complete before this skill activates.CompletionTool— the tool name whose invocation marks this skill as complete. When the agent calls this tool, the harness records the skill as done, unlocking any dependents.
---
id: analysis
name: Analysis Agent
prerequisites: [data-gathering] # must complete first
completion_tool: submit_report # calling this marks analysis as done
allowed-tools:
- submit_report
- file_system
---
Prerequisites enforce that one skill finishes before another starts — they don't control the order of tool calls within a skill. The agent still decides how to use its tools on each turn. For fine-grained step control, use the DAG plan executor instead.
Multi-skill agents
An agent's AGENT.md can reference multiple skills in its skills: list.
At context assembly time, AgentExecutionContextFactory merges instructions from all
referenced skills and combines their tool lists into a single AgentExecutionContext.
---
skills:
- data-gathering # prerequisite — runs first
- analysis # depends on data-gathering
- code-reviewer # independent — can run anytime
---
The factory handles each skill according to its SkillMode: Managed skills get
explicit tool resolution; Injected skills get pass-through MCP tools. Prerequisites are
enforced across the merged set. The result is one agent with the combined capabilities of
all its skills.
The context budget tracker
While all of the above is happening, IContextBudgetTracker is keeping score. For
the current turn, it tracks how many tokens have been committed to each of four things:
- the system prompt
- the skill content that has been loaded
- the tool schemas
- the conversation history
Those four compete for the same finite budget. As it starts to run out, the assembler stops promoting full Tier-2 skill content and serves only the short Tier-1 descriptions instead — the agent still knows the skill exists, it just no longer carries the full text.
IContextBudgetTracker governs a single turn, sized by
AppConfig:AI:AgentFramework:DefaultTokenBudget (default
200,000 tokens). A separate IConversationBudgetTracker
(a singleton) enforces a cross-turn ceiling for the whole conversation
(ConversationTokenBudget, default 1,000,000
tokens) and gracefully breaks the loop when a long session runs out. If you're seeing skills "disappear" mid-conversation, the
per-turn tracker is preserving budget; if a run stops cleanly after many turns, that's
the conversation tracker. (Note: there is no MaxTurnsPerConversation
property — budgets are token-based, not turn-count-based.)
Adding a new skill
The end-to-end recipe — covered in more detail on Extending the Harness:
- Create a folder under
skills/<your-skill-id>/. - Add a
SKILL.mdwith frontmatter (id, name, description, allowed-tools) and a Markdown body with the instructions. - Add any Tier-3 references under
references/,templates/, orscripts/in that folder. - Make sure every tool in
allowed-toolsis registered in DI under that key. - (Optionally) reference the skill from an
AGENT.mdso an orchestrator can route to it.
You don't need to redeploy or recompile. No restart needed — the harness watches skill paths for changes by default (AI:Skills:WatchForChanges) and picks up an added, edited, or removed SKILL.md automatically within ChangeDebounceMilliseconds (500ms default). If WatchForChanges is disabled, restart the host or call the operator refresh endpoint (POST /api/skill-registry/refresh) to confirm immediately.
Skill amendments and learnings
Skills aren't static. A SkillAmendment is a short, plain-language note attached to one skill — what was learned, and what triggered it. Unlike a SKILL.md change, an amendment doesn't need a host restart: AgentExecutionContextFactory loads a skill's amendments from the knowledge graph, entirely independent of the file-based skill loader, whenever that skill's agent is (re)built — immediately for any brand-new conversation, and for an already-open one once its cached agent is rebuilt, the same cadence the rest of that agent's static instructions already follow. It is not re-checked on every turn of an ongoing conversation.
The harness also tracks, automatically, whether each skill succeeds — broken down by what kind of request it was handling — so the data needed to decide "this skill needs an amendment" is being collected on every turn. Turning that track record into a written amendment is currently a manual step: nothing in the harness yet decides on its own when a pattern is strong enough to act on.