← Phase 8 Phase 9 · Monitoring Next: Phase C →

Home › Phase 9 · Monitoring

Phase 9 · Monitoring

Monitoring, explained How a live system becomes a watched one — and how the engagement writes down what it learned — the idea beside the real example (expand any section for the full method), the complete worked example, and a quick reference.

How it works the idea beside the real example — expand any section for the full method · Example the complete Harbor artifacts · Steps the same procedure, no company, just the plugin · Reference the quick mechanics

The system is live. The hard part isn't keeping it running today — it's making sure that when something breaks at 3 a.m., the right person wakes up, and knows exactly what to do.

Phase 9 is where a running system becomes a watched one — and where the whole engagement finally writes down what it learned.

Every idea below is paired with the real thing — examples come from a fictional but fully worked engagement, Harbor Mutual, a regional insurer.

01

Why this phase exists at all

Live is not the same as watched

Going live proves the system can run. It says nothing about whether anyone would notice if it quietly stopped. A system nobody is watching is an outage waiting to be discovered by a furious customer instead of an alert.

The idea

Phase 9 turns "it's running" into "we'd know within minutes if it weren't." That means dashboards someone actually reads, alerts that wake the right person at the right urgency, and a written plan for what to do when one fires.

And it does this on purpose, during hypercare — the agreed window after go-live when the client's operators are at the controls and the pod is still beside them. That timing isn't an accident: it's the only moment when the system is finally producing the one thing every alert needs — real numbers about what "normal" looks like.

At Harbor Mutual

Harbor's claims system went live on a Thursday. By the following Monday it was handling real FNOL intake — but nothing paged anyone. The dashboards didn't exist yet; the alerts were still a wish.

Where the phase starts
"Live since 7/23. Hypercare running. The first real numbers exist; nothing pages anyone yet."

Two weeks — 7/27 to 8/7 — ran inside that hypercare window, ending the day it closed.

Skip this phase and…

The first time you learn the system is broken is when a customer calls — and by then it's been broken for hours, with no record of how to fix it.

Go deeper — the full method

Going live proved the system can run; it said nothing about whether anyone would know if it quietly stopped. Phase 9 makes production observable, makes the alerts real, and writes down what the engagement learned. It runs inside the hypercare window on purpose — the system is finally producing the one thing alert thresholds must be built from (real baseline data), and the pod is still in the room while the client's operators take the controls.

The phase has a second job that matters as much as the first: the engagement retrospective. The weekly Retro+ answered "which check should have caught it?" all through Build; Phase 9 asks the cumulative version — what worked, what didn't, what the next engagement inherits — and produces the raw material for the harvest that closes the loop on the standard itself.

Skip this phase and…

The first time you learn the system is broken is when a customer calls — and by then it has been broken for hours, with no record of how to fix it. A system nobody watches is an outage waiting to be discovered the most expensive way.

02

The conversation that shapes everything

Decide what "healthy" means — together

You can't alert on "something's wrong" until you've agreed what "fine" looks like. And that question can't be answered by either side alone: the pod knows what the system does; the client's operators know what 3 a.m. is like in their company.

The idea

The defining event of the phase is a session where the pod and the client's operations walk through every way the system can fail and every journey that matters, and answer three things for each: what healthy looks like, what degraded looks like, and who gets woken versus who just gets told in the morning.

The operators' answers win ties — they're the ones being paged. Knowing what must not page is as important as knowing what must: an alarm that cries wolf every night trains people to ignore the one that matters.

At Harbor Mutual

The deepest decision of the whole phase came from Harbor's side of the table. Their system checks claims against a copy of the core records that refreshes overnight — and is unavailable from 2 to 4:30 a.m. while it does.

The trap they spotted
"An alert that fires every night at 02:00 trains the on-call to ignore the one that matters."
What they decided instead
"Suppress during the refresh window — but add a check that the system recovered after 04:30. That's the part actually worth waking for."
If operations can't be in the room

You'll ship your idea of healthy in their tooling — which becomes noise the week you leave. Protect this session; it is the phase.

Go deeper — the full method

Phase 9 is the Pod Lead's phase jointly with the client's operations — because the central session of the week, "what does healthy mean?", cannot be answered by either side alone. The pod knows what the system does; operations knows what 3 a.m. is like in this company.

The session that shapes the phase: the pod and the client's operations walk the RUNBOOK's failure scenarios and the top-priority user journeys and answer, for each, what healthy looks like, what degraded looks like, and who should be woken for what. Operations' answers win ties — they are the ones being paged. Capture one line per scenario and journey — healthy looks like, degraded looks like, who is woken, who is told in the morning — in the monitoring configuration. That table is what the alert definitions are written from.

The monitoring scope lands in writing: which stack (theirs — we wire into what their team already watches), which dashboards, which alert channels, who owns each dashboard. A dashboard without a named owner is decoration.

If operations cannot co-author the thresholds

The phase produces our alerts in their tooling — which become noise the week we leave. The session is the phase; protect it like Phase 7 protected the cold runs.

03

What the two weeks must answer

Four questions — and deliberately nothing else

Phase 9 doesn't add features and doesn't migrate the client to fancier tooling. It exists to answer exactly four questions, and everything it produces serves one of them.

The idea — the four questions
  1. Can the client see the system? Dashboards with owners, every important feature observable, business numbers next to system numbers.
  2. Will the right person find out, at the right urgency? Alerts built from real numbers, routed to named people.
  3. Does the response actually work? A written plan, proven by drill — not assumed.
  4. What did the engagement learn? An honest retrospective: what worked, what didn't, what the next project inherits.
At Harbor Mutual — the four answers
  1. Three dashboards in Harbor's own tooling — system, application, and the business clock — each with a named owner.
  2. Six alerts, each routed to Harbor's on-call rotation, each tied to a real failure mode.
  3. A drill where every critical alert fired for real and Harbor's own on-call answered it.
  4. A retrospective that sent four patterns home to improve the next project.
The temptation to resist

It feels helpful to add "just one more dashboard" or swap in a better monitoring tool. Don't. You wire into the stack the client's team already watches — anything else is a thing they won't maintain.

Go deeper — the full method

Phase 9 answers four questions, and nothing else:

  1. Can the client see the system? Dashboards with owners, every top-priority feature observable, business metrics alongside system metrics.
  2. Will the right person find out, at the right urgency? Alerts derived from measured baselines, routed to named people, covering every critical failure mode the RUNBOOK describes.
  3. Does the response actually work? The incident playbook, proven by drill — detected, diagnosed, communicated by the client's own on-call.
  4. What did the engagement learn? The honest retrospective: product, process, debt, and the patterns the kit inherits.

New features, alert tooling migrations, and the formal handover of the harness are out of scope — the backlog stays closed, the client's existing monitoring stack is the one we wire into, and Close & Transfer is the next phase's job. Defects surfaced by hypercare ride the loop, as ever.

04

The number behind every alarm

Thresholds come from measured reality — not from a number that "felt right"

"Alert me if response time goes over 500 milliseconds" sounds rigorous. It's a guess. Unless you know what response time actually is on a normal day, that number is arbitrary — and arbitrary thresholds either cry wolf or stay silent through real trouble.

The idea

Before any alert is written, you capture the baseline: what normal looks like for each key metric, measured from real production traffic, each number recorded with the period it was measured over. Every threshold is then derived from it — "twice the measured normal, sustained ten minutes" — and the derivation is written down next to the number.

And when production hasn't yet exercised a path — the rare storm, the seasonal surge — you don't guess. You derive from the engagement's earlier modeled data, flag it as modeled, and set a date to revisit. Honest beats tidy.

At Harbor Mutual

The team measured the real baseline from the first hypercare week — and caught that production ran slightly better than the old design-phase estimate:

~180ms165ms the design-phase estimate → the measured production p95 for verification

One path had no real numbers: late July gave Harbor no storm. The surge thresholds were derived from a modeled 2024 catastrophe dataset and flagged modeled, revisit at the first real storm — with a named person on the revisit.

The headline failure mode of Phase 9

Thresholds set by intuition. "500ms feels right" is not engineering. A threshold that can't explain where it came from can't be tuned later — it just gets ignored or deleted.

Go deeper — the full method

The production baseline gets captured from the first hypercare weeks' real traffic: request rates, latency percentiles, error rates, queue depths, dependency health — each recorded with the period it was measured over. Where production hasn't yet exercised a path (the seasonal surge, the rare failure), the threshold derives from the engagement's modeled data instead (whatever the engagement already measured: the hardening passes' load-test results, the design phase's spikes) — flagged as modeled, with a revisit date, never silently presented as baseline. Claude measures, the Setup Owner runs the capture in the client's tooling, and the Quality Engineer owns the resulting artifact — including the honesty of every modeled flag.

Every threshold is then derived from the baseline — alert at a stated multiple of normal, never at a number that felt right — and the derivation is written next to the number. "2x the measured p95 over 10 minutes" survives an argument; "500ms" does not.

Thresholds by intuition

"500ms feels right" is not engineering. Measure the baseline, alert at a stated multiple, write the derivation down. A threshold that can't explain itself can't be tuned later — it just gets ignored or deleted.

05

Why fewer alerts are better

An alert nobody acts on is noise wearing a badge

Fifty alerts feels thorough. It guarantees the team ignores all of them within a month — and then misses the real one. Every alert must demand a human do something; anything else is just decoration that erodes trust in the whole system.

The idea

Before any alert ships, it's tested against history: replay each proposed condition over the hypercare data and count how often it would have fired. The standing rule — anything that would fire more than once a week without demanding action gets its threshold raised or gets cut. Now, before the team learns to tune it all out.

This is the discipline of restraint. Six alerts people answer beat eight they've trained themselves to swipe away.

At Harbor Mutual

The fatigue review cut two proposed alerts before they shipped:

Per-instance CPU alert — cut. Not actionable; the platform autoscales, and capacity is already on a dashboard.

Portal latency warning — cut. It duplicated the signal the intake error-rate alert already gave.

Six actionable alerts beat eight ignorable ones.

Go deeper — the full method

An alert is a condition that demands a human response. Every alert is actionable; a notification nobody acts on is noise wearing an alert's badge. Alert fatigue is what happens when alerts fire often and mean little: the team learns to ignore them, and then misses the real one.

The alert-fatigue review runs on everything proposed: replay each proposed condition against the hypercare metrics history — the same data the baseline came from — and count the firings. The standing rule: anything that would have fired more than once a week during hypercare without demanding action gets raised or cut now, not after the team has learned to ignore it. Two thresholds per alert make this precise: warning means investigate during working hours; critical means wake someone up. The difference is who suffers if it waits until morning.

Fifty alerts feels thorough and guarantees the team ignores all of them within a month. Fewer, actionable, derived, reviewed — six alerts people answer beat eight they've trained themselves to swipe away.

06

What to do when it fires

An alert tells you something's wrong; the playbook tells you what to do about it

A pager going off at 3 a.m. is useless if the person it woke doesn't know what the alert means, how to diagnose it, who to call, or what to tell upset customers while they work. The alarm is half the system; the response plan is the other half.

The idea

The incident playbook is the detect-diagnose-communicate companion to the existing repair guide. For each alert it spells out what it means, the first diagnosis steps, who escalates to whom, and the templates for what users and leadership hear while it's happening.

It deliberately doesn't repeat the repair steps — it points to the existing guide for those. One source of truth per failure, two lenses on it: the playbook detects and communicates; the runbook resolves.

At Harbor Mutual

Harbor's on-call lead corrected the draft playbook line by line. The escalation names were all Harbor's — the pod appears only as a last resort:

Where the pod sits in the playbook
"Escalation of last resort until Close."

It spelled out exactly what the claims office hears from the intake team when intake degrades, and what leadership hears on a sev-1 — and who says it.

Go deeper — the full method

The playbook gets written against the alert table: per alert, what it means, the first three diagnosis steps, the severity classification, the escalation names, and the communication templates — what users and stakeholders are told, by whom, while it's happening. It cross-references the RUNBOOK rather than duplicating it: the playbook detects and communicates; the RUNBOOK resolves.

The incident playbook is the detect-diagnose-communicate companion to the RUNBOOK: what each alert means, first diagnosis steps, who escalates to whom, and what to tell users while it's happening. One source of truth per failure, two lenses on it — which is why the playbook never repeats the repair steps; it points into the cold-verified RUNBOOK for resolution.

07

Proving it actually works

An alert that has never fired is a wish

Routing rules, paging chains, and playbook steps all look correct on paper. The only way to know they work is to make each one happen on purpose — and watch the client's own on-call respond. Discover the broken pieces by appointment, at drill prices, not at 3 a.m. during a real incident.

The idea

Every critical alert is triggered deliberately, in a controlled way, and the client's on-call responds from the playbook while the pod watches in silence. "Controlled" means designed: before drill day, you write down one safe synthetic trigger per alert — a test lane, flagged test data, a deliberately blocked dependency — that makes the real condition true without touching real traffic. No agreed trigger, no drill.

What breaks gets fixed and re-drilled until it's clean. This is the same discipline that proved the rollback in the previous phase: the difference between a procedure and a hope.

At Harbor Mutual

The drill earned its keep. Two findings, both exactly what a drill is for:

The silent-night bug
"VERIFY-DEGRADED routed to the general ops channel instead of the pager rotation — a routing-key typo that would have meant a silent night during a real incident."
Caught and fixed at drill prices
"Routing key corrected by reviewed PR. Re-fired 11:05: paged in 38 seconds, responder followed the playbook, closed clean."

A second finding — a playbook step that opened a dashboard needing pod permissions — was caught and re-drilled clean the same day.

Go deeper — the full method

Each critical alert fires for real, triggered in a controlled way, and the client's on-call responds from the playbook while the pod observes silently — the same discipline as Phase 7's cold runs and Phase 8's rehearsal. Controlled means designed: before drill day, the Quality Engineer and the client's platform engineer write down one synthetic trigger per critical alert — a test lane, flagged test data, a test toggle, a deliberately blocked dependency — that makes the real alert condition true without touching real traffic or real data. No trigger agreed in writing, no drill.

Routing that goes to the wrong channel, a playbook step that assumes pod access, a threshold that doesn't actually trigger: all of it fails here, by appointment, at drill prices. What the drill breaks gets fixed through the loop and re-drilled. The Quality Engineer records, per alert: the trigger used, trigger time, detection time, where it routed, who responded, and the outcome — pass, or the finding and its fix. That record goes in the gate packet.

The drill is to alerts what the rehearsal was to rollback

An alert that has never fired is a wish. The difference between a procedure and a hope is that someone has run it — deliberately, in advance, at drill prices.

08

What the whole engagement learned

The honest retrospective — argued from receipts, not vibes

Phase 9's second job matters as much as the first. With the engagement nearly over, this is the moment to ask the cumulative question: what worked, what didn't, and what the next project should inherit. The trap is the feelings-meeting version — "great teamwork, communicate better next time" — written for management and useless to everyone.

The idea

The AI assembles the evidence — the metrics history, the weekly logs, every escaped bug with its answer — so the humans argue from facts, not vibes. The AI never softens the retro; a flattering drafter produces a worthless one. The candor is the humans' job.

It produces three concrete outputs: a technical debt log (every known shortcut, with a priority and a date — logged debt is managed, unlogged debt is next year's crisis), client-facing improvements, and the harvest list — the patterns and fixes the next engagement inherits.

At Harbor Mutual

The retro argued from receipts, not adjectives. Not "security review was slow" but the number, and a structural fix instead of an exhortation:

The finding, with its evidence
"The security-review queue drifted all engagement — 2.1-day median against a 0.9-day general queue."
The concrete fix
"A twice-weekly committed security-review slot in Dan's calendar, written into the Close handoff."

Four patterns went onto the harvest list to improve the next project — among them the suppression-window-plus-recovery-check alert pattern from this very phase.

Go deeper — the full method

A half-day, the pod plus the client's core people, arguing from the assembled evidence: what worked (with the receipts), what didn't (without blame), which gates caught real issues and which were rubber-stamped, which artifacts earned their cost and which were written and never read. Claude assembles the evidence base — the metrics dashboard history, the Retro+ log, the escaped-bug answers, the gate records — so the humans argue from facts, not vibes. An honest retro written by a flattering drafter is worthless; the humans own the candor.

Three outputs, all concrete: the technical debt log (every known shortcut, with a priority and a suggested timing — logged debt is managed, unlogged debt is a future crisis), the client-facing improvements (what the client's team changes about how they run the system and the loop), and the harvest list — the patterns, skills, hook improvements, and template corrections the next engagement inherits through the kit. Phase C opens the harvest PR; this day decides what goes in it.

The platitude retrospective

"Great teamwork, communicate better next time" — written for management, useless to everyone. Honest, specific, evidence-backed, or skip the meeting and admit it.

09

Now watch the whole thing happen

The two weeks, end to end

You've got the ideas; here's the actual rhythm at Harbor, running inside the hypercare window. Week one makes the system visible — agree what healthy means, capture the baseline, build the dashboards and the alerts. Week two proves the response and learns — write the playbook, drill it, run the retrospective, hand over the watch. Step through it.

Days 1–2 · what does healthy mean?

The session that shapes the phase

The pod and the client's operations walk every failure scenario and every top-priority journey, answering for each: what healthy looks like, what degraded looks like, who gets woken, who's told in the morning. Operations' answers win ties. The monitoring scope — which stack, which dashboards, which owners — lands in writing.

Days 3–4 · baseline & dashboards

Measure normal; make it visible

Capture the production baseline from the first hypercare week — request rates, latency, error rates, queue depths — each with its measurement period. Paths production hasn't exercised get modeled values, flagged with a revisit date. Dashboards go live in the client's stack: system, application, and the business layer in their own reporting language.

Day 5 · the alert definitions

Every threshold derived, and confirmed

Each critical failure mode gets an alert with a baseline-derived condition, a severity, a named recipient, and a playbook link — the derivation written next to the number. The operators confirm every threshold; that's the named human stop of the phase. The fatigue review cuts anything that would fire weekly without demanding action.

Days 6–7 · the incident playbook

Write down what to do when it fires

Per alert: what it means, the first diagnosis steps, the severity classification, the escalation names, and the communication templates — what users and stakeholders are told, by whom, while it's happening. It cross-references the repair guide rather than duplicating it.

Day 8 · the drill

Fire every alert on purpose

Each critical alert fires for real through a pre-agreed synthetic trigger, and the client's on-call responds from the playbook while the pod observes in silence. Wrong routing, a broken playbook step, a threshold that doesn't trigger — all of it fails here, by appointment. What breaks gets fixed through the loop and re-drilled.

Day 9 · the retrospective

The honest cumulative look

Half a day, the pod plus the client's core people, arguing from assembled evidence: what worked with receipts, what didn't without blame, which gates caught real issues. Three concrete outputs — the technical debt log, the client-facing improvements, and the harvest list the next engagement inherits.

Day 10 · hypercare ends, the gate

Hand over the watch; read the metric honestly

Hypercare closes on its agreed date with a deliberate handover: the dashboards, the pager, and the playbook are formally the client's, with the pod one escalation away until Close. The gate runs; the sponsor signs at steering — and hears the outcome metric's first honest production read, caveats welded on.

1 / 7
Go deeper — the full method

The default calendar is two weeks, deliberately inside hypercare — week one needs the production baseline to accumulate; week two needs the pod still present for the drill. The gate falls at hypercare's end, closing both together.

Week one — see the system

  • Days 1–2, what does healthy mean? The session that shapes the phase: the pod and operations walk the RUNBOOK's failure scenarios and the top journeys, capturing one line each — healthy, degraded, who is woken, who is told in the morning. The monitoring scope lands in writing: which stack (theirs), which dashboards, which channels, who owns each dashboard.
  • Days 3–4, the baseline and the dashboards. The production baseline is captured from real traffic, each number with its measurement period; unexercised paths derive from modeled data, flagged with a revisit date. Dashboards go live in the client's stack: system health, application health, and the business layer co-built with their data lead in their own reporting language.
  • Day 5, the alert definitions. Every critical RUNBOOK failure mode gets an alert with a baseline-derived condition, a severity, a named recipient, a response expectation, and a playbook link — the derivation written next to the threshold. Confirmation is a recorded act: the operations names go on the review of the change that ships the alert rules. The alert-fatigue review runs on everything proposed.

Week two — prove the response, and learn

  • Days 6–7, the incident playbook. Written against the alert table: per alert, meaning, first three diagnosis steps, severity classification, escalation names, and communication templates. It cross-references the RUNBOOK rather than duplicating it.
  • Day 8, the alert drill. Each critical alert fires for real through a pre-agreed synthetic trigger; the client's on-call responds from the playbook while the pod observes silently. What the drill breaks gets fixed through the loop and re-drilled; the Quality Engineer records every alert's trigger, detection, routing, responder, and outcome for the gate packet.
  • Day 9, the retrospective. A half-day arguing from assembled evidence, with three concrete outputs: the technical debt log, the client-facing improvements, and the harvest list the next engagement inherits.
  • Day 10, hypercare ends, the gate. Hypercare closes on its agreed date with a deliberate handover of the watch. The gate check runs; Claude drafts the Close & Transfer handoff. Steering: the gate sign-off, the billing milestone, and the outcome metric's first honest production read — stated with its caveats.
When the two weeks stretch

A too-thin baseline ships thresholds as modeled-with-revisit-date rather than waiting for a storm. A badly failed drill is fixed, re-drilled, and — if the gap is people, not config — flagged at steering. A political retrospective is held to "for the teams, not for performance reviews." A defect-heavy hypercare is a quality finding for the retro, not a reason to extend forever.

10

How the phase — and hypercare — ends

A gate, not a calendar — and it closes the watch with it

Phase 9 doesn't end because two weeks passed. It ends when a specific list is true and a named human on each side signs to advance — and because the gate falls at hypercare's end, signing it is also the moment the pod formally hands over the watch.

The idea

This is also a billing milestone, and it carries the engagement's most delicate moment: the outcome metric's first honest production read. The credibility protected all engagement gets spent or banked here — so the number is stated with every caveat attached, never oversold.

The headline checklist items have teeth: a drill where every critical alert actually fired and was answered by the client's own on-call, and a retrospective with real receipts, not platitudes.

At Harbor Mutual

The outcome metric read spectacularly — and was deliberately underclaimed:

11.41.9 days, all-claims baseline → fast-path claims (61% of volume) over the two live weeks

Steering heard it with the caveat welded on: complex claims are mostly still in flight, so the overall median can't be read fairly until a full quarter. The sponsor signed — billing milestone 7. The watch became Harbor's.

The full checklist — tick it:

  • Every top-priority feature has at least one metric on a dashboard with a named owner
  • Every critical failure mode has an alert, each threshold derived from the baseline (or flagged modeled) and confirmed by the client's operations
  • The alert-fatigue review ran — nothing ships that would fire weekly without demanding action
  • The incident playbook covers every critical alert — detection, diagnosis, escalation names, templates
  • The drill happened — every critical alert fired and was answered by the client's own on-call; what failed was fixed and re-drilled
  • The watch was handed over — dashboards, paging, and the playbook are formally the client's
  • The retrospective exists with receipts — findings, the debt log, and the harvest list
  • The outcome metric has its first production read on the scorecard, caveats stated
  • A named human on each side approved the advance
Go deeper — the full method

A phase does not end because two weeks went by. It ends when a specific list is true and a named human on each side signs to advance — gates report, humans decide. Phase 9 closes — and hypercare ends with it — when all of these are true:

  • Every top-priority feature has at least one observable metric on a dashboard with a named owner.
  • Every critical RUNBOOK failure mode has an alert; every threshold is derived from the measured baseline (or explicitly flagged as modeled, with a revisit date) and confirmed by the client's operations. (teeth)
  • The alert-fatigue review ran: nothing ships that would fire weekly without demanding action.
  • The incident playbook covers every critical alert — detection, diagnosis, escalation names, communication templates.
  • The drill happened: every critical alert fired in a controlled way and was answered by the client's own on-call from the playbook; what failed was fixed and re-drilled. (teeth)
  • The watch was handed over: dashboards, paging, and the playbook are formally the client's, with the pod one escalation away until Close.
  • The retrospective exists with receipts: findings, the technical debt log, and the harvest list — concrete, honest, owned.
  • The outcome metric has its first production read on the scorecard, caveats stated.
  • The Close & Transfer handoff exists: monitoring inventory, drill record, debt log, open items with owners.
  • A named human on each side approved the advance.

This is also a billing milestone, and it carries the engagement's most delicate moment: the outcome metric's first honest production read. The credibility protected all engagement gets spent or banked here — so the number is stated with every caveat attached, never oversold.

11

How it goes wrong

The failure modes, and the defense against each

Every one of these has happened to someone. Knowing them by name is half the defense — and Harbor's structure caught two of them in the act.

The trapWhat it looks likeThe defenseAt Harbor
Thresholds by intuition"500ms feels right" set without a measured baselineMeasure normal; alert at a stated multiple; write the derivation down.Every threshold derived; modeled ones flagged with a revisit
Our monitoring, their pagerAlerts built in a stack the client's team doesn't watchTheir stack, their names, their session.All wired into Harbor's own Azure Monitor workspace
Alert fatigue on day oneFifty alerts; the team ignores all of them within a monthReplay against history; cut anything that fires weekly without action.Caught — two proposals cut at the fatigue review
The undrilled pagerRouting and playbooks that have never actually firedFire each one on purpose; the client's on-call responds.Caught — the day-8 silent-night routing typo
The dashboard nobody ownsA beautiful screen no one reviewsEvery dashboard has a named owner or it doesn't ship.Three dashboards, each with a named Harbor owner
The platitude retrospective"Great teamwork, communicate better" — written for managementArgue from receipts; honest and specific or skip it."2.1 days vs 0.9," a calendar slot, not an exhortation
Hypercare that never endsQuietly extending the watch because closing feels riskyExtend deliberately with the sponsor, or close on schedule.Closed on its agreed date — two weeks, as written
Go deeper — the full method

Named failure modes, because every one of these has happened to somebody. Knowing them by name is half the defense.

  • Thresholds by intuition. "500ms feels right" is not engineering. Measure the baseline, alert at a stated multiple, write the derivation down. A threshold that can't explain itself can't be tuned later.
  • Our monitoring, their pager. The pod builds alerts in its own image, in a stack the client's team doesn't watch, routed by assumptions. The week we leave, it's noise. Their stack, their names, their session.
  • Alert fatigue shipped on day one. Fifty alerts feels thorough and guarantees the team ignores all of them within a month — and then misses the real one. Fewer, actionable, derived, reviewed.
  • The undrilled pager. Routing and playbooks that have never fired, discovered broken during the first real incident. The drill is to alerts what the rehearsal was to rollback: the difference between a procedure and a wish.
  • The dashboard nobody owns. A beautiful screen no one reviews is decoration. Every dashboard has a named owner or it doesn't ship.
  • Monitoring the system but not the business. All RED metrics, no outcome metric — the client can see requests but not whether the thing they bought is working. The business layer is co-built with their data lead, in their language.
  • The platitude retrospective. "Great teamwork, communicate better next time" — written for management, useless to everyone. Honest, specific, evidence-backed, or skip the meeting and admit it.
  • Debt left unlogged. The shortcuts everyone knows about but nobody wrote down become next year's crisis with no paper trail. Logged debt is managed debt; the log is a gate item for a reason.
  • Hypercare that never ends. Quietly extending the watch because closing feels risky. Extend deliberately with the sponsor and a new end date, or close on schedule — an open-ended hypercare is a handoff that's failing in slow motion.

When Phase 9 is done

The system can be seen, the response is proven, and the lessons are written down

Monitoring closes with a watched system in the client's own hands, a response drilled until it works, and an honest account of what the engagement learned — debt logged, patterns harvested. One thing remains: proving the client can run all of it without you. Where to go next:

01

Before a single alarm is wired, three things arrive — and the promises Phase 2 made come due

What Phase 9 received

Harbor Mutual — a fictional regional insurer — rebuilt how property-insurance claims get reported and decided. A claim took a median of 11.4 days from FNOL (first notice of loss) to a coverage decision; the target is 5 days or less. Production went live Thursday 2026-07-23; hypercare is running on its two-week window. Phase 9 does not start from a blank page — it starts from a live system, a handoff on disk, and a plan written months ago about where every number would be read.

Inherited — the three inputs Phase 9 is built on phase9-handoff.md RUNBOOK.md — the failure scenarios nfr-proving-plan.md — from Phase 2; the promises now due the live system — producing real numbers, not a document
The through-line: Phase 2 said where each number would be read

Back in design week, the quality engineer wrote an NFR proving plan — for every quality target, the method that would prove it and the named place its number would be read. Phase 9 is where those promises come due.

  • NFR-01 (ingestion p95 < 5s) — "read from the monitoring dashboard." That dashboard gets built this week.
  • NFR-02 (the 10x surge) — "read from the load-test report." Late July gave Harbor no storm, so this one stays modeled.
  • ADR-004 eval gate (≥95% email-extraction accuracy) — "read from the eval suite in CI." Phase 9 turns that into a live production watch.

A proving plan that named the reading spot but never got read would be a promise quietly broken. This is the week it gets kept.

Where Phase 9 starts

The system: portal / phone / email FNOL intake → a buffered claim queue → coverage verification against PolicyOne's nightly snapshot replica (unavailable 02:00–04:30 during refresh) → fast-path recommendations → acknowledgment dispatch through the postal vendor.

From phase9-handoff.md — where we are
"Live since 7/23 (cutover at intake; legacy fallback warm). Hypercare running. The first real numbers exist; nothing pages anyone yet."

Nobody new joins the story. The people who will answer Harbor's pager are the people who sat in every session — which is the whole point of the two weeks.

Our pod

Maya ChenPod Lead — this phase is hers, jointly with Harbor's operations
Nadia BrooksQuality Engineer — designs and owns the alert drill
Rob FeldSetup Owner — wires monitoring into Harbor's own stack
Jonah KimOrchestrator — drafts alert and dashboard config; fixes through the loop
Sara WhitfieldOrchestrator — drives the playbook draft; fixes through the loop

Harbor Mutual

On-call lead & operatorsOwn everything this phase produces; answer the drill from the playbook
Tom ReillyPlatform engineer — wires channels and paging; agrees the synthetic triggers
Priti ShahData & reporting — the business dashboards; named on the surge-threshold revisit
Dan KowalskiIT security — the security alert lane
Luis OrtegaProduct owner — confirms the business metrics measure what the business means
Dee AlvarezIntake supervisor — the intake-degraded communication path
Karen VossVP Claims Operations — sponsor; hears the first honest metric read
The ID codes, decoded

Every artifact in this engagement carries a stable identifier, so a promise made in design week can still be traced when its number comes due in production. The prefixes Phase 9 touches:

PrefixMeansBorn inExample here
NFR-NNA non-functional requirement — a quality target with a numberPhase 1NFR-01 ingestion p95 · NFR-02 the 10x surge
REQ-NNNA functional requirementPhase 1REQ-014 same-business-day coverage status
ADR-NNNAn architecture decision record — a signed choicePhase 2ADR-004 email-extraction eval gate
C-NNA constraint — a hard limit the system must honorPhase 0C-04 the acknowledgment regulatory clock
D-NNA product decisionPhase 0/1D-09 the fast path (61% of claims)
Q-NNAn open question with an owner and a due dateanyQ-18 the surge load-test dataset

The thread tying design to production: Phase 2's proving plan named, per NFR, the place its number would be read. Phase 9 either reads it — on a dashboard, in an alert, on the scorecard — or flags it modeled with a revisit date. A proving plan whose numbers never got read is homework nobody graded.

02

The plugin's seven steps, the pod's ten days, braided

The procedure, step by step

Phase 9 is seven numbered steps in claude-code-sdlc and ten working days in this standard, run inside the hypercare window. Below, they're braided: what the tool runs, what the humans do that the tool cannot, and the file each beat leaves behind. Two of the most important beats — the drill and the what-healthy session — have no command at all. Step through it.

Legend a command does it — and writes the file a person does it — and it is recorded a person does it — and nothing records it
Days 1–2 — Mon–Tue 7/27–28 · plugin Step 0 opens

Agree what "healthy" means, together, before anything is wired

The plugin's Step 0 is a blocking human gate: before configuring any monitoring, Claude asks the client, via AskUserQuestion, the top failure modes, who gets paged and at what threshold, and which stack the alerts live in. But those four questions are the thin version. The phase's defining session is deeper and has no command: the pod and Harbor's operations walk every RUNBOOK failure scenario and every top-priority journey and answer, for each, what healthy looks like, what degraded looks like, and who gets woken versus who's told in the morning.

Tooling /sdlc-next Step 0 HITL gate the what-healthy session — humans in a room, no command
Out the healthy / degraded / who-is-woken table — decided in the room, not yet on disk
At Harbor

The deepest decision of the phase came from Harbor's side of the table. The replica is unavailable 02:00–04:30 nightly by design (REQ-014 degrades to "pending verification") — so an alert that fires every night at 02:00 would train the on-call to ignore the one that matters. They chose a suppression window plus a post-04:30 recovery check: don't page for the expected outage; do page if it fails to come back. Scope landed in writing — everything wires into Harbor's Azure Monitor workspace, where Tom's team already lives, paging through their existing rotation.

The gap you should know about

The session that shapes the whole phase — the healthy/degraded/who-is-woken table — is the standard's core human work, and no command captures it. The plugin's doc-updater later writes monitoring-config.md from the measured baseline, not from this session's transcript. The table only reaches disk if a human types it there. The phase's most important input has no receipt of its own.

Days 3–4 — Wed–Thu 7/29–30 · plugin Step 1

Measure what normal actually is; make it visible

Now the machine runs. Claude spawns the performance-benchmarker agent to establish the production baseline against the Phase 1 NFR targets, then the doc-updater agent to write it up as the monitoring configuration. Every number is recorded with the period it was measured over. Dashboards go live in Harbor's own workspace: system health, application health, and — built with Priti in Harbor's reporting language — the business layer.

Tooling Agent performance-benchmarker Agent doc-updater monitoring-config.md
Out — under .sdlc/artifacts/09-monitoring/ monitoring-config.md baseline measurements (folded in)
At Harbor

The measured baseline caught that production ran slightly better than the design-phase estimate: replica verification p95 came in at 165ms against Phase 2's ~180ms spike number. Portal submit p95 410ms, intake error rate 0.3%, queue depth steady under 40. One honest gap, written down not papered over: the surge path has no baseline — late July gave Harbor no storm — so its thresholds derive from Q-18's modeled 2024 CAT dataset, flagged modeled with Priti named on the revisit.

Day 5 — Fri 7/31 · plugin Step 2

Turn each failure mode into an alert — and cut the ones nobody would act on

Claude drafts the alert table against the phase template — every critical failure mode gets a condition derived from the baseline, a severity, a named recipient, and a playbook link, with the derivation written next to the number. Then two pieces of human work the plugin does not perform: the operators confirm every threshold (the named human stop of the phase), and the alert-fatigue review replays each proposed condition over the hypercare history and cuts anything that would fire weekly without demanding action.

Tooling /sdlc-coach alert-definitions.md threshold confirmation & fatigue review — human work, no command
Out alert-definitions.md operations' confirmation on every threshold the fatigue-review record — the method requires it; nothing writes it
At Harbor

Six alerts shipped; the fatigue review cut two before they went live — a per-instance CPU alert (not actionable; the platform autoscales) and a portal-latency warning (it duplicated the intake error-rate signal). "2x the measured p95 over 10 minutes" survives an argument in six months; "500ms" would not. Six actionable alerts beat eight the on-call learns to swipe away.

Days 6–7 — Mon–Tue 8/3–4 · plugin Step 3

Write down what to do when it fires

Claude drafts the incident playbook against the alert table and the Phase 7 RUNBOOK: per alert, what it means, the first three diagnosis steps, the P1/P2/P3 classification, the escalation names, and the communication templates — what users and leadership hear while it's happening, and who says it. It cross-references the RUNBOOK for resolution rather than repeating it. Harbor's on-call lead corrects it line by line.

Tooling /sdlc-coach incident-response.md the on-call lead's line-by-line correction — human
Out incident-response.md
At Harbor

Every escalation name in the playbook was Harbor's own; the pod appears once, at the bottom of the chain — "escalation of last resort until Close." The playbook spelled out exactly what the claims office hears from Dee's intake team when intake degrades, and what leadership hears on a P1. One source of truth per failure, two lenses on it: the playbook detects and communicates, the RUNBOOK resolves.

Day 8 — Wed 8/5 · no plugin step exists

Fire every critical alert on purpose — and watch the client answer it

Each critical alert is triggered for real, one at a time, through a synthetic trigger agreed with Tom in advance — replica reads blocked outside the window, a retry storm on the dispatch test lane, flagged test messages pushed past the queue threshold, a staleness clock wound forward. Harbor's on-call responds from the playbook while Nadia observes in silence. What breaks gets fixed through the loop and re-drilled until clean.

Tooling no command — the drill is human work, and Step 4 says so: Claude prepares the plan and writes the record, but cannot page anyone
Out drill-record.md — required, and Step 4 produces it
The widest gap in the phase — now closed

The exit gate carried a teeth condition — "Alert drill executed: every critical alert fired and answered from the playbook" — while the workflow ran Step 0 through Step 6 and never mentioned a drill, no command triggered one, and drill-record.md was listed optional. The single most valuable act of the phase was required by the gate and produced by nothing. The plugin now ships Step 4: Alert Drill — placed after the playbook, so the drill tests incident-response.md as much as it tests the alert — with a spec, a template, and the artifact promoted to required. This is exactly the work that caught Harbor's silent-night bug.

At Harbor

The drill earned its keep. VERIFY-DEGRADED routed to the general ops channel instead of the pager rotation — a routing-key typo that would have meant a silent night during a real incident. Fixed by reviewed PR, re-fired 11:05, paged in 38 seconds, closed clean. A second finding — a playbook step that opened a dashboard needing pod permissions — was caught and re-drilled clean the same day. Both fixes rode the full loop: specs, the grader, a non-author Checker, in a monitoring week.

Day 9 — Thu 8/6 · plugin Step 4

Ask the cumulative question: what did the whole engagement learn?

Claude spawns the feedback-synthesizer agent in the background to assemble the hypercare findings and intake-team feedback, and /visual-explainer renders the evidence pack the room argues from. Then the half-day itself, which no command can do: the pod plus Harbor's core people argue from receipts — the metrics history, the Retro+ log, every escaped bug with its "which check should have caught it?" answer, the gate records. The AI assembles the evidence; the humans own the candor.

Tooling Agent feedback-synthesizer /visual-explainer the evidence pack the candor — human; a flattering drafter produces a worthless retro
Out project-retrospective.md
At Harbor

The retro argued from receipts, not adjectives. Not "security review was slow" but 2.1-day median against a 0.9-day general queue — and a structural fix, a twice-weekly committed slot in Dan's calendar, not an exhortation. Three concrete outputs: the technical debt log, the client-facing improvements, and a four-item harvest list — among them the suppression-window-plus-recovery-check pattern born this very week.

Day 10 — Fri 8/7 · plugin Steps 5–6 · the gate

Hand over the watch, run the gate, read the metric honestly

The machine renders the visual report and runs the gate. check_gates.py reports; advance_phase.py will not move the engagement to Close without --confirmed — a named human's sign-off. Because the gate falls at hypercare's end, signing it is also the moment the pod formally hands over the watch: the dashboards, the pager rotation, and the playbook become Harbor's, with the pod one escalation away until Close.

Tooling /visual-explainer .sdlc/reports/phase09-visual.html /sdlc-gate check_gates.py /sdlc-phase-report generate_phase_report.py /sdlc-next advance_phase.py --confirmed
Out .sdlc/reports/phase09-report.html close-handoff.md — the registry requires it; the phase body never says what it is the sponsor's signature — billing milestone 7
What the gate actually checks

check_gates.py verifies that five files exist, are non-empty, and contain no placeholder text. The two conditions with real teeth — every threshold derived from a measured baseline and the drill executed — are free-text check: lines the script never evaluates. And one of the five required files, close-handoff.md, has no Artifact Specification anywhere in the phase body — the registry demands a file the instructions never describe. The human gate is real; what it verifies is thinner than what the standard asks.

At Harbor

The outcome metric read spectacularly — and was deliberately underclaimed.

11.41.9 days, all-claims baseline → fast-path claims (61% of volume) over the two live weeks

Steering heard it with the caveat welded on: complex claims are mostly still in flight, so the overall median can't be read fairly until a full quarter. C-04 acknowledgment compliance: 100%, system-enforced. Karen signed — billing milestone 7. The watch became Harbor's.

1 / 7
03

Everything that exists on Friday and didn't two weeks earlier

What Phase 9 produced

The monitoring phase's whole output, named. Blue rows are written by a command and checked by the gate. Amber rows are the method's human work — required by this standard (and in one case by the plugin's own gate), produced by no tool, and today leaving no file behind. Those are the rows to argue about.

ArtifactWhat it actually isWritten bySigned byLives atFeeds
monitoring-config.mdDashboard inventory, metrics catalog, P0 coverage assessment, and the measured baseline — each number with its measurement periodperformance-benchmarker measures; doc-updater writesQuality Engineer.sdlc/artifacts/09-monitoring/Close — the monitoring inventory Harbor inherits
alert-definitions.mdThe six alerts: condition and derivation, severity, recipient, playbook link; per-critical detail; the alert philosophyClaude drafts; operations confirm every thresholdClient operations.sdlc/artifacts/09-monitoring/The playbook; the drill
incident-response.mdPer alert: meaning, first diagnosis steps, P1/P2/P3, escalation names, communication templates; cross-references the RUNBOOKClaude drafts; on-call lead correctsClient operations.sdlc/artifacts/09-monitoring/The drill; Close
project-retrospective.mdWhat worked and didn't with receipts, the SDLC review, the technical debt log, and the harvest list — "the most important Phase 9 artifact"feedback-synthesizer assembles; the humans own the candorPod Lead.sdlc/artifacts/09-monitoring/The harvest PR (Phase C)
phase09-report.html
phase09-visual.html
The gate result and artifact inventory, self-contained — the document a sponsor actually reads before signinggenerate_phase_report.py · /visual-explainer.sdlc/reports/The manual sign-off gate
drill-record.mdPer critical alert: the trigger, detection time, where it routed, who responded, the outcome — pass, or the finding and its fix. The one proof the pager worksQuality Engineer, from the Step 4 drillQErequired; Step 4 runs the drill and the template shipsThe gate packet; Close
the what-healthy tablePer failure scenario and journey: healthy, degraded, who is woken, who is told in the morning. The session's entire outputPod + client operationsOn-call lead + Pod Leadno path — folded into monitoring-config only if a human types itEvery alert definition
the fatigue-review recordEach proposed alert replayed over hypercare history; anything firing weekly without action raised or cut, with the countQuality EngineerQEno path — nothing writes itThe shipped alert set
the outcome-metric first readThe engagement's headline number, read honestly for the first time in production, caveats attachedPod Lead + sponsorSponsorno path — on the business dashboard and spoken at steeringClose — the final scorecard
close-handoff.mdMonitoring inventory, drill record, debt log, open items with owners — the package Phase C opens onClaude drafts — but the phase body never says howPod Leadgate demands the path; no step writes itPhase C, directly
Read the amber rows again

Five of the ten things Phase 9 is supposed to produce have nowhere the tooling puts them — and two of them are the phase's whole reason to exist. The drill that caught Harbor's silent-night routing typo is required by the plugin's own exit gate and produced by no step and no command. The what-healthy session that shapes every alert reaches disk only if someone types it up. Human work is not the problem. Human work without a receipt is.

Deliberately not produced in Phase 9: new features, a parallel monitoring stack of the pod's own (we wire into Harbor's), an alert-tooling migration, and the formal harness handover — that is Close's job. The backlog stays closed; hypercare defects ride the loop, as ever.

04

An exhibit from alert-definitions.md — the six that shipped

An alert nobody acts on is noise wearing a badge

Every row derives its threshold from the measured baseline or says on its face that it's modeled. Every row pages a named recipient and links a playbook step. The measured baseline itself caught that production ran slightly better than the design-phase estimate — replica verification came in at 165ms against Phase 2's ~180ms spike. The two proposals that couldn't survive the fatigue review aren't here; that's the point.

AlertCondition (derivation)SeverityPages
VERIFY-DEGRADEDReplica read failures sustained 5 min outside 02:00-04:30 (suppression window + post-window recovery check)CriticalOn-call rotation
SYNC-MISSEDReplica staleness > 24 h (nightly sync failed)CriticalOn-call + Priti
QUEUE-DEPTHClaim queue > 400 sustained 15 min (10x measured normal of ~40; modeled vs Q-18 surge curve — revisit at first CAT event)Warning → critical at 1,200On-call rotation
INTAKE-ERROR-RATE> 2% sustained 30 min (matches the Phase 8 fallback trigger — one number, two documents)CriticalOn-call + Dee notified
ACK-RETRY-BURSTDispatch retries > 3x baseline per 10 min: warning; sustained 60 min: critical (hypercare day-one finding)Warning → criticalOn-call; vendor escalation path
EVAL-GATE-DRIFTEmail extraction accuracy < 95% on the golden set (the ADR-004 gate, now watched in production)WarningQuality owner (Harbor)

Cut at the fatigue review: a per-instance CPU alert (not actionable under autoscale; capacity already on Tom's dashboard) and a portal-latency warning (it duplicated INTAKE-ERROR-RATE's signal). Six actionable alerts beat eight the on-call trains itself to ignore. No number here "felt right" — every one can explain itself, and the one set from modeled data says so on its face, with a named revisit.

05

The promises from design week, cashed — reading nfr-proving-plan.md against production

Every number the design promised now has a place it's actually read

Phase 2's proving plan named, per quality target, the method and the exact place its number would be read. This is the column that came due. Some numbers are now live on a dashboard; one is still modeled and honestly flagged; one became a standing production watch.

NFR / gatePhase 2 said it would be read……and in Phase 9 it is
NFR-01 (ingestion p95 < 5s)The monitoring dashboardLive on the application dashboard — portal submit p95 410 ms, comfortably under the 5s target
NFR-02 (10x surge)The load-test report; hardening pass 1No real storm yet — QUEUE-DEPTH ships modeled from Q-18, revisit at the first CAT event (Priti)
NFR-03 (99.5% business hours)The ops dashboard, monthlyUptime monitor on the intake endpoints, on Tom's system-health dashboard, read monthly
NFR-05 (PII boundary)PR security-gate recordsCarried by Dan's security alert lane — what pages security directly, bypassing the general queue
NFR-07 (audit completeness)Test results; quarterly compliance sampleThe event-log completeness suite stays in CI; the first compliance sample is on the quarter-read calendar
ADR-004 eval gate (≥95% accuracy)The eval suite in CI; regression blocks changesExtended into production as EVAL-GATE-DRIFT — the CI gate now has a live counterpart watching real extractions

This is the through-line of the whole standard in one table: a number named in design week, with the place it would be read written down, is either read here or flagged modeled with a date — never quietly dropped. The one still modeled says so on its face and carries a named owner. That is the difference between a proving plan and a wish list.

06

An exhibit — the artifact the plugin doesn't require but the phase can't do without

An alert that has never fired is a wish

The drill record is the receipt the registry marks optional and the gate's teeth depend on. Nadia kept it by hand. Every critical alert, its synthetic trigger, when it fired, where it routed, who answered, and the outcome — pass, or the finding and its fix. It went into the gate packet.

Drill record — Wed 8/5, observed N. Brooks; attached to the gate packet
Wed 8/5 09:12 VERIFY-DEGRADED synthetic replica block (outside window) FAIL — alert fired but routed to #harbor-ops channel, not the pager rotation (routing-key typo). On a real night: nobody woken. Fixed (routing key corrected via reviewed PR), re-fired 11:05: paged in 38 s, responder followed playbook to RUNBOOK scenario 1, closed clean. 13:20 ACK-RETRY-BURST synthetic retry storm, test lane PASS — warning at 10 min as designed; escalated to critical on sustained simulation; vendor escalation contact confirmed current. 14:40 QUEUE-DEPTH flagged test messages past threshold PASS — paged; responder's playbook step 2 opened a dashboard link that required pod permissions. Link replaced with Harbor-scoped dashboard. Re-drilled clean. 15:55 SYNC-MISSED staleness clock advanced via test toggle PASS — paged on-call + Priti; diagnosis path correct. All critical alerts fired and answered by Harbor's on-call from the playbook. Two findings, both fixed and re-drilled same day.

The playbook the on-call answered from was Harbor's own — every escalation name theirs, the pod appearing once as "escalation of last resort until Close." Both drill fixes went through the full loop — specs, the grader, a non-author Checker — in a monitoring week. The loop is simply how changes happen now, which is exactly what Harbor inherits.

07

An exhibit from project-retrospective.md — the honest cumulative look

Argued from receipts, not vibes

The retro argued from evidence, not adjectives — not "the grader was valuable" but spec 0016, the bug eleven green tests hid, dead on a branch; not "security review was slow" but 2.1 days against 0.9. Accepted-as-is ended the engagement at 84% and rising. Three concrete outputs, priorities and timing attached, feed straight into Close.

OutputItems
Technical debt logNo un-merge path for a wrong claim merge (accepted in spec 0016 by design; revisit Q4 2026) · legacy intake fallback decommission at day 30 (owner: Tom) · modeled surge thresholds (revisit at first CAT event, Priti)
Client-facing improvementsA twice-weekly committed security-review slot in Dan's calendar, written into the Close handoff — the structural fix for the 2.1-day drift
The harvest list (4 items for the kit PR)Config versioned with the release artifact (spec 0046 → kit pipeline starters) · the timezone-boundary test pattern (Build's escaped-bug answer → kit test-writer skill) · the suppression-window-plus-recovery-check alert pattern (this week's replica decision → kit alert template) · the vendor-blip warning/sustained-critical split. Phase C opens the PR.

The harvest list closes the loop on the standard itself: the suppression-window pattern that came from Harbor's what-healthy session becomes a kit alert template the next engagement inherits. A lesson learned once, paid forward.

08

The handoff — and the open items that travel with it under their original IDs

What Phase C receives

A phase ends by handing the next one a package, not a feeling. Everything below crosses the boundary into Close & Transfer: the watched system now formally Harbor's, the drilled response, the honest retrospective, and the questions still open — carried forward under their original IDs, never silently dropped.

Crosses into Phase C monitoring-config.md alert-definitions.md incident-response.md project-retrospective.md drill-record.md close-handoff.md — required, unspecified the harvest list → the Phase C PR

The Close & Transfer handoff (summary)

Drafted day 10 by Claude for the Pod Lead to own — against no template the plugin provides.

SectionContents
Monitoring inventoryThree dashboards (system / application / business), each with a named Harbor owner; six alerts, all drilled; the playbook cross-referenced to the RUNBOOK
The watchFormally Harbor's as of 8/7; the pod one escalation away until Close
Debt logUn-merge path (revisit Q4 2026) · legacy fallback decommission (Tom, day 30) · modeled surge thresholds (Priti, first CAT event)
Open itemsThe security-review slot change (Dan's calendar, structural fix from the retro) — to be observed working during Close
Harvest listFour items for the kit PR (config-with-artifact, timezone test pattern, suppression-window alert pattern, vendor-blip severity split) — Phase C opens it

The questions that travel with it

IDOpen questionOwnerDue
Q-18Surge thresholds are modeled from the 2024 CAT dataset — revalidate against the first real catastrophe eventPriti ShahFirst real CAT event
The overall (all-claims) median can't be read fairly until complex claims clear — the first honest readingKaren Voss + Pod LeadEnd of first full quarter
Legacy intake fallback still warm — decommission once production is proven stableTom ReillyDay 30 post-go-live

The engagement's credibility was banked, not spent: 1.9 days on fast-path stated with the overall-median caveat welded on. Phase C is where the client proves it can run all of this without the pod — the last thing to transfer is the watch itself.

All names, numbers, and documents are invented but internally consistent — the 11.4-day baseline, the six alerts, the drill record, and the Phase 2 proving-plan promises trace through every artifact on this page.

The system is already live. You've been handed it and told to run Phase 9. This page is what you actually type, in order, and what you do between the typing.

Phase 9 takes about two weeks and runs inside the hypercare window on purpose — week one exists so real production traffic can pile up, because you cannot write a single honest alert threshold until you have measured what normal looks like. Read this end to end before you start: the number that unlocks the whole phase takes days to accumulate, not minutes.

Before you type anything

What you need first

Six things. Three of them are people, one of them is time, and none of them can be arranged the morning you need them.

  • The system live in production, and Phase 8 closed. That's the entry condition. No production traffic means no baseline, and no baseline means every threshold you write is a guess.
  • The plugin, installed. If you arrived here through Phase 8 it already is. If /sdlc-status doesn't run, install it: /plugin marketplace add MCKRUZ/claude-code-sdlc then /plugin install claude-code-sdlc@mckruz. You also need uv (pip install uv) — the plugin runs its Python checks through it.
  • phase9-handoff.md, read. Phase 8 wrote it. The plugin's first step in this phase is to read it back to you and ask questions about it.
  • Read and write access to the client's monitoring stack. Whatever their team already watches — that's the stack. You are wiring into theirs, never standing up your own. Alerts built in a tool the client doesn't open become noise the week you leave.
  • The client's operations / on-call people, booked twice. Once in week one for the session that decides what "healthy" means, once on day 8 for the drill. They are the thread of the phase; they're the ones being paged.
  • RUNBOOK.md from Phase 7, open. Its failure scenarios are the list your alerts get derived from. You don't brainstorm alerts; you walk that list.
The mistake new people make

Writing alerts on day one. It feels productive and it produces numbers nobody measured. Week one is for measuring normal and getting operations in a room. If you skip straight to thresholds you will spend week two arguing about numbers that can't explain themselves.

01

Type this — first thing in the phase

Orient, and agree the scope

You type /sdlc

What happens: it prints the Phase 9 guidance from the plugin's phase definition — what to do next, the required artifacts with an exists/missing mark against each, and the exit criteria. The phase's own first step is then a stop: before any monitoring gets configured, Claude reads phase9-handoff.md and puts questions in front of you — what are the top three things that could go wrong in production, who gets paged and at what threshold, and what monitoring infrastructure already exists (are we adding to their Grafana / Datadog / CloudWatch, or starting from scratch).

What you do: answer the third question from the client's reality, not your preferences. And do not answer the second one yourself — "who gets paged" is operations' answer, in their names. If you haven't asked them yet, that's the next thing you do.

Check the project type before anything else

Phase 9 reads project_type from state.yaml and changes shape completely. For service / app it's the full thing: dashboards, alerting, on-call, runbook. For library / cli it's package health — download counts, open issues, version adoption — and "alerts" means issue-triage criteria. For skill there is no server and no metrics pipeline at all: monitoring is GitHub Issues plus user feedback. Configuring dashboards for a skill project is a wasted phase.

Don't move on until: the monitoring scope is written down and names a stack that already exists at the client, with a named owner for every dashboard you intend to build.

02

The step everything else rests on

Measure what normal looks like

You type nothing new — this happens inside the /sdlc session

What happens: the phase tells Claude to spawn an agent called performance-benchmarker to establish the production baseline — response times, throughput, error rates and resource usage, measured against the targets in non-functional-requirements.md. Those measurements are the "normal" that every alert threshold later gets derived from.

The agent this step names does not ship with the plugin

performance-benchmarker appears in the plugin's agent roster and in Phase 9's own step text, but there is no agent definition for it anywhere in the plugin. The same is true of doc-updater (step 4) and feedback-synthesizer (step 10). If your install doesn't have them from somewhere else, the spawn simply won't find them. That does not excuse you from the step — read the numbers out of the client's own stack yourself. The measurement is the point; the agent was only ever a convenience.

What you do: pull the real numbers over the hypercare window that has actually elapsed — request rate, latency percentiles, error rate, queue depth, dependency health — and write each one down with the period it was measured over. A number without its window is not a baseline; it's an anecdote.

When production hasn't exercised a path: the seasonal surge, the rare dependency failure. Do not invent a normal for it. Use what the engagement already measured — the hardening passes' load-test results, the design phase's spikes — mark that value modeled on its face, and give it a revisit date and a named owner. Honest and modeled beats confident and wrong.

Where this lands baseline-data.md — optional; the gate never asks for it baseline measurements also belong inside monitoring-config.md

Write it down even though nothing forces you to. Every threshold in step 5 has to cite this data, and a reviewer who can't see the baseline can't review a single alert.

03

Nothing to type — get operations in a room

Decide what "healthy" means, together

Half a day with the client's operations people. Walk every failure scenario in RUNBOOK.md and every top-priority user journey, and answer the same four questions for each one: what does healthy look like, what does degraded look like, who gets woken, and who is merely told in the morning.

What you do: ask, capture, and let operations win ties. You know what the system does; they know what 3 a.m. is like in this company. One line per scenario and per journey is enough. That table is what the alert definitions get written from — the whole of step 5 is just this session, formalised.

This session's output has no required home

what-healthy-table.md is an optional artifact. The gate never asks for it. Phase 9's exit conditions don't mention it either, so — unlike the optional receipts in the final phase — nothing puts the question in front of the person who signs. If nobody types the table up, the entire design session survives only in somebody's memory. Fold it into monitoring-config.md or write the file. Waking the wrong person at 3 a.m. is a design defect, and this is where that design is recorded.

04

Back at the keyboard — the first required file

Write the monitoring configuration

You type nothing new — you draft this with Claude in the /sdlc session

What happens: the phase hands the baseline output to a doc-updater agent (see the warning in step 2 — it may not exist on your install) to write monitoring-config.md. The file is required and it must contain four things: the dashboard inventory (what exists and what each one shows), the metrics catalog (every metric collected, its source, its meaning), the coverage assessment (is every top-priority feature observable, and what's the gap), and the baseline measurements from step 2.

What you do: two things nobody will make you do. First, put a named owner on every dashboard — a dashboard nobody reviews is decoration, and it doesn't ship. Second, build the business layer, not just system metrics: the one number the client hired you to move, expressed in their own reporting language, co-built with their data lead. A client who can see request rates but not whether the thing they bought is working has been given a monitoring stack and no answer.

You now have monitoring-config.md — required; the gate blocks without it
05

The spine of the whole phase

Derive the alerts from the baseline

You type nothing new — you draft this with Claude in the /sdlc session

What happens: you and Claude write alert-definitions.md. It's required, and the specification is exact. An alert table — name, condition, severity, recipient, response time, link to its playbook entry. Per-alert detail for every critical alert — the exact query or threshold, why this threshold, what to do when it fires, how to resolve it. An alert philosophy — how you decided what deserves an alert versus what you merely watch. And a baseline reference — how each threshold was derived from step 2's numbers.

What you do: walk the RUNBOOK's critical failure modes and give each one an alert. Give every alert two thresholds: warning means investigate during working hours, critical means wake someone up. The difference between them is simply who suffers if it waits until morning. Then write the derivation next to the number, every time. “2x the measured p95 over 10 minutes” survives an argument six months from now. “500ms” does not — and a threshold that can't explain itself can never be tuned, only deleted.

Why a made-up number is worse than no alert

A guessed threshold fails in one of two ways. Too low: it fires constantly, the team mutes it, and now you have something worse than no alert — a team trained to ignore the pager, which is the state they'll be in when the real one fires. Too high: it never fires at all, and you discover that during the incident. Neither failure is visible in any gate, any report, or any file listing. The only thing standing between you and both of them is having actually measured normal first.

Get the confirmation recorded, not remembered: the people being paged confirm every threshold, and their names go on the review of the change that ships the alert rules. A meeting where everyone nodded is not a record.

You now have alert-definitions.md — required; the gate blocks without it
06

Nothing to type — and nothing checks it

Run the fatigue review before anything ships

Take every alert you just proposed and replay its condition against the hypercare metrics history — the same data your baseline came from. For each one, count two numbers: how many times it would have fired over that period, and how many of those firings would have demanded that somebody actually do something.

What you do: anything that would have fired more than once a week without demanding action gets its threshold raised or gets cut. Now — not in three months, after the team has already learned to ignore it. Record the decision per alert: kept, raised, or cut, and why.

Why this is a separate pass: every alert looks reasonable in isolation. The damage is cumulative. Fifty alerts feels thorough and guarantees all fifty get ignored within a month. Fewer, actionable, derived, reviewed.

Where this lands fatigue-review-record.md — optional; nothing asks whether you did this
07

Only if any part of what you built is LLM-powered

The three alerts the plugin never mentions

If a spec's deliverable was an agent, a prompt, or anything else where the behaviour is probabilistic, the standard requires three extra alerts here: cost spikes (token cost per task), refusal rates, and eval drift in production samples — plus tracing (what the agent saw, decided, and called) and failure-mode logging.

What you do: add all three to alert-definitions.md like any other alert — condition, severity, named recipient, playbook link. The eval-drift one already has its number: /sdlc-evals authored a versioned golden set and a pass threshold for that spec during the build. That threshold is the alert condition. Don't invent a second one; extend the CI gate into a live watch on real production samples.

Nothing in the tooling will remind you

This requirement lives in the delivery standard, not in the plugin. Phase 9's definition file never mentions agents, tokens, refusals, or evals. No artifact field asks for them, no gate check looks for them, and no command produces them. If your system has an LLM-powered feature and you forget this step, every automated check in the phase will still come back green.

08

Back at the keyboard — the third required file

Write the incident playbook

You type nothing new — you draft this with Claude in the /sdlc session

What happens: incident-response.md gets written against the alert table. Required, and it must contain: incident classification (P1/P2/P3, with definitions), response procedures per alert type — detect, diagnose, resolve, communicate — an escalation matrix naming who to contact at each severity, communication templates for what you tell users and stakeholders while it's happening, and the post-incident process for writing the post-mortem.

What you do: cross-reference RUNBOOK.md instead of duplicating it. Same failure modes, different lens — the RUNBOOK resolves; the playbook detects, classifies, and communicates. And make every escalation entry a person's name. A team alias nobody owns is how an escalation dies at 3 a.m.

You now have incident-response.md — required; the gate blocks without it
09

Nothing to type — and the plugin gives you no step for it

Fire every critical alert on purpose

Before drill day, write down one synthetic trigger per critical alert — a test lane, flagged test data, a test toggle, a deliberately blocked dependency — something that makes the real alert condition true without touching real traffic or real data. No trigger agreed in writing, no drill.

What you do: fire each critical alert and let the client's own on-call respond from the playbook while you watch, silently. Record, per alert: the trigger used, the time it fired, the time it was detected, where it routed, who responded, and the outcome — pass, or the finding and its fix. Whatever the drill breaks gets fixed through the normal loop and re-drilled.

Why it's non-negotiable: routing that goes to the wrong channel, a playbook step that quietly assumes pod access, a threshold that doesn't actually trigger — all of it fails here, by appointment, at drill prices. An alert that has never fired is a wish.

The receipt, and where its shape comes from

drill-record.md is on the plugin's required-artifact list for Phase 9 — the gate's integrity and completeness checks block if it's missing or empty. Step 4 runs the drill and specifies the file, and a template ships at templates/phases/09-monitoring/drill-record.md, so you are filling in a shape rather than inventing one. The registry also carries "alert drill executed" as an exit condition, put in front of the human who signs. It goes in the gate packet.

You now have drill-record.md — required, with a step and a template behind it
10

Half a day — the pod and the client's core team

Write the retrospective honestly

You type nothing new — Claude assembles the evidence, the humans supply the candour

What happens: the phase spawns feedback-synthesizer in the background to look through whatever user feedback the launch produced — issues, support requests, survey results — for patterns. (Same caveat as step 2: that agent isn't shipped with the plugin. If it isn't available, read the feedback yourself.) Meanwhile you write project-retrospective.md, which is required and has two halves that both have to be there.

The product half: which technical decisions aged well, which created debt, what collaboration patterns worked.

The process half (this one is explicitly required, and it's the one people skip): which phases were worth their cost and which felt like overhead; which human sign-off gates caught real issues and which got rubber-stamped; which artifacts were actually referenced later and which were written and never read; what you'd skip next time and what you'd add; whether the profile's thresholds were right.

What you do: produce three concrete outputs. The technical debt log — every known shortcut, with a priority and a suggested timing. Logged debt is managed; unlogged debt is next year's crisis with no paper trail. The client-facing improvements — what their team changes about how they run this. And the harvest list — the patterns, corrections and improvements the next engagement should inherit.

No gate can tell candid from polite

"Everything went great, communicate better next time" passes every automated check in the phase — the file exists, it isn't empty, it has no leftover TODO in it. Write items concrete enough to act on: not "communicate better" but "add a daily async standup during the build loop". And keep it for the teams, not for management — the retro records findings, not names attached to blame.

You now have project-retrospective.md — required; the gate blocks without it
11

Nothing to type — the number the engagement was bought for

Read the outcome metric out loud, with its caveats

The one number fixed back in Phase 0 gets its first honest production read here. Sit with the sponsor, read it off the live system, and state every caveat alongside it: which cohort it covers, whether the period is partial, and what else changed at the same time.

Why the caveats are the artifact: a first read published without them gets quoted for a year without them. Record the number, the window, where it was read from, every caveat, and who was in the room.

Also today: hypercare formally ends. The dashboards, the pager and the playbook become the client's, with the pod one escalation away until the final phase. Make the handover a deliberate moment with a date on it — a hypercare that quietly never ends is a handover failing in slow motion.

Where this lands outcome-metric-first-read.md — optional; nothing asks for it
12

The last required file — and it's all names

Write the handoff to the final phase

You type nothing new — you draft this with Claude in the /sdlc session

What happens: close-handoff.md is required, and the next phase opens by reading it. It must state: transfer readiness plainly, including if the answer is no; the named client engineers who will run real work at the close gate, and their availability; the named client Setup Owner who takes the harness, and whether they have already merged a harness change themselves; candidate backlog items — real work, not toy work invented for the test; the access inventory to revoke (every seat, token, repo permission, environment role and vault policy the pod holds); the operational state at handover — what's live, what's monitored, what's still manual, and any alert or incident open right now; and the known debt and open risks carried from the retrospective, with owners on the client's side.

What you do: the phase stops and asks you to confirm every name individually. Do that literally — each one is a real person who has agreed, not a plausible name from an org chart. "Will be identified later" is not a name. If you can't fill one in, record the gap explicitly. The next phase needs to know what's missing far more than it needs a tidy-looking list.

You now have close-handoff.md — required; the gate blocks without it
13

Type this — the machine checks your work

Run the gate

You type /sdlc-gate

What happens: it runs the seven gate checks, records the results in state.yaml, then writes and opens an HTML report in your browser. The blocking part is narrow and mechanical: it confirms the six required files exist, aren't empty, and contain no leftover placeholder text (TODO, TBD, PLACEHOLDER, [INSERT). The last gate renders the phase's declared exit conditions — including "every alert threshold derived from a measured baseline" and "alert drill executed" — as review lines for the human who signs. Those never block. Gates report; humans decide.

What you do: fix whatever it flags and run it again until it's clean. A stray TODO in one file is the usual culprit.

Optional: /sdlc-phase-report regenerates the HTML at any time, and /sdlc-enhance writes plain-language companions to the artifacts — worth doing if your sponsor isn't technical.

What the gate does NOT check

It can confirm alert-definitions.md exists and has no placeholder text in it. It cannot read a threshold and tell you whether anybody measured anything — a file full of numbers pulled from thin air passes exactly as cleanly as one derived from real traffic. It cannot tell whether the drill actually happened or whether drill-record.md was written from memory the morning of the gate — Step 4 tells you to write it as you go, and nothing can enforce that. It cannot see operations' names on the change review. And the fatigue review, the what-healthy table and the outcome metric's first read have no check and no sign-off question behind them at all. A green gate is not a finished phase.

On the escape hatch: a line reading WAIVED: <name> — <reason> inside a required artifact lets it pass the completeness check, printed with the name attached in the report the approver signs. It exists so a receipt that genuinely doesn't apply doesn't get worked around invisibly. It is not a way past a step you skipped.

14

Type this — last thing in the phase

Take the sign-off and advance

You type /sdlc-next

Before you type it: run the steering session with the sponsor. Walk the drill record, the debt log, the watch handover, and the outcome metric's first read with its caveats stated. This is a billing milestone and it's where the engagement's credibility gets spent or banked — it is not ceremonial.

What happens: it re-runs the gates, then stops and asks you to confirm before moving anything. Say yes and it advances the project to the final phase. Then it surfaces every open question from close-handoff.md and makes you answer or explicitly defer each one before any work in the next phase starts.

What you do: answer them properly. An open question waved through at a handover is a question the client rediscovers after you've gone.

Phase 9 is done when state.yaml says Close & Transfer every critical alert has fired once, on purpose, and was answered the dashboards, pager and playbook are formally the client's

Keep these handy

Commands you'll use constantly

Type thisWhen
/sdlc-statusAny time you're lost. Shows what phase you're in and what's missing.
/sdlcStart of every work session. Tells you what to do next.
/sdlc-gateAny time you want to know how far off you are. It records results but never advances the phase.
/sdlc-coachYou're stuck and want to be walked through it conversationally instead of following a list.

Rule of thumb for the whole phase: a threshold that can't explain itself can't be tuned. If you can't say which measurement a number came from and over what window, you haven't set an alert — you've set a coin flip that someone will eventually mute.

01

Reference · Phase 9 · Monitoring

Live is not the same as watched

One job, two halves: make the live system observable and the response proven, and write down what the engagement learned. Phase 9 runs deliberately inside the hypercare window — week one needs the production baseline to accumulate; week two needs the pod present for the drill. The gate falls at hypercare's end, closing both together.

The precise mechanics — the exact roles, calendar, artifacts, cadences, and gate, including the tooling specifics the How-it-works view leaves out. The full prose method sits under each section's “Go deeper” on the How it works tab; for the complete Harbor artifacts (every ID and command), see Example.

02

The rule that governs everything

Decide what "healthy" means — together

Human drivesClaude doesMandatory human stops
Operations co-author what healthy means and confirm every threshold against real baseline data; the Pod Lead owns the retrospective's candor; the client's on-call responds to the drill.Establishes the baseline from production data; drafts alert definitions with each threshold derived from the baseline; drafts the incident playbook from the RUNBOOK; assembles the retrospective's evidence base.Confirming a threshold (the named stop — every threshold confirmed against real baseline data by the people being paged); declaring the drill passed; softening the retrospective. Phase advance.
03

What the two weeks must answer

Four questions — and deliberately nothing else

Phase 9 answers four questions, and nothing else. New features, alert-tooling migrations, and the formal handover of the harness are out of scope.

  1. Can the client see the system? (dashboards with owners, every top-priority feature observable, business metrics alongside system metrics)
  2. Will the right person find out, at the right urgency? (alerts derived from measured baselines, routed to named people, covering every critical RUNBOOK failure mode)
  3. Does the response actually work? (the incident playbook, proven by drill — detected, diagnosed, communicated by the client's own on-call)
  4. What did the engagement learn? (the honest retrospective: product, process, debt, and the patterns the kit inherits)

Phase 9 runs deliberately inside the hypercare window: week one needs the production baseline to accumulate; week two needs the pod present for the drill. The gate falls at hypercare's end, closing both together.

Who is involved — our side

PersonLoadWorkstream
Pod Lead60–80%Runs the what-healthy-means session, owns the retrospective, routes threshold decisions to named owners, runs the gate and steering
Quality Engineer60–80%Observability coverage, the alert-fatigue review, and the drill design
Setup Owner50–70%Wires dashboards and alert rules into the client's stack; owns the baseline capture; versions every monitoring change like a harness change
Orchestrators20–40%Drive the drafting (alert definitions, playbook, dashboard config); fix what hypercare and the drill surface — through the loop

Client side

PersonNeeded forHow much
Operations / on-callCo-author what healthy means; receive the routing; respond to the drill from the playbook — they own all of this in a few weeksThe thread of the phase
Platform engineerWires alert channels and paging into their tooling; confirms dashboards live where their team already looks3–4 hours
Data / reporting leadThe business-metric dashboards: the outcome metric and its caveats, in the client's reporting language2–3 hours
SecurityThe security alert lane: what pages security directly, bypassing the general queue1–2 hours
Product OwnerConfirms the business metrics measure what the business means by them~1 hour
SponsorThe gate steering: the first honest production read of the outcome metric, caveats included45 min

If operations cannot co-author the thresholds, the phase produces our alerts in their tooling — which become noise the week we leave. The session is the phase; protect it like Phase 7 protected the cold runs.

04

The number behind every alarm

Thresholds come from measured reality

Before any alert is written, the baseline is captured — what normal looks like for each key metric, measured from real production traffic, each number recorded with its measurement period. Every threshold is derived from it, the derivation written next to the number. Where production hasn't exercised a path, the threshold derives from modeled data — flagged as modeled, with a revisit date and a named owner, never silently presented as baseline.

FieldMeans
Measured normalThe number read from real production traffic, with the period it was measured over
ThresholdA stated multiple of normal ("2x p95 over 10 min"), with the derivation written next to it
Modeled flagFor unexercised paths: derived from the engagement's modeled data, flagged, with a revisit date and named owner
05

Why fewer alerts are better

An alert nobody acts on is noise wearing a badge

Every alert is a condition that demands a human response. The alert-fatigue review replays each proposed condition against the hypercare metrics history and counts the firings; anything that would fire more than once a week without demanding action gets its threshold raised or gets cut.

TermMeans
WarningInvestigate during working hours
CriticalWake someone up — the difference is who suffers if it waits until morning
Alert fatigueWhat happens when alerts fire often and mean little: the team learns to ignore them, and then misses the real one
The standing ruleFires more than once a week without action → raise the threshold or delete the alert
06

What to do when it fires

The incident playbook

The detect-diagnose-communicate companion to the RUNBOOK: per alert, what it means, the first three diagnosis steps, severity classification, escalation names, and the communication templates — what users and stakeholders are told, by whom, while it's happening. It cross-references the RUNBOOK rather than duplicating it: the playbook detects and communicates; the RUNBOOK resolves. One source of truth per failure, two lenses on it.

07

Proving it actually works

The alert drill

Deliberately triggering each critical alert in a controlled way and having the client's on-call respond from the playbook. An alert that has never fired is a wish — the same rule the rollback rehearsal enforced in Phase 8. Controlled means designed: before drill day, one synthetic trigger per critical alert is written down — a test lane, flagged test data, a test toggle, a deliberately blocked dependency — that makes the real condition true without touching real traffic. No trigger agreed in writing, no drill. The Quality Engineer records, per alert: the trigger used, trigger time, detection time, where it routed, who responded, and the outcome; the record goes in the gate packet. What breaks gets fixed through the loop and re-drilled.

08

What the whole engagement learned

The engagement retrospective

A half-day, pod plus client core, arguing from assembled evidence — the metrics history, the Retro+ log, the escaped-bug answers, the gate records. Claude assembles the evidence base; the humans own the candor (an honest retro written by a flattering drafter is worthless). Three concrete outputs:

OutputWhat it is
Technical debt logEvery known shortcut, with a priority and a suggested timing — logged debt is managed, unlogged debt is a future crisis
Client-facing improvementsWhat the client's team changes about how they run the system and the loop
The harvest listThe patterns, skills, hook improvements, and template corrections the next engagement inherits through the kit; Phase C opens the harvest PR
09

The two weeks, end to end

The two-week calendar

DaysFocusWhat happensTooling
1–2What does healthy mean?Pod + operations walk the RUNBOOK failure scenarios and top journeys; healthy/degraded/who-is-woken captured per scenario; monitoring scope (stack, dashboards, owners) lands in writing— (human session)
3–4Baseline & dashboardsProduction baseline captured per metric with its measurement period; unexercised paths flagged modeled with a revisit date; dashboards live in the client's stack — system, application, business/sdlc → performance-benchmarker, doc-updater
5Alert definitionsEvery critical failure mode gets a baseline-derived alert with severity, recipient, response expectation, playbook link; operations confirm every threshold; alert-fatigue review runs/sdlc-coach
6–7Incident playbookPer alert: meaning, first three diagnosis steps, severity classification, escalation names, communication templates; cross-references the RUNBOOK/sdlc-coach
8The alert drillEach critical alert fired via a pre-agreed synthetic trigger; client's on-call responds from the playbook; pod observes silently; failures fixed through the loop and re-drilled; drill record captured— (drill)
9The retrospectiveHalf-day, pod + client core; arguing from assembled evidence; three outputs: technical debt log, client-facing improvements, harvest list/sdlc → feedback-synthesizer, /visual-explainer
10Gate & handoverHypercare closes; watch formally handed over; gate check; Close & Transfer handoff drafted; steering — sign-off, billing milestone, the outcome metric's first honest read/sdlc-gate, /sdlc-phase-report, /sdlc-next

When the two weeks stretch: a too-thin baseline ships thresholds as modeled-with-revisit-date rather than waiting for a storm; a badly failed drill is fixed, re-drilled, and — if the gap is people, not config — flagged at steering; a political retrospective is held to "for the teams, not for performance reviews"; a defect-heavy hypercare is a quality finding for the retro, not a reason to extend forever.

Cadences

RhythmWhoWhat
Daily 15-min pod syncWhole podHypercare findings, baseline capture state, drill prep
Hypercare watchClient operators driving, pod besideContinues throughout; its findings feed thresholds and the debt log
What-healthy-means sessionPod + client operationsThe phase's defining event, week one: healthy, degraded, and who gets woken, per failure mode
The alert drillClient on-call responding, QE observingWeek two: every critical alert fired and answered from the playbook
The retrospectivePod + client core teamWeek two: the honest cumulative look, from evidence — produces the debt log and the harvest list
SteeringSponsor + Pod LeadFalls at the gate: the watch handover, the drill record, the outcome metric's first honest read

The artifacts

ArtifactDrafted byOwned byDone means
Monitoring configurationClaude (drafts), Setup Owner (wires)Setup Owner → clientDashboards live in the client's stack, each with a named owner; every top-priority feature observable
Production baselineClaude (measures)Quality EngineerNormal recorded per key metric with its measurement period; modeled values flagged with revisit dates
Alert definitionsClaude (drafts), ops (confirm)Client operationsEvery critical failure mode covered; every threshold derived from baseline and confirmed by the people being paged
Incident playbookClaude (drafts), ops (correct)Client operationsDetect-diagnose-escalate-communicate per alert; templates included; cross-referenced to the RUNBOOK
The drill recordQuality EngineerQuality EngineerEvery critical alert fired and answered by the client's on-call from the playbook; failures fixed and re-drilled
Engagement retrospectiveClaude (evidence base), humans (the candor)Pod LeadProduct and process findings with receipts; debt log; harvest list — concrete items, not platitudes
Close & Transfer handoffClaude (drafts)Pod LeadMonitoring inventory, drill record, debt log, open items with owners

Deliberately not produced: a parallel monitoring stack of our own (we wire into theirs), alerts for things nobody would act on, and a sanitized retrospective for external consumption — the steering gets the summary; the teams keep the honest version.

10

What "done" actually requires

The exit gate

Phase 9 closes — and hypercare ends with it — when all of these are true:

  • Every top-priority feature has at least one observable metric on a dashboard with a named owner
  • Every critical RUNBOOK failure mode has an alert; every threshold is derived from the measured baseline (or explicitly flagged as modeled, with a revisit date) and confirmed by the client's operations (teeth)
  • The alert-fatigue review ran — nothing ships that would fire weekly without demanding action
  • The incident playbook covers every critical alert — detection, diagnosis, escalation names, communication templates
  • The drill happened — every critical alert fired in a controlled way and was answered by the client's own on-call from the playbook; what failed was fixed and re-drilled (teeth)
  • The watch was handed over — dashboards, paging, and the playbook are formally the client's, with the pod one escalation away until Close
  • The retrospective exists with receipts — findings, the technical debt log, and the harvest list; concrete, honest, owned
  • The outcome metric has its first production read on the scorecard, caveats stated
  • The Close & Transfer handoff exists — monitoring inventory, drill record, debt log, open items with owners
  • A named human on each side approved the advance — gates report, humans decide
11

The failure modes to watch

What goes wrong

  • Thresholds by intuition. "500ms feels right" is not engineering. Measure the baseline, alert at a stated multiple, write the derivation down. A threshold that can't explain itself can't be tuned later.
  • Our monitoring, their pager. The pod builds alerts in a stack the client's team doesn't watch, routed by assumptions. The week we leave, it's noise. Their stack, their names, their session.
  • Alert fatigue shipped on day one. Fifty alerts feels thorough and guarantees the team ignores all of them within a month. Fewer, actionable, derived, reviewed.
  • The undrilled pager. Routing and playbooks that have never fired, discovered broken during the first real incident. The drill is to alerts what the rehearsal was to rollback.
  • The dashboard nobody owns. A beautiful screen no one reviews is decoration. Every dashboard has a named owner or it doesn't ship.
  • Monitoring the system but not the business. All system metrics, no outcome metric — the client can see requests but not whether the thing they bought is working. The business layer is co-built with their data lead, in their language.
  • The platitude retrospective. "Great teamwork, communicate better next time" — written for management, useless to everyone. Honest, specific, evidence-backed, or skip the meeting.
  • Debt left unlogged. The shortcuts everyone knows about but nobody wrote down become next year's crisis. Logged debt is managed debt; the log is a gate item for a reason.
  • Hypercare that never ends. Quietly extending the watch because closing feels risky. Extend deliberately with the sponsor and a new end date, or close on schedule.