The system is live. The hard part isn't keeping it running today — it's making sure that when something breaks at 3 a.m., the right person wakes up, and knows exactly what to do.
Phase 9 is where a running system becomes a watched one — and where the whole engagement finally writes down what it learned.
Every idea below is paired with the real thing — examples come from a fictional but fully worked engagement, Harbor Mutual, a regional insurer.
Why this phase exists at all
Live is not the same as watched
Going live proves the system can run. It says nothing about whether anyone would notice if it quietly stopped. A system nobody is watching is an outage waiting to be discovered by a furious customer instead of an alert.
Phase 9 turns "it's running" into "we'd know within minutes if it weren't." That means dashboards someone actually reads, alerts that wake the right person at the right urgency, and a written plan for what to do when one fires.
And it does this on purpose, during hypercare — the agreed window after go-live when the client's operators are at the controls and the pod is still beside them. That timing isn't an accident: it's the only moment when the system is finally producing the one thing every alert needs — real numbers about what "normal" looks like.
Harbor's claims system went live on a Thursday. By the following Monday it was handling real FNOL intake — but nothing paged anyone. The dashboards didn't exist yet; the alerts were still a wish.
"Live since 7/23. Hypercare running. The first real numbers exist; nothing pages anyone yet."
Two weeks — 7/27 to 8/7 — ran inside that hypercare window, ending the day it closed.
The first time you learn the system is broken is when a customer calls — and by then it's been broken for hours, with no record of how to fix it.
Go deeper — the full method
Going live proved the system can run; it said nothing about whether anyone would know if it quietly stopped. Phase 9 makes production observable, makes the alerts real, and writes down what the engagement learned. It runs inside the hypercare window on purpose — the system is finally producing the one thing alert thresholds must be built from (real baseline data), and the pod is still in the room while the client's operators take the controls.
The phase has a second job that matters as much as the first: the engagement retrospective. The weekly Retro+ answered "which check should have caught it?" all through Build; Phase 9 asks the cumulative version — what worked, what didn't, what the next engagement inherits — and produces the raw material for the harvest that closes the loop on the standard itself.
The first time you learn the system is broken is when a customer calls — and by then it has been broken for hours, with no record of how to fix it. A system nobody watches is an outage waiting to be discovered the most expensive way.
The conversation that shapes everything
Decide what "healthy" means — together
You can't alert on "something's wrong" until you've agreed what "fine" looks like. And that question can't be answered by either side alone: the pod knows what the system does; the client's operators know what 3 a.m. is like in their company.
The defining event of the phase is a session where the pod and the client's operations walk through every way the system can fail and every journey that matters, and answer three things for each: what healthy looks like, what degraded looks like, and who gets woken versus who just gets told in the morning.
The operators' answers win ties — they're the ones being paged. Knowing what must not page is as important as knowing what must: an alarm that cries wolf every night trains people to ignore the one that matters.
The deepest decision of the whole phase came from Harbor's side of the table. Their system checks claims against a copy of the core records that refreshes overnight — and is unavailable from 2 to 4:30 a.m. while it does.
"An alert that fires every night at 02:00 trains the on-call to ignore the one that matters."
"Suppress during the refresh window — but add a check that the system recovered after 04:30. That's the part actually worth waking for."
You'll ship your idea of healthy in their tooling — which becomes noise the week you leave. Protect this session; it is the phase.
Go deeper — the full method
Phase 9 is the Pod Lead's phase jointly with the client's operations — because the central session of the week, "what does healthy mean?", cannot be answered by either side alone. The pod knows what the system does; operations knows what 3 a.m. is like in this company.
The session that shapes the phase: the pod and the client's operations walk the RUNBOOK's failure scenarios and the top-priority user journeys and answer, for each, what healthy looks like, what degraded looks like, and who should be woken for what. Operations' answers win ties — they are the ones being paged. Capture one line per scenario and journey — healthy looks like, degraded looks like, who is woken, who is told in the morning — in the monitoring configuration. That table is what the alert definitions are written from.
The monitoring scope lands in writing: which stack (theirs — we wire into what their team already watches), which dashboards, which alert channels, who owns each dashboard. A dashboard without a named owner is decoration.
The phase produces our alerts in their tooling — which become noise the week we leave. The session is the phase; protect it like Phase 7 protected the cold runs.
What the two weeks must answer
Four questions — and deliberately nothing else
Phase 9 doesn't add features and doesn't migrate the client to fancier tooling. It exists to answer exactly four questions, and everything it produces serves one of them.
- Can the client see the system? Dashboards with owners, every important feature observable, business numbers next to system numbers.
- Will the right person find out, at the right urgency? Alerts built from real numbers, routed to named people.
- Does the response actually work? A written plan, proven by drill — not assumed.
- What did the engagement learn? An honest retrospective: what worked, what didn't, what the next project inherits.
- Three dashboards in Harbor's own tooling — system, application, and the business clock — each with a named owner.
- Six alerts, each routed to Harbor's on-call rotation, each tied to a real failure mode.
- A drill where every critical alert fired for real and Harbor's own on-call answered it.
- A retrospective that sent four patterns home to improve the next project.
It feels helpful to add "just one more dashboard" or swap in a better monitoring tool. Don't. You wire into the stack the client's team already watches — anything else is a thing they won't maintain.
Go deeper — the full method
Phase 9 answers four questions, and nothing else:
- Can the client see the system? Dashboards with owners, every top-priority feature observable, business metrics alongside system metrics.
- Will the right person find out, at the right urgency? Alerts derived from measured baselines, routed to named people, covering every critical failure mode the RUNBOOK describes.
- Does the response actually work? The incident playbook, proven by drill — detected, diagnosed, communicated by the client's own on-call.
- What did the engagement learn? The honest retrospective: product, process, debt, and the patterns the kit inherits.
New features, alert tooling migrations, and the formal handover of the harness are out of scope — the backlog stays closed, the client's existing monitoring stack is the one we wire into, and Close & Transfer is the next phase's job. Defects surfaced by hypercare ride the loop, as ever.
The number behind every alarm
Thresholds come from measured reality — not from a number that "felt right"
"Alert me if response time goes over 500 milliseconds" sounds rigorous. It's a guess. Unless you know what response time actually is on a normal day, that number is arbitrary — and arbitrary thresholds either cry wolf or stay silent through real trouble.
Before any alert is written, you capture the baseline: what normal looks like for each key metric, measured from real production traffic, each number recorded with the period it was measured over. Every threshold is then derived from it — "twice the measured normal, sustained ten minutes" — and the derivation is written down next to the number.
And when production hasn't yet exercised a path — the rare storm, the seasonal surge — you don't guess. You derive from the engagement's earlier modeled data, flag it as modeled, and set a date to revisit. Honest beats tidy.
The team measured the real baseline from the first hypercare week — and caught that production ran slightly better than the old design-phase estimate:
One path had no real numbers: late July gave Harbor no storm. The surge thresholds were derived from a modeled 2024 catastrophe dataset and flagged modeled, revisit at the first real storm — with a named person on the revisit.
Thresholds set by intuition. "500ms feels right" is not engineering. A threshold that can't explain where it came from can't be tuned later — it just gets ignored or deleted.
Go deeper — the full method
The production baseline gets captured from the first hypercare weeks' real traffic: request rates, latency percentiles, error rates, queue depths, dependency health — each recorded with the period it was measured over. Where production hasn't yet exercised a path (the seasonal surge, the rare failure), the threshold derives from the engagement's modeled data instead (whatever the engagement already measured: the hardening passes' load-test results, the design phase's spikes) — flagged as modeled, with a revisit date, never silently presented as baseline. Claude measures, the Setup Owner runs the capture in the client's tooling, and the Quality Engineer owns the resulting artifact — including the honesty of every modeled flag.
Every threshold is then derived from the baseline — alert at a stated multiple of normal, never at a number that felt right — and the derivation is written next to the number. "2x the measured p95 over 10 minutes" survives an argument; "500ms" does not.
"500ms feels right" is not engineering. Measure the baseline, alert at a stated multiple, write the derivation down. A threshold that can't explain itself can't be tuned later — it just gets ignored or deleted.
Why fewer alerts are better
An alert nobody acts on is noise wearing a badge
Fifty alerts feels thorough. It guarantees the team ignores all of them within a month — and then misses the real one. Every alert must demand a human do something; anything else is just decoration that erodes trust in the whole system.
Before any alert ships, it's tested against history: replay each proposed condition over the hypercare data and count how often it would have fired. The standing rule — anything that would fire more than once a week without demanding action gets its threshold raised or gets cut. Now, before the team learns to tune it all out.
This is the discipline of restraint. Six alerts people answer beat eight they've trained themselves to swipe away.
The fatigue review cut two proposed alerts before they shipped:
Portal latency warning — cut. It duplicated the signal the intake error-rate alert already gave.
Six actionable alerts beat eight ignorable ones.
Go deeper — the full method
An alert is a condition that demands a human response. Every alert is actionable; a notification nobody acts on is noise wearing an alert's badge. Alert fatigue is what happens when alerts fire often and mean little: the team learns to ignore them, and then misses the real one.
The alert-fatigue review runs on everything proposed: replay each proposed condition against the hypercare metrics history — the same data the baseline came from — and count the firings. The standing rule: anything that would have fired more than once a week during hypercare without demanding action gets raised or cut now, not after the team has learned to ignore it. Two thresholds per alert make this precise: warning means investigate during working hours; critical means wake someone up. The difference is who suffers if it waits until morning.
Fifty alerts feels thorough and guarantees the team ignores all of them within a month. Fewer, actionable, derived, reviewed — six alerts people answer beat eight they've trained themselves to swipe away.
What to do when it fires
An alert tells you something's wrong; the playbook tells you what to do about it
A pager going off at 3 a.m. is useless if the person it woke doesn't know what the alert means, how to diagnose it, who to call, or what to tell upset customers while they work. The alarm is half the system; the response plan is the other half.
The incident playbook is the detect-diagnose-communicate companion to the existing repair guide. For each alert it spells out what it means, the first diagnosis steps, who escalates to whom, and the templates for what users and leadership hear while it's happening.
It deliberately doesn't repeat the repair steps — it points to the existing guide for those. One source of truth per failure, two lenses on it: the playbook detects and communicates; the runbook resolves.
Harbor's on-call lead corrected the draft playbook line by line. The escalation names were all Harbor's — the pod appears only as a last resort:
"Escalation of last resort until Close."
It spelled out exactly what the claims office hears from the intake team when intake degrades, and what leadership hears on a sev-1 — and who says it.
Go deeper — the full method
The playbook gets written against the alert table: per alert, what it means, the first three diagnosis steps, the severity classification, the escalation names, and the communication templates — what users and stakeholders are told, by whom, while it's happening. It cross-references the RUNBOOK rather than duplicating it: the playbook detects and communicates; the RUNBOOK resolves.
The incident playbook is the detect-diagnose-communicate companion to the RUNBOOK: what each alert means, first diagnosis steps, who escalates to whom, and what to tell users while it's happening. One source of truth per failure, two lenses on it — which is why the playbook never repeats the repair steps; it points into the cold-verified RUNBOOK for resolution.
Proving it actually works
An alert that has never fired is a wish
Routing rules, paging chains, and playbook steps all look correct on paper. The only way to know they work is to make each one happen on purpose — and watch the client's own on-call respond. Discover the broken pieces by appointment, at drill prices, not at 3 a.m. during a real incident.
Every critical alert is triggered deliberately, in a controlled way, and the client's on-call responds from the playbook while the pod watches in silence. "Controlled" means designed: before drill day, you write down one safe synthetic trigger per alert — a test lane, flagged test data, a deliberately blocked dependency — that makes the real condition true without touching real traffic. No agreed trigger, no drill.
What breaks gets fixed and re-drilled until it's clean. This is the same discipline that proved the rollback in the previous phase: the difference between a procedure and a hope.
The drill earned its keep. Two findings, both exactly what a drill is for:
"VERIFY-DEGRADED routed to the general ops channel instead of the pager rotation — a routing-key typo that would have meant a silent night during a real incident."
"Routing key corrected by reviewed PR. Re-fired 11:05: paged in 38 seconds, responder followed the playbook, closed clean."
A second finding — a playbook step that opened a dashboard needing pod permissions — was caught and re-drilled clean the same day.
Go deeper — the full method
Each critical alert fires for real, triggered in a controlled way, and the client's on-call responds from the playbook while the pod observes silently — the same discipline as Phase 7's cold runs and Phase 8's rehearsal. Controlled means designed: before drill day, the Quality Engineer and the client's platform engineer write down one synthetic trigger per critical alert — a test lane, flagged test data, a test toggle, a deliberately blocked dependency — that makes the real alert condition true without touching real traffic or real data. No trigger agreed in writing, no drill.
Routing that goes to the wrong channel, a playbook step that assumes pod access, a threshold that doesn't actually trigger: all of it fails here, by appointment, at drill prices. What the drill breaks gets fixed through the loop and re-drilled. The Quality Engineer records, per alert: the trigger used, trigger time, detection time, where it routed, who responded, and the outcome — pass, or the finding and its fix. That record goes in the gate packet.
An alert that has never fired is a wish. The difference between a procedure and a hope is that someone has run it — deliberately, in advance, at drill prices.
What the whole engagement learned
The honest retrospective — argued from receipts, not vibes
Phase 9's second job matters as much as the first. With the engagement nearly over, this is the moment to ask the cumulative question: what worked, what didn't, and what the next project should inherit. The trap is the feelings-meeting version — "great teamwork, communicate better next time" — written for management and useless to everyone.
The AI assembles the evidence — the metrics history, the weekly logs, every escaped bug with its answer — so the humans argue from facts, not vibes. The AI never softens the retro; a flattering drafter produces a worthless one. The candor is the humans' job.
It produces three concrete outputs: a technical debt log (every known shortcut, with a priority and a date — logged debt is managed, unlogged debt is next year's crisis), client-facing improvements, and the harvest list — the patterns and fixes the next engagement inherits.
The retro argued from receipts, not adjectives. Not "security review was slow" but the number, and a structural fix instead of an exhortation:
"The security-review queue drifted all engagement — 2.1-day median against a 0.9-day general queue."
"A twice-weekly committed security-review slot in Dan's calendar, written into the Close handoff."
Four patterns went onto the harvest list to improve the next project — among them the suppression-window-plus-recovery-check alert pattern from this very phase.
Go deeper — the full method
A half-day, the pod plus the client's core people, arguing from the assembled evidence: what worked (with the receipts), what didn't (without blame), which gates caught real issues and which were rubber-stamped, which artifacts earned their cost and which were written and never read. Claude assembles the evidence base — the metrics dashboard history, the Retro+ log, the escaped-bug answers, the gate records — so the humans argue from facts, not vibes. An honest retro written by a flattering drafter is worthless; the humans own the candor.
Three outputs, all concrete: the technical debt log (every known shortcut, with a priority and a suggested timing — logged debt is managed, unlogged debt is a future crisis), the client-facing improvements (what the client's team changes about how they run the system and the loop), and the harvest list — the patterns, skills, hook improvements, and template corrections the next engagement inherits through the kit. Phase C opens the harvest PR; this day decides what goes in it.
"Great teamwork, communicate better next time" — written for management, useless to everyone. Honest, specific, evidence-backed, or skip the meeting and admit it.
Now watch the whole thing happen
The two weeks, end to end
You've got the ideas; here's the actual rhythm at Harbor, running inside the hypercare window. Week one makes the system visible — agree what healthy means, capture the baseline, build the dashboards and the alerts. Week two proves the response and learns — write the playbook, drill it, run the retrospective, hand over the watch. Step through it.
The session that shapes the phase
The pod and the client's operations walk every failure scenario and every top-priority journey, answering for each: what healthy looks like, what degraded looks like, who gets woken, who's told in the morning. Operations' answers win ties. The monitoring scope — which stack, which dashboards, which owners — lands in writing.
Measure normal; make it visible
Capture the production baseline from the first hypercare week — request rates, latency, error rates, queue depths — each with its measurement period. Paths production hasn't exercised get modeled values, flagged with a revisit date. Dashboards go live in the client's stack: system, application, and the business layer in their own reporting language.
Every threshold derived, and confirmed
Each critical failure mode gets an alert with a baseline-derived condition, a severity, a named recipient, and a playbook link — the derivation written next to the number. The operators confirm every threshold; that's the named human stop of the phase. The fatigue review cuts anything that would fire weekly without demanding action.
Write down what to do when it fires
Per alert: what it means, the first diagnosis steps, the severity classification, the escalation names, and the communication templates — what users and stakeholders are told, by whom, while it's happening. It cross-references the repair guide rather than duplicating it.
Fire every alert on purpose
Each critical alert fires for real through a pre-agreed synthetic trigger, and the client's on-call responds from the playbook while the pod observes in silence. Wrong routing, a broken playbook step, a threshold that doesn't trigger — all of it fails here, by appointment. What breaks gets fixed through the loop and re-drilled.
The honest cumulative look
Half a day, the pod plus the client's core people, arguing from assembled evidence: what worked with receipts, what didn't without blame, which gates caught real issues. Three concrete outputs — the technical debt log, the client-facing improvements, and the harvest list the next engagement inherits.
Hand over the watch; read the metric honestly
Hypercare closes on its agreed date with a deliberate handover: the dashboards, the pager, and the playbook are formally the client's, with the pod one escalation away until Close. The gate runs; the sponsor signs at steering — and hears the outcome metric's first honest production read, caveats welded on.
Go deeper — the full method
The default calendar is two weeks, deliberately inside hypercare — week one needs the production baseline to accumulate; week two needs the pod still present for the drill. The gate falls at hypercare's end, closing both together.
Week one — see the system
- Days 1–2, what does healthy mean? The session that shapes the phase: the pod and operations walk the RUNBOOK's failure scenarios and the top journeys, capturing one line each — healthy, degraded, who is woken, who is told in the morning. The monitoring scope lands in writing: which stack (theirs), which dashboards, which channels, who owns each dashboard.
- Days 3–4, the baseline and the dashboards. The production baseline is captured from real traffic, each number with its measurement period; unexercised paths derive from modeled data, flagged with a revisit date. Dashboards go live in the client's stack: system health, application health, and the business layer co-built with their data lead in their own reporting language.
- Day 5, the alert definitions. Every critical RUNBOOK failure mode gets an alert with a baseline-derived condition, a severity, a named recipient, a response expectation, and a playbook link — the derivation written next to the threshold. Confirmation is a recorded act: the operations names go on the review of the change that ships the alert rules. The alert-fatigue review runs on everything proposed.
Week two — prove the response, and learn
- Days 6–7, the incident playbook. Written against the alert table: per alert, meaning, first three diagnosis steps, severity classification, escalation names, and communication templates. It cross-references the RUNBOOK rather than duplicating it.
- Day 8, the alert drill. Each critical alert fires for real through a pre-agreed synthetic trigger; the client's on-call responds from the playbook while the pod observes silently. What the drill breaks gets fixed through the loop and re-drilled; the Quality Engineer records every alert's trigger, detection, routing, responder, and outcome for the gate packet.
- Day 9, the retrospective. A half-day arguing from assembled evidence, with three concrete outputs: the technical debt log, the client-facing improvements, and the harvest list the next engagement inherits.
- Day 10, hypercare ends, the gate. Hypercare closes on its agreed date with a deliberate handover of the watch. The gate check runs; Claude drafts the Close & Transfer handoff. Steering: the gate sign-off, the billing milestone, and the outcome metric's first honest production read — stated with its caveats.
A too-thin baseline ships thresholds as modeled-with-revisit-date rather than waiting for a storm. A badly failed drill is fixed, re-drilled, and — if the gap is people, not config — flagged at steering. A political retrospective is held to "for the teams, not for performance reviews." A defect-heavy hypercare is a quality finding for the retro, not a reason to extend forever.
How the phase — and hypercare — ends
A gate, not a calendar — and it closes the watch with it
Phase 9 doesn't end because two weeks passed. It ends when a specific list is true and a named human on each side signs to advance — and because the gate falls at hypercare's end, signing it is also the moment the pod formally hands over the watch.
This is also a billing milestone, and it carries the engagement's most delicate moment: the outcome metric's first honest production read. The credibility protected all engagement gets spent or banked here — so the number is stated with every caveat attached, never oversold.
The headline checklist items have teeth: a drill where every critical alert actually fired and was answered by the client's own on-call, and a retrospective with real receipts, not platitudes.
The outcome metric read spectacularly — and was deliberately underclaimed:
Steering heard it with the caveat welded on: complex claims are mostly still in flight, so the overall median can't be read fairly until a full quarter. The sponsor signed — billing milestone 7. The watch became Harbor's.
The full checklist — tick it:
- Every top-priority feature has at least one metric on a dashboard with a named owner
- Every critical failure mode has an alert, each threshold derived from the baseline (or flagged modeled) and confirmed by the client's operations
- The alert-fatigue review ran — nothing ships that would fire weekly without demanding action
- The incident playbook covers every critical alert — detection, diagnosis, escalation names, templates
- The drill happened — every critical alert fired and was answered by the client's own on-call; what failed was fixed and re-drilled
- The watch was handed over — dashboards, paging, and the playbook are formally the client's
- The retrospective exists with receipts — findings, the debt log, and the harvest list
- The outcome metric has its first production read on the scorecard, caveats stated
- A named human on each side approved the advance
Go deeper — the full method
A phase does not end because two weeks went by. It ends when a specific list is true and a named human on each side signs to advance — gates report, humans decide. Phase 9 closes — and hypercare ends with it — when all of these are true:
- Every top-priority feature has at least one observable metric on a dashboard with a named owner.
- Every critical RUNBOOK failure mode has an alert; every threshold is derived from the measured baseline (or explicitly flagged as modeled, with a revisit date) and confirmed by the client's operations. (teeth)
- The alert-fatigue review ran: nothing ships that would fire weekly without demanding action.
- The incident playbook covers every critical alert — detection, diagnosis, escalation names, communication templates.
- The drill happened: every critical alert fired in a controlled way and was answered by the client's own on-call from the playbook; what failed was fixed and re-drilled. (teeth)
- The watch was handed over: dashboards, paging, and the playbook are formally the client's, with the pod one escalation away until Close.
- The retrospective exists with receipts: findings, the technical debt log, and the harvest list — concrete, honest, owned.
- The outcome metric has its first production read on the scorecard, caveats stated.
- The Close & Transfer handoff exists: monitoring inventory, drill record, debt log, open items with owners.
- A named human on each side approved the advance.
This is also a billing milestone, and it carries the engagement's most delicate moment: the outcome metric's first honest production read. The credibility protected all engagement gets spent or banked here — so the number is stated with every caveat attached, never oversold.
How it goes wrong
The failure modes, and the defense against each
Every one of these has happened to someone. Knowing them by name is half the defense — and Harbor's structure caught two of them in the act.
| The trap | What it looks like | The defense | At Harbor |
|---|---|---|---|
| Thresholds by intuition | "500ms feels right" set without a measured baseline | Measure normal; alert at a stated multiple; write the derivation down. | Every threshold derived; modeled ones flagged with a revisit |
| Our monitoring, their pager | Alerts built in a stack the client's team doesn't watch | Their stack, their names, their session. | All wired into Harbor's own Azure Monitor workspace |
| Alert fatigue on day one | Fifty alerts; the team ignores all of them within a month | Replay against history; cut anything that fires weekly without action. | Caught — two proposals cut at the fatigue review |
| The undrilled pager | Routing and playbooks that have never actually fired | Fire each one on purpose; the client's on-call responds. | Caught — the day-8 silent-night routing typo |
| The dashboard nobody owns | A beautiful screen no one reviews | Every dashboard has a named owner or it doesn't ship. | Three dashboards, each with a named Harbor owner |
| The platitude retrospective | "Great teamwork, communicate better" — written for management | Argue from receipts; honest and specific or skip it. | "2.1 days vs 0.9," a calendar slot, not an exhortation |
| Hypercare that never ends | Quietly extending the watch because closing feels risky | Extend deliberately with the sponsor, or close on schedule. | Closed on its agreed date — two weeks, as written |
Go deeper — the full method
Named failure modes, because every one of these has happened to somebody. Knowing them by name is half the defense.
- Thresholds by intuition. "500ms feels right" is not engineering. Measure the baseline, alert at a stated multiple, write the derivation down. A threshold that can't explain itself can't be tuned later.
- Our monitoring, their pager. The pod builds alerts in its own image, in a stack the client's team doesn't watch, routed by assumptions. The week we leave, it's noise. Their stack, their names, their session.
- Alert fatigue shipped on day one. Fifty alerts feels thorough and guarantees the team ignores all of them within a month — and then misses the real one. Fewer, actionable, derived, reviewed.
- The undrilled pager. Routing and playbooks that have never fired, discovered broken during the first real incident. The drill is to alerts what the rehearsal was to rollback: the difference between a procedure and a wish.
- The dashboard nobody owns. A beautiful screen no one reviews is decoration. Every dashboard has a named owner or it doesn't ship.
- Monitoring the system but not the business. All RED metrics, no outcome metric — the client can see requests but not whether the thing they bought is working. The business layer is co-built with their data lead, in their language.
- The platitude retrospective. "Great teamwork, communicate better next time" — written for management, useless to everyone. Honest, specific, evidence-backed, or skip the meeting and admit it.
- Debt left unlogged. The shortcuts everyone knows about but nobody wrote down become next year's crisis with no paper trail. Logged debt is managed debt; the log is a gate item for a reason.
- Hypercare that never ends. Quietly extending the watch because closing feels risky. Extend deliberately with the sponsor and a new end date, or close on schedule — an open-ended hypercare is a handoff that's failing in slow motion.
When Phase 9 is done
The system can be seen, the response is proven, and the lessons are written down
Monitoring closes with a watched system in the client's own hands, a response drilled until it works, and an honest account of what the engagement learned — debt logged, patterns harvested. One thing remains: proving the client can run all of it without you. Where to go next: