The code was done weeks ago. The system is live, watched, and quiet. So why isn't the engagement over?
Because "done" was never about the code. It's about whether the client can run all of this — the system, the build loop, the judgment — without us in the room.
The engagement ends when the client can run this without us — proven by use, observed, unassisted. Every idea below is paired with the real thing: the close of the Harbor Mutual engagement, a fictional but fully worked project.
Why this phase exists at all
Leaving is a deliverable, not an afterthought
Most projects end by drifting: the team rolls off, a few people keep answering questions on the side, and one day everyone notices it's over. That's not a close. That's a slow leak — and it leaves the client depending on people who are already gone.
The whole engagement was built so the client's own people learned to run it as they went. Phase C is where that claim stops being an intention and gets proven the same way everything else got proven — by use, observed, without help.
So this phase has a job nobody enjoys: make ourselves unnecessary, on purpose, and prove it. Hands off the keyboard early, out of the room by the test, out of the building at the end.
By the time Phase C started, Harbor had been doing the handover all along without a ceremony for it. Their lead engineer co-signed every architecture decision back in March. Their newest hire cold-verified the README as a total stranger. Their on-call answered the incident drill.
Phase C wasn't a handover event — it was the moment the handover that had been happening for five months finally got tested, and then proven on the record.
You leave a dependency, not a capability. The client owns a system they can't actually run — and the people who could are billing somewhere else now.
Go deeper — the full method
The engagement does not end when the code is done — the code was done weeks ago. It ends when the client can run all of this without us: the system, the harness, the loop, and the judgment. Phase C is where that claim stops being an intention and gets proven the same way everything else in this standard gets proven — by use, observed, without help.
The whole engagement built toward this phase. The kit was adapted in the open in Foundation so the client's engineers saw how it works. Their lead engineer co-signed every ADR. Their platform engineer executed every promotion. Their newest hire cold-verified the README. Their on-call answered the drill. Phase C is not a handover event; it is the moment the handover that has been happening all along is tested — and then we leave, cleanly, and the standard itself learns from the engagement before the door closes.
Where the engagement is when this phase opens: live in production, observable, drilled, retrospected. The client's operators hold the watch. One thing remains unproven — that the client's team can run the build loop, intent through deploy, without us in the driver's seat. That is the single question Phase C exists to answer.
The rule that governs the whole method
The harness was always the operating system — never us
Throughout the engagement, the AI did the drafting and most of the typing while humans made every decision and carried the accountability. Phase C asks the question that rule was always building toward: was the thing running this project the harness, or was it secretly us?
The harness — the project's instructions, its specs, its skills and guardrails, all versioned in the client's own repo — was the operating system the whole time. The AI ran inside it. The pod ran inside it.
If that's true, then the client's people can step into the same seats and the machine keeps working. If it's not true — if something only lived in a pod member's head — this phase is where that lie surfaces, while there's still time to fix it.
The proof was continuity. During the transfer, the same AI agent kept working inside the same harness — only now the client's engineers were the ones driving it.
"Could Harbor operate this if the pod vanished tonight?"
Nothing about the agent changed when the hands on the wheel did. That's how you know it was the harness running the project, not the people leaving.
Go deeper — the full method
Claude is not handed over — Claude was never ours. The harness, the agents, the skills, and the keys have been the client's since Foundation; what transfers in Phase C is the human judgment around them. Claude's work this phase is bounded and specific:
- Audits the harness, read-only. A sweep of the repo for anything that would strand the client: skills with pod-specific assumptions, hooks referencing paths or permissions only we had, conventions that live in pod memory instead of CLAUDE.md. Every finding becomes a documented fix the client Setup Owner merges.
- Drafts the final handoff report from the engagement's own records: every gate, every phase report, the metrics history, the debt log, the open items with owners.
- Drafts the harvest PR against our standard repo from the Phase 9 retrospective's list — client specifics stripped, patterns generalized — for the pod to review and our standard's deputy to merge.
- Keeps working for the client. During the shadow flip and the close gate, Claude is the same agent inside the same harness — driven now by the client's Orchestrators. That continuity is the proof the harness, not the pod, was the operating system.
What Claude never does: drive the close-gate spec (nobody from our side does — that is the test), approve the client team's work as the gate evidence, or write the retrospective's candor for them.
What these three weeks must answer
Four questions — and deliberately nothing else
Everything Phase C does exists to answer exactly four questions. New features are out of scope unless they're the vehicle for the test — and a "quick follow-on" is a brand-new agreement, not a quiet extension of this one.
- Can their people run the loop? First with us checking, then one real change end to end with nobody from our side driving.
- Is the harness fully theirs? A sweep finds nothing only we understand, and a named client owner has merged changes themselves.
- Is the record complete and delivered? Every report, the outcomes dashboard, the debt log — in their hands, in their tooling.
- Did we leave cleanly — and did the standard learn? Our access revoked and audited; the lessons sent home before they evaporate.
- Ines ran a HIGH-risk change end to end, hit a guardrail, escalated correctly, shipped to production — zero pod hands.
- The audit found two small things; Wes, Harbor's owner, merged the fix for each — and a change of his own on top.
- Every report 0 through C, the dashboard re-pointed to Priti, the debt log in Harbor's tracker with Harbor owners.
- Every pod seat and key revoked and confirmed by Harbor's security; four patterns sent back into our own standard.
At the very end, "could you just build one more thing while you're here?" feels generous to say yes to. It isn't. Scope at close is still scope — it burns the calendar the test needs, and it belongs to a new SOW with its own Phase 0.
Go deeper — the full method
Phase C answers four questions, and nothing else:
- Can their people run the loop? Their engineers orchestrate real specs with our Checkers first; then one real spec end-to-end with nobody from the pod driving.
- Is the harness fully theirs? The audit finds nothing undocumented; the client Setup Owner is named and has merged harness changes themselves.
- Is the record complete and delivered? Every phase report, the outcomes dashboard, the debt log — handed over, in their hands, in their tooling.
- Did we leave cleanly — and did the standard learn? Access revoked and audited; the harvest PR opened against our own repo before the engagement's lessons evaporate.
New features for the client are out of scope unless they are the vehicle — the specs run during this phase are real work from the client's own backlog, chosen because the close gate must be run on something that matters. A follow-on engagement, if there is one, is a new SOW with its own Phase 0, not a quiet extension of this one. Scope at close is still scope; it burns the calendar the gate needs, and the Pod Lead holds that line precisely because everyone else in the room has an incentive not to.
How the transfer actually starts
The shadow flip — their hands, our eyes
You can't prove someone can drive by describing the car. The transfer opens by reversing the roles that built the whole project: the client's engineers take the driver's seat on real work, and the pod is demoted to checking.
This is the inverse of how the build began — back then we drove and they watched; now they drive and we watch. The pod becomes Checkers only: review the work, coach by question ("what does the spec say about that path?"), and never, ever grab the keyboard.
And every time a Checker wanted the keyboard gets written down — because each one is a gap in the transfer with a name on it. The point isn't to look good; it's to find what hasn't transferred yet, while it's still cheap to fix.
Three real changes from Harbor's own backlog rode the loop with Harbor driving and the pod checking:
0048 — intake supervisor daily digest (Ines)
0050 — claims-search timeout messaging (Wes)
The bar didn't move an inch: plan-then-build, the automated grader, a non-author reviewer, merge deploys. Jonah and Sara coached by question and logged every keyboard itch as a transfer gap.
These aren't exercises invented for the occasion. They're ordinary backlog items, mixed difficulty, each sized to finish within the week. A rehearsal with no stakes teaches the rehearsal, not the job.
Go deeper — the full method
The shadow flip is week one's work. The client's engineers take the Orchestrator seat on real specs from their own backlog — triage with their PO, spec writing, bounds, plan approval, driving the agent — while the pod serves only as Checkers. Coaching happens by question ("what does the spec say about that path?"), never by taking the keyboard. At least three real specs ride the loop this way, and the bar is the same bar: the grader runs, a non-author approves, merge deploys.
The specs come out of the client's own intent triage like any other work — ordinary backlog items, mixed tiers, each sized to merge within the week. The Pod Lead confirms each one is real work, not an exercise built for the occasion. A rehearsal with no stakes teaches the rehearsal, not the job.
The pod Checkers log every place they had to coach — every moment they wanted the keyboard — because each one is a transfer gap with a name. The Quality Engineer rolls those logs into the shadow-flip spec record: per spec, who orchestrated, who checked, the grader and Checker outcomes, and the gaps.
Finding what only we understand
Audit the harness for anything living in our heads
The most dangerous thing you can leave behind is knowledge that never made it into the repo — a convention "everyone just knows," a script that only works on a pod laptop, a setting nobody wrote down. It works fine right up until the people who knew it are gone.
So the harness gets swept — the AI reads the whole repo looking for assumptions that only the pod could satisfy, and a human walks every skill, hook, and convention asking one blunt question: could their team operate this if we vanished tonight?
Every finding becomes a documented fix — and here's the clever part: the client's owner merges each one. The gap closes, and the client owner's real merge history starts to exist at the same time. Two birds.
After a five-month engagement, the sweep found only two things — because every harness change had been a reviewed PR all along, in the open, where Wes could see it:
The grader's prompt said "review per MCKRUZ standards" — a phrase only the pod could resolve. Rewritten to state the actual rule inline.
A hook read a lint config from a path that existed on pod machines but wasn't in the repo. Pinned into the repo; path made repo-relative.
Go deeper — the full method
The harness audit runs in parallel with the shadow flip: Claude's read-only sweep plus the Setup Owner's walk of every skill, hook, and convention, asking one question — could their team operate this if we vanished tonight? It is a sweep for the indispensable: a skill with pod-specific assumptions, a hook referencing a path or permission only we had, knowledge living in pod memory instead of CLAUDE.md.
Findings get fixed by documented PRs that the client Setup Owner merges, which kills two birds: the gap closes, and the client owner's merge history starts being real. "Ask the pod" must return zero results before the gate — that is the bar the audit exists to clear.
One person's head still holds how something really works, discovered the week after they leave. The harness audit exists for exactly this. Two findings after a five-month engagement is short precisely because the open-adaptation habit from Foundation made every harness change a reviewed PR the client could already see.
An owner, not a label
The harness gets a real owner — proven by their merge history
It's easy to put a name in a slide: "Harbor's Setup Owner is Wes." It means nothing. A harness with a named owner who has never actually changed it doesn't have an owner — it has a label.
The test for real ownership is the git log, not the org chart. Before we leave, the client's Setup Owner must have merged harness changes themselves — the audit fixes, and at least one improvement of their own.
And the rule that protected us protects them too: no role without a deputy. A named client engineer reviews the owner's harness changes, the same way our deputy reviewed ours. The safety practice transfers along with the keys.
Wes didn't just merge the two audit fixes. He shipped a harness change of his own, unprompted:
Tom became his named deputy. "Harbor has a Setup Owner" stopped being an org-chart claim and became a git log.
Go deeper — the full method
The client Setup Owner is the named client engineer who owns the harness after close — and who must have merged harness changes themselves before we leave, not merely been told about them. Merge history is the test: the audit fixes, plus at least one change of their own, or the harness has no owner — it has a label.
By week two the client Setup Owner ships at least one harness change of their own — their improvement, their PR, their deputy arrangement on their side. The no-role-without-a-deputy rule transfers too: a named client engineer reviews the client Setup Owner's harness changes, the way our deputy reviewed ours.
A client owner who was named in a deck but never merged anything. The defense is mechanical: the git log, not the org chart. The audit fixes plus one change of their own, or it didn't happen.
The test the whole phase exists for
The close gate — one real change, solo, with us silent in the room
This is the moment everything was building toward. The client team runs one real change end to end — deciding it's worth doing, writing the spec, driving the agent, grading it, reviewing it, merging, deploying — with nobody from our side driving. We sit in the room, silent, taking notes.
The change has to be real: something the client genuinely needs, with a real risk level and real consequences if it goes wrong. A toy change picked because it can't fail proves nothing, and everyone in the room knows it.
And the one unbreakable rule: help voids the run. Any answer, hint, or keyboard touch from our side, and it doesn't count — the gap that question revealed becomes a finding, gets fixed, and the gate re-runs on a different real change. "They got through it with a little help" is a fail, for the same reason it was a fail at the README checkout.
The gate ran on a HIGH-risk change with a deadline: 0049 — decommission the legacy intake fallback, due at day 30 of the rollout. Ines drove the whole thing.
The wobble the observers were waiting for arrived right on cue:
The agent's plan drifted toward editing a protected deploy file. The hook blocked it cold.
Ines stopped, escalated the change to Wes as its own reviewed PR — exactly the rule she'd learned eight weeks earlier. The pod said nothing.
The best moment of the gate wasn't a flawless run — it was the agent drifting, the rails catching it, and a Harbor engineer responding with the right judgment, unprompted. Mechanics transfer in documents; judgment only shows up under observation.
Go deeper — the full method
The close gate is week two's defining event. One real spec — something the client genuinely needs, with a real risk tier — runs the loop end to end with nobody from the pod driving: their triage, their spec, their bounds, their plan approval, their Checker, their merge, the automatic deploy. The client's triage picks the spec; the Pod Lead confirms it clears the real-spec bar — real need, real tier, real consequences — before the run is scheduled.
The pod observes the way the QE observed the cold runs: present, silent, taking notes. The Quality Engineer's record captures each loop step with timestamps and names, every stall, every guardrail event and how the team responded, the merge and deploy that closed it, and the verdict. Stalls are data. Help — any answer, hint, or keyboard touch from our side — voids the run.
A failed or wobbly run is information, not embarrassment: the gap gets named, fixed — more reps, a harness clarification, a missing playbook line — and the gate is re-run on a different real spec. "They got through it with a little help" fails the close gate for the same reason it failed the README checkout.
A copy change with a LOW tier and no stakes, chosen so it cannot fail. It proves nothing, and everyone in the room knows it. Real spec, real tier, real consequences — or the gate did not run.
Leaving so you provably can't come back
The clean exit — revoke our access and prove it
A close that ends with the pod still able to reach production isn't a close. It's a "we'll clean up the seats next sprint" that, months later, becomes a finding on somebody's security audit. Leaving cleanly is the last deliverable of the security posture the whole engagement ran on.
Every pod seat, token, repo permission, environment role, and vault key gets removed — as a dated checklist item, confirmed against the client's own audit trail by their security. Not a cleanup intention; a checklist with a date and a signature.
The list isn't improvised on the last day. It's drafted in week one from everything the engagement was ever granted, so week three is execution, not archaeology. The engagement ends with the pod provably unable to touch the system — which is not distrust, it's the posture being honored to the end.
Rob and Dan walked the revocation as a checklist, item by item, then confirmed it against Harbor's audit trail. By end of day the pod provably could not touch the system it built.
"We'll sort the seats out next sprint" un-revokes the whole thing. Revocation is a gate item with a date and a security signature, or it didn't happen.
Go deeper — the full method
Access revocation is the deliberate, audited removal of every pod credential, seat, and permission — a checklist item with a date, not an eventual cleanup. Every pod seat, token, repo permission, environment role, and vault access is removed on a checklist, confirmed against the client's audit trail by their security.
The Setup Owner drafts the checklist in week one from everything the engagement was ever granted — the Phase 0/1 access checklist, CI secrets, environment roles, vault policies — so that week three is execution, not discovery. The engagement ends with the pod provably unable to touch the system, which is not distrust; it is the last deliverable of the security posture the engagement ran on, bookended with Phase 8's secrets rotation: we never held production secrets, and now we hold nothing at all.
"We'll clean up the seats next sprint." Months later the pod can still reach production, which is a finding on somebody's audit eventually. Revocation is a dated checklist item confirmed by their security, not a cleanup intention.
The two things that quietly un-do a close
Capability is a paid deliverable — not a parting gift
There are two warm, generous instincts at the end of a good engagement, and both of them quietly cancel the close. The first leaks the standard's value away for free. The second un-transfers everything you just proved, one Slack message at a time.
The capability was sold, not given. The client's team can now run this loop because a training workstream was priced into the contract from day one. That's why the close gate can be passed at all — it was paid for. If the client wants more help after close, that's a new agreement, made in daylight.
And the close means the close. The reflex to keep answering questions "just for a while" feels kind and slowly un-does the transfer. Future help is a future agreement, not an indefinite off-the-record retainer.
The engagement also paid the standard back. Four patterns Harbor's project surfaced went home through a single PR against our own repo — client specifics stripped, patterns generalized:
An engagement that ends without a harvest taught the standard nothing. The next pod starts where Harbor finished — which is the whole reason the kit exists.
Hypercare reflexes that outlive their window are the most common way a clean close rots. It feels generous, and it un-transfers the engagement one free answer at a time.
Go deeper — the full method
The harvest is the mandatory improvements PR against our own standard and kit, carrying what this engagement taught: generalized skills, corrected templates, patterns worth repeating. Opened in this phase from the Phase 9 retrospective's list, reviewed by the pod, merged by the standard's deputy. This is the compounding asset doing its compounding; an engagement that ends without a harvest taught the standard nothing, and the next pod re-discovers this engagement's lessons at a client's expense. The harvest is a gate item, not a virtue.
The capability the client now holds was priced into the SOW from Phase 0 — the training workstream is why the close gate can be passed at all. If the client cannot name engineers to run the loop, the close gate cannot be passed, and that conversation belongs at steering weeks before this phase. Phase C reveals the transfer's state; it cannot manufacture one.
Hypercare reflexes outlive their window and the pod keeps answering questions for free, indefinitely, off the record. It feels generous and it un-transfers the engagement one message at a time. The close means the close; future help is a future agreement made in daylight.
Now watch the whole thing happen
The three weeks, end to end
You've got the ideas; here's the actual rhythm at Harbor. The pod gets deliberately less necessary as the weeks go: hands off the keyboard in week one, out of the room by the gate in week two, out of the building by the end of week three. Step through it.
Their hands take the wheel
The client's engineers take the driver's seat on real backlog work; the pod drops to Checkers only. The flow check and the intent triage are run by the client now — the queue numbers are theirs to read. In parallel, the harness audit's read-only sweep starts.
Coach by question, never by keyboard
At least three real changes ride the loop with the client driving and the pod checking. Every place a Checker wanted to grab the keyboard gets logged — each one a named gap in the transfer. The bar doesn't move: plan, grader, non-author reviewer, merge.
Findings fixed by the client's owner
The harness audit's findings — anything only the pod understood — get fixed as documented PRs that the client's Setup Owner merges. At Harbor, two findings, both merged by Wes, whose real ownership history starts here.
A real change, with real stakes
The client's triage picks the close-gate change from their own backlog; the Pod Lead confirms it clears the bar — real need, real risk tier, real consequences. At Harbor: decommission the legacy fallback, HIGH risk, with a hard deadline two days out.
Solo, end to end, observed and silent
The client runs the whole loop with nobody from the pod driving. The pod observes — present, silent, taking notes. A guardrail blocks the agent; the client engineer escalates correctly, unprompted. Stalls are data; any help from our side voids the run.
The client owner's own change
The client's Setup Owner ships at least one harness change of their own — their improvement, their PR, their named deputy reviewing it. The "no role without a deputy" rule transfers along with the keys.
Everything, in their tooling
The final handoff report; every phase report from 0 through the close; the outcomes dashboard re-pointed to client ownership with its caveats intact and the quarter-read date on their calendar; the debt log with owners and dates. Delivered into their tooling, not left in ours.
Provably unable to touch it
Every pod seat, token, permission, role, and key is removed on the checklist and confirmed against the client's audit trail by their security. By end of day, the pod provably cannot reach the system it built. Their security signs the record.
Send the lessons home, then close
The harvest PR opens against our own standard repo with the engagement's lessons generalized, plus a retro file. The close steering hands the sponsor the record, the gate evidence, and the metric's read — and the last milestone bills. The engagement ends.
Go deeper — the full method
The default calendar is three weeks — long enough for the role flip to be real, short enough that the pod's presence doesn't quietly become a dependency again.
Week one — their hands, our eyes
The shadow flip: the client's engineers take the Orchestrator seat on real specs from their own backlog while the pod serves only as Checkers, coaching by question, never by keyboard. At least three real specs ride the loop this way, the bar unchanged. The pod Checkers log every transfer gap. The harness audit runs in parallel — Claude's read-only sweep plus the Setup Owner's walk — with findings fixed by PRs the client Setup Owner merges.
Week two — the close gate
One real spec, real risk tier, runs the loop end to end with nobody from the pod driving: their triage, their spec, their plan approval, their Checker, their merge, the automatic deploy. The pod observes — present, silent, taking notes. Help voids the run; a wobbly run is re-run on a different real spec. The client Setup Owner ships one harness change of their own this week, with their own named deputy reviewing it.
Week three — the clean exit
The record hands over: the final handoff report; every phase report 0 through 9; the outcomes dashboard re-pointed to client ownership; the debt log with owners and dates — delivered into their tooling. Access revokes on an audited checklist confirmed against the client's audit trail. The harvest PR opens against our standard repo. The close steering gives the sponsor the engagement record, the outcome metric's current read with its caveats, and the formal end of the SOW; the last billing milestone lands with the close gate's evidence attached.
The close gate fails → name the gap honestly at steering, fix it, re-run on a different real spec — leaving on schedule with an unpassed close gate is abandoning, with paperwork. The backlog has no real specs → that is a triage problem the client's PO now owns; solving it together is transfer work. "One more feature" arrives dressed as closure → a follow-on conversation with its own SOW, not the gate window. The pod becomes the pager again → hold the Phase 9 handover: their watch, the pod one escalation away, an escalation not a reflex.
How it goes wrong
The failure modes, and the defense against each
Every one of these is a way a close quietly fails while looking finished. Knowing them by name is half the defense — and Harbor's structure caught the ones that count.
| The trap | What it looks like | The defense | At Harbor |
|---|---|---|---|
| The toy close gate | A no-stakes copy change picked so it can't fail — proving nothing | Real change, real tier, real consequences, or the gate didn't run. | HIGH-risk fallback decommission, dated |
| The helpful observer | Someone from the pod answers "one little question" mid-run | Help voids the run — same rule as the cold checkout. The question is a finding. | Held — pod silent through the wobble |
| The indispensable pod member | One head still holds how something really works — found the week after they leave | The audit; "ask the pod" must return zero results before the gate. | Two findings, both documented & fixed |
| The owner in name only | A client owner named in a deck who never merged anything | Merge history is the test: audit fixes plus one of their own. | Wes merged both fixes + a skill of his own |
| Access that lingers | "We'll clean up the seats next sprint" | A dated checklist confirmed by their security, not an intention. | Revoked & audit-confirmed, Dan signed |
| Scope dressed as closure | "Before you go, could you just…" | New work is a new SOW. The Pod Lead holds the line. | Held — follow-on = a new Phase 0 |
| The skipped harvest | The pod rolls off and the lessons-PR never opens | The harvest is a gate item, not a virtue. | Four patterns sent home |
| The quiet retainer | Free answers continuing indefinitely, off the record | The close means the close; future help is a future agreement. | Clean break — who they call now: their own team |
Go deeper — the full method
- The indispensable pod member. One person's head still holds how something really works, discovered the week after they leave. The harness audit exists for exactly this; "ask the pod" must return zero results before the gate.
- The toy close gate. The solo spec is a copy change with a LOW tier and no stakes, chosen so it cannot fail. It proves nothing, and everyone in the room knows it. Real spec, real tier, real consequences — or the gate did not run.
- The helpful observer. Forty minutes into the solo spec, someone from the pod answers one little question. The run is void — same rule as the cold checkout, same reason. The gap the question revealed is the finding; record it, fix it, re-run.
- The Setup Owner in name only. A client owner who was named in a deck but never merged anything. Merge history is the test: the audit fixes, plus at least one change of their own, or the harness has no owner — it has a label.
- Access that lingers. "We'll clean up the seats next sprint." Months later the pod can still reach production, which is a finding on somebody's audit eventually. Revocation is a dated checklist item confirmed by their security, not a cleanup intention.
- Scope dressed as closure. "Before you go, could you just—" at close burns the calendar the gate needs. New work is a new conversation with a new SOW; the Pod Lead holds that line precisely because everyone else in the room has an incentive not to.
- The skipped harvest. The pod rolls off to the next engagement and the PR never opens; the standard learns nothing, and the next pod re-discovers this engagement's lessons at a client's expense. The harvest is a gate item, not a virtue.
- The quiet retainer. Hypercare reflexes outlive their window and the pod keeps answering questions for free, indefinitely, off the record. It feels generous and it un-transfers the engagement one Slack message at a time. The close means the close; future help is a future agreement.