Troubled Project Turnaround: A Practical Guide
Fire Recovery Playbook — From Crisis to Controlled Recovery
Using an SAP Project as the Case Study — A Practical Guide for PMOs, Project Managers, and Executive Leadership
July 2026 Paddy&Water Inc.
Introduction: Three Truths Common to Troubled Projects
| To the Reader Holding This Guide |
|
If you are reading this guide, you are probably right now in the middle of a troubled project, standing at its threshold, or sensing a premonition that “this might turn into a troubled project.” This guide is not a “theory of recovery” but a “recovery playbook.” It describes, at the level of concrete action, what to do in the first 72 hours starting today, what to decide within one week, and what to achieve within one month. Recovering a troubled project does not mean “saving the project” — it means “rebuilding the project.” You cannot achieve a turnaround by simply “trying harder” with the same members, the same plan, and the same structure. |
Three Truths of Troubled Projects
| Truth | Content | Why It Matters |
| Truth 1: The problem is not the “visible problem” | Surface-level problems (schedule delays, failed tests) are symptoms, not causes. The root cause is one of “failed business design,” “absent decision-making,” or “unbounded scope expansion” | If you only address the symptoms while leaving the root cause untouched, the same symptoms will recur even after the turnaround |
| Truth 2: Adding people does not fix it | Throwing additional staff into a troubled project is like “pouring oil on a fire.” The catch-up cost of new joiners drains the team. If the root cause is a design quality problem, headcount is irrelevant | What a troubled-project turnaround needs is “quality,” not “quantity.” Cutting scope must come first |
| Truth 3: Recovery is a series of decisions | Recovery is not a technical exercise but a continuous series of decisions. The moment uncomfortable decisions — “what to give up,” “who to remove,” “whether to postpone Go-Live” — are deferred, the fire reignites | The single most important skill for a recovery leader is not “precise technical knowledge” but “fast decision-making backed by strong communication skills” |
Chapter 1: The First 72 Hours — From Confirming the Crisis to Initial Action
1.1 ”Forced Agreement on the Current Situation” Before Declaring Recovery
The first thing to do in a troubled-project turnaround is neither “making a plan” nor “hunting for the guilty party.” In the first 24 hours, everyone must come to share the same understanding of “what state the project is actually in.” Without this, even a well-crafted recovery plan will fail, because each member will act on different assumptions and discussions will never converge.
| The 72-Hour Timeline (Standard Action Plan at the Start of Recovery) |
|
[Day 1 (0–24 hours): Freeze the Current State and Gather Facts] 09:00 Executive leadership formally approves and announces the assignment of the recovery leader 10:00 Take a “frozen snapshot” of all issue and risk logs (all subsequent changes are logged from this point) 11:00 30-minute interview with the PMO/PM: “What is the problem, and who has failed to decide what?” 13:00 30-minute interview with the lead consultant/SI leader (conducted separately from the PM) 14:00 30-minute interview with on-site key users (business-side staff), conducted without consultants present 16:00 Answer the “Five Questions” (see below) based on the information gathered 17:00 Request executive leadership to convene a steering committee for the following day [Day 2 (24–48 hours): Report to Executive Leadership and Demand a Decision] 09:00 Steering committee: report the current state “without sugarcoating” (see the “Five-Slide Principle” below) 10:30 Obtain documented “delegation of recovery authority” from executive leadership 13:00 One-on-one meetings with key stakeholders (client-side PMO, IT department head, key business unit heads) 15:00 Draft the composition of the recovery team (who to remove, who to place at the core) [Day 3 (48–72 hours): Provisional Decision on the Recovery Policy and Its Declaration] 09:00 Present three recovery options (scope reduction plans A/B/C) 11:00 Executive leadership and client-side PMO select a recovery option 14:00 Formally notify all project members of the recovery policy 16:00 Formulate and distribute the action plan for the next two weeks (Week 1–2 sprint plan) |
1.2 The “Five Questions” for Diagnosing the Current State
If you can answer the following five questions within 30 minutes after the interviews, the essence of the problem and the direction of recovery become visible. The more items you cannot answer honestly, the more serious the problem.
| # | Question | If YES | If NO or Unknown |
| Q1 | Is the Go-Live date an immovable constraint? | Clarify the reason for the date constraint (legal, contractual, or executive decision). If the constraint is real, treat the date as fixed and focus on compressing scope | Question why the date is set where it is. “Because it’s already decided” is not a reason. In many cases it can be changed — no one has simply asked |
| Q2 | Is the full scope listed out? | Classify the list into Open and Closed, and agree that Closed items cannot be reopened without going through the Change Control Board (CCB) | There is no scope list — meaning no one knows what remains unfinished. Conduct a scope inventory as the top priority (1–2 days) |
| Q3 | Is it clear “whose YES makes something decided”? | Document the decision-maker and communicate it to everyone | The decision-maker is unclear — equivalent to having no escalation path. The first job of recovery is to decide “who decides” |
| Q4 | Is Go-Live achievable with the current team? | Estimate the scope achievable with the current structure and present it to executive leadership (be prepared to accept trade-offs) | The most common case is that the project continues while executive leadership remains unaware that it is “impossible with the current structure.” If it is impossible, present “what must be cut to make it possible” |
| Q5 | Is this a quality (testing) problem or a design problem? | If it is a quality problem: address it through the test regime, automation, and prioritization | If it is a design problem: repeating tests will not raise quality. What is needed first is the decision and the time to redo the design |
1.3 The “Five-Slide Principle” for Reporting to Executive Leadership
| The Five-Slide Principle for Crisis Reporting — A Fact-Based Reporting Format Free of Sugarcoating |
|
Slide 1: Current Position (Facts Only) “Original Plan” vs. “Actual Results”: planned vs. actual completed task counts, weeks of delay, quality KPIs Rule: 80% of the slide should be numbers. Do not include opinions or excuses Slide 2: Root Cause (Narrow It to Three) State clearly, “The root causes of the delay are X,” narrowed to no more than three The phrase “there are multiple, complex contributing factors…” is a sign that the root cause has not actually been identified The most common failure on this slide: writing “the cause of the schedule delay is that the schedule is delayed” (circular reasoning) Slide 3: Options (Three) Option A: Reduce scope to protect the Go-Live date (state the specific reductions) Option B: Postpone Go-Live to protect scope (state the postponement period and additional cost) Option C: Pause the project and redesign it (state the redesign period and cost) “Please think of Option D” is offloading the problem onto executive leadership. The options must always be prepared in advance by you Slide 4: Recommendation and Rationale Clearly state the recovery leader’s recommendation among the three options Failing to make a recommendation is equivalent to passing the problem back to executive leadership Slide 5: What We Need Decided Today (Request for a Decision) List, in bullet points, “the items we want YES/NO decided at this meeting today” For each decision item, specify “who, what, by when, and in what format” |
Chapter 2: A True Diagnosis of the Troubled Project’s Current State
2.1 The Four-Domain Diagnostic Framework
To accurately grasp the reality of the crisis, diagnose independently across four domains: “Scope,” “Schedule,” “Quality,” and “Organization/Structure.” If the recovery leader cannot diagnose these four domains within the first 48 hours, the precision of the recovery plan cannot be guaranteed.
Diagnostic Domain 1: Scope Diagnosis
| Scope Diagnosis | Priority | |
| ☐ | A scope list (WBS/feature list) is documented, and everyone refers to the same definition | H |
| ☐ | Scope additions and changes go through the CCB (Change Control Board) or a clearly defined approver | H |
| ☐ | A list of scope added since the original baseline, together with its approver and approval date, is on record | H |
| ☐ | Features that are essential for Go-Live (Must Have) are clearly distinguished from those that can wait until after Go-Live (Nice to Have) | H |
| ☐ | The current completion rate (%) against total scope is calculated objectively (not as a gut feeling) | M |
| ☐ | It has been confirmed whether “out-of-scope” work is effectively occurring within the project | M |
Diagnostic Domain 2: Schedule Diagnosis
| Schedule Diagnosis | Priority | |
| ☐ | The current schedule delay is accurately tracked “in weeks” (not as a feeling or “slightly behind”) | H |
| ☐ | Tasks on the critical path are identified, and the cascading impact of delays has been calculated | H |
| ☐ | The remaining volume of work to Go-Live (remaining task count/remaining effort) is recorded as an actual measured value | H |
| ☐ | The projected Go-Live date “at the current pace” is calculated, and the gap against the original plan is clear | H |
| ☐ | Milestone achievement/non-achievement is documented, and the number of plan revisions can be traced | M |
| ☐ | Realistic estimates for testing, cutover, and training effort are included in the remaining plan | M |
Diagnostic Domain 3: Quality Diagnosis
| Quality Diagnosis | Priority | |
| ☐ | The number of test cases, pass rate, and number of unresolved bugs are tracked with up-to-date data | H |
| ☐ | The number of bugs that “block Go-Live” (Blockers) is identified, and there is a clear path to resolving them | H |
| ☐ | The quality of design documents (business process flows, screen designs, interface specs) has been assessed as sufficient for use in the next phase of work | H |
| ☐ | It has been confirmed that testing verifies “business scenarios” rather than merely “clicking through screens” | M |
| ☐ | A data migration rehearsal has been conducted, and the time required and error rate have been actually measured | M |
| ☐ | Performance testing (load testing) against the production environment has been conducted, and response times under realistic business volume have been confirmed | M |
Diagnostic Domain 4: Organization/Structure Diagnosis
| Organization/Structure Diagnosis | Priority | |
| ☐ | An authority matrix (DOA) documenting “who can decide what” exists and is known to everyone | H |
| ☐ | A person capable of making project decisions is substantively assigned on the client side and is devoting sufficient time to the project | H |
| ☐ | The consultant/SI-side project leader has been assessed as capable of judgment in both business design and system design | H |
| ☐ | Escalation items are not piling up, and are being decided on a weekly basis | H |
| ☐ | The risk of key members leaving (job change, extended sick leave, transfer) is tracked | M |
| ☐ | It has been confirmed whether on-site key users (business-side staff) can devote sufficient time to the project, and whether they are over-committed to other duties | M |
| ☐ | Whether meetings function as “decision meetings” rather than “status-report meetings”: whether decisions are documented at every meeting | M |
Chapter 3: Building the Recovery Plan — Scope Compression and a Realistic Plan
3.1 The Iron Rule of Recovery: Cutting Scope Is the First Job
The most important decision in recovering a troubled project is “cutting scope.” “Doing everything” is a decision to keep the fire burning. “Giving something up” is a decision to end it. The recovery leader’s job is to decide “what to abandon” and to present that pain to stakeholders in a form they can accept.
| Criteria for Scope Reduction — Priorities for “What to Give Up” |
|
[Step 1: Classify all scope into an MHN list (time required: 2–4 hours)] M (Must Have) — Legally, contractually, or operationally essential for Go-Live H (High Value) — Its absence makes operations inefficient, but it can be worked around manually N (Nice to Have) — Convenient to have, but can be added within one year after Go-Live [Step 2: Estimate whether Go-Live is achievable with M alone] Remaining effort for M-scope only ÷ team productivity → earliest possible Go-Live date → If this date is within the acceptable range → go live with M-only scope → If this date is outside the acceptable range → decide whether to trim M further or postpone the date [Step 3: From H, select items that can be “added in the near term”] Divide H-scope into phases of within 30 days, within 90 days, and within one year after Go-Live, and turn it into a roadmap. Agree that this is “doing it later,” not “abandoning it” [Step 4: Formally exclude N from the project] Pass a formal exclusion resolution through the CCB. Record it as “under consideration for Phase 2.” When excluding N, record the reason it was classified as N (to prevent later arguments that “you promised”) |
3.2 Components of the Recovery Plan
| Plan Component | Content | Owner | Target Completion |
| Recovery Declaration Document | A one-page “fresh-start declaration” summarizing what is happening, what will change and how, and who holds decision-making authority. Distributed to all project members | Recovery Leader | Day 3 |
| Revised Scope List (Revised WBS) | New scope list after MHN classification. Clarifies what is in scope for Go-Live versus deferred to later phases. Attach a before/after comparison to visualize “what was cut” | Recovery Leader + PMO | Week 1 |
| Realistic Master Schedule (Recovery Schedule) | A schedule that is “genuinely achievable,” incorporating remaining work volume, team productivity, and buffer. Premised on weekly updates | PMO | Week 1–2 |
| Quality Recovery Plan | Includes the plan for resolving Blocker bugs, the plan for re-running tests, and the schedule for data migration rehearsals | QA Lead + Test Lead | Week 1–2 |
| Communication Plan | Designs who is told what, when, and in what format — designed separately for executive leadership, the client, the team, and suppliers | Recovery Leader | Day 5 |
| Updated Risk Register | Risk register updated with new risks identified after the recovery began. Each risk lists its mitigation, owner, and trigger conditions | PMO | Week 2 |
3.3 How to Build a “Realistic Master Schedule”
The remaining schedule for a troubled project must always be built in three stages. Trying to create a detailed schedule for all remaining work from the outset will produce a schedule that lacks precision and, worse, erodes trust.
| Stage | Time Horizon | Content | Precision |
| Sprint Plan (Short-Term, Fixed) | Next two weeks | A fixed plan broken down to the task level. Updated for the following week every Monday | ±2 days approx. |
| Phase Plan (Mid-Term Outlook) | Next 1–2 months | Schedule at the milestone level. Reviewed and updated biweekly | ±1 week approx. |
| Roadmap (Long-Term Direction) | Through Go-Live and beyond | The overall flow of phases and the target Go-Live date. Reviewed monthly | ±2–4 weeks (acceptable) |
| Four Prohibitions in Schedule Design |
|
Prohibition 1: Building a schedule that assumes “everyone working at full capacity” → Realistically, effective utilization is 60–70% of the nominal figure. Account for overtime, meetings, and interruptions Prohibition 2: Building without buffer (contingency time) → In a troubled project there is no day without an unannounced problem → Build in a minimum buffer of 10–20% between milestones Prohibition 3: Parallelizing tasks without regard to dependencies → Ignoring dependency relationships such as “B can only start once A is complete” when building the plan → produces a schedule that will not complete on time even if everyone starts as planned Prohibition 4: Building the schedule on “hope” → “We can do it if we try hard” or “we’ll make it if things go well” is not a schedule → Build it starting from an “actuals-based projected completion date,” calculated by dividing remaining effort by actual measured productivity |
Chapter 4: Team Reset and Rebuilding the Structure
4.1 Criteria for “Who to Remove” and “Who to Keep”
One of the hardest decisions in a recovery is “who to remove.” If the cause of the crisis lies in the ability or attitude of team members, the project cannot be turned around without replacing people. However, if this decision is made based on emotion or personal relationships, trust in the recovery leader is lost. Judge based on whether the person is “able to function,” using the following criteria.
| Judgment Category | Specific Criteria | Outcome of the Judgment |
| Skill Shortfall for the Role | A consultant who cannot design business processes; a PM who cannot manage a schedule; a leader who cannot make decisions | Replace with someone skilled, or add support. Continuing without the required skill becomes an obstacle to recovery |
| Insufficient Commitment to the Role | Not attending meetings, unable to produce materials, effectively unable to work due to other duties, consistently delaying on things they said “I’ll do it” | Either secure commitment or replace. Returning the person to their primary duties is also the fair choice for them |
| Negative Impact on the Team | Only criticizing or rejecting in meetings with no constructive proposals; behavior that lowers other members’ morale; publicly undermining the PM’s or leader’s decisions | A direct one-on-one conversation first — if there is no improvement, move to a different role. Left unaddressed, it spreads through the whole team |
| Role Mismatch | “Assigned role” and “actual capability” do not match (e.g., someone in charge of quality management who cannot make quality judgments) | Reassign to a more suitable role. Present it as “specialization,” not as a “demotion” |
| Principles for Communicating When “Removing Someone” |
|
1. Decide fast, deliver gently Once you reach the conclusion to “remove” someone, the longer you delay, the more damage accrues to both the individual and the team Speak with the person one-on-one within 48 hours of the decision 2. Frame it as a “mismatch between role and capability,” not “you are at fault” “The skill set this project needs and your current expertise are not aligned” — present it as a fact of mismatch, not as criticism 3. Present the next step Be clear about whether you are asking them to “take on a different role in this project” or to “step away for now” Ending the conversation with an ambiguous “I’ll think about it” wears the person down further 4. Bring in a third party In difficult cases, have HR or a senior sponsor sit in Delivering this alone as the recovery leader carries risk |
4.2 The “Four Roles” Needed for Recovery
| Role | What They Do | Required Skills | Common Staffing Mistakes |
| Recovery Leader | Directs the overall project recovery. Bridges executive leadership, the client, and the team. Decides and declares “what to give up” | Decisiveness and interpersonal communication. Grasp of the overall project picture. The ability to say No | The mistake of assigning someone strong in technology. A conflict of interest in assigning the head of the department where the problem occurred |
| PMO (Planning and Control) | Keeps the current state of schedule, scope, and issues always up to date and visible through numbers. Reports weekly on whether things are on plan | Schedule and issue management tools. The meticulousness to record and track data accurately | The mistake of having a consultant hold this role concurrently (they go easy on their own scope). Operating with a bias toward either the client or the vendor |
| Functional Architect (Business Design) | Can design and judge the To-Be business process. Can propose not just what users “want” but what they “should” do | Hands-on experience in the relevant business domain. Deep understanding of SAP standard processes. The communication skill to “persuade toward standardization” | Assigning a junior consultant with no business experience as the functional architect. Treating “someone who interviews and records” as a designer |
| QA Lead (Quality Management) | Leads test planning, execution, and quality judgment. Defines “passing criteria” and judges “pass” or “fail.” The quality-side owner of the Go-Live decision | Experience in test design (scenario design). The ability to prioritize bugs. The standing and capability to make a Go/No-Go call | The mistake of treating a tester as the QA lead. “Someone who runs a lot of tests” and “someone who judges quality” are different roles |
Chapter 5: Communication Design During Recovery
5.1 Communication Design by Stakeholder
What drains a troubled project the most is the absence of a design for “who is told what and how,” leaving the team stuck in constant ad hoc reporting and fielding questions. Designing a communication plan in the first week of recovery dramatically lowers the cost of reporting.
| Stakeholder | What to Communicate | Frequency/Format | Points of Caution |
| Executive Leadership (Sponsor/C-Suite) | Current numbers (completion rate, quality, remaining issue count). This week’s decision items (requiring a YES or NO). Executive decisions needed for recovery | 30 minutes weekly (in person or video). A 1–2 page summary document | Design it as a “decision-making forum,” not a status report. Convey the severity of the problem without sugarcoating. “Things are generally on track” is the worst possible report |
| Client-Side PMO (Project Owner) | Detailed progress, issues, and changes. A weekly list of items requiring judgment. Current risk status and countermeasures | 60 minutes weekly (regular meeting). A one-page progress report | Rather than listing every issue, narrow it to “items that need a decision.” Convey the depth and breadth of issues accurately |
| Business Unit Heads (Key Users’ Managers) | Business impact (what will change, what won’t). Requests for cooperation from key users. Post-Go-Live business changes | 30 minutes biweekly. A business-impact summary | Frame it as “a business story,” not “a system story.” Avoid IT jargon and speak in business terms |
| All Project Team Members | This week’s sprint plan and target achievements. Latest status of issues and countermeasures. Good news (even small wins matter) | 30 minutes weekly (all-hands), plus a 15-minute daily standup | Do not create “information gaps” — everyone should hold the same information. During a crisis, “I hadn’t heard that” is a seed of distrust |
| Consultant/SI Lead | Detailed issues and technical decisions needed. Obstacles to achieving the schedule | 60 minutes weekly (technical regular meeting), an issue tracker | The recovery leader attends directly and does clear fact-checking, not allowing an “excuses mode” |
5.2 Meeting Design: Making It a “Decision Meeting,” Not a “Status Meeting”
| Meeting Design Principles During Recovery |
|
Principle 1: Label every agenda item as either a “decision item” or “information sharing” Agenda items labeled “decision” must reach a YES/NO at that meeting “Carrying it over to next time” is prohibited in principle (during recovery, carrying items over prolongs the crisis) Principle 2: Limit the number of agenda items (a maximum of three per meeting) The more items you cram in, the less gets decided Decide the three most important items, and report the rest in writing Principle 3: Use the last 5 minutes of every meeting to confirm action items Read out and confirm, in front of everyone, “who, does what, by when” Distribute the minutes (agenda, decisions, action items) within 30 minutes of the meeting ending Principle 4: If “there are too many issues and not enough time,” set up a separate issue-resolution working group The regular meeting must not become a venue for surfacing issues Resolve issues in a separate working group; the regular meeting reports only what has been resolved and what remains unresolved Principle 5: Do not gather only “people without decision authority” A meeting attended only by listeners is an information exchange, not a decision meeting Do not hold a meeting as a “decision meeting” unless at least one decision-maker is present |
Chapter 6: Managing the Recovery Execution Phase
6.1 Execution Management Through Two-Week Sprints
During the recovery phase, short-cycle PDCA through “two-week sprints” is more effective than long-term planning. Cycling through plan → execute → review → next-sprint planning every two weeks enables the plan to adapt to reality.
| Sprint Schedule | Activity | Participants | Output |
| Sprint Day 1 (Monday) | Sprint planning meeting. Everyone commits together to the tasks to be achieved over the two weeks | All team members, PMO, leader | Sprint backlog (task list, owners, completion criteria) |
| Sprint Days 1–9 (Monday through the following Friday) | Daily standup (15 minutes): what to do today, what was done yesterday, obstacles. Obstacles are escalated the same day | All team members | Daily progress log, obstacle resolution log |
| Sprint Day 10 (following Friday) | Sprint review: demo and confirm completed tasks. Sprint retro: a 30-minute look back at what went well and what to improve | All team members, client-side key users | Sprint completion report, improvement items for the next sprint |
6.2 Quality Management During Recovery
Principles of Test Prioritization
Recovery starts from the reality that “there is no time to re-run every test.” Test prioritization achieves the maximum quality assurance possible within limited time.
| Test Priority | Scope | Criteria | Approximate Time Allocation |
| P0 (Blocker — Mandatory) | Post-fix testing of bugs blocking Go-Live. Scenarios tied to legal/compliance requirements. Core scenarios for financial calculations, inventory processing, and order management | Go-Live is not possible unless this testing passes | 50% of remaining time |
| P1 (Mandatory) | Major business scenarios (the top 20 most-used transactions). Monthly/annual processing (batch, closing processes). External system integrations (EDI, banking, e-commerce) | If this testing fails, there is a risk of business disruption | 30% of remaining time |
| P2 (Important) | Sub-scenarios and exception handling. Confirmation of report/form output. Performance testing | If it fails, operations are hindered but a manual workaround exists | 15% of remaining time |
| P3 (Deferred) | Low-frequency scenarios, future features. “Nice to have”-level checks | Can be addressed sequentially after Go-Live. Managed as a post-Go-Live backlog | 5% of remaining time (deferred as needed) |
6.3 Go-Live Decision Criteria (Go/No-Go Checklist)
The Go-Live decision should be judged not by “is everything ready” but by “does it meet the minimum conditions needed to go live and keep the business running.” This is especially true after recovering from a troubled state — without agreeing on these criteria in advance, Go-Live will be repeatedly postponed by emotional arguments of “I’m still worried.”
| Go / No-Go Decision Checklist | Priority | |
| ☐ | All P0 tests (Blocker — mandatory) have passed, and there are zero unresolved Blocker bugs | H |
| ☐ | The pass rate for P1 tests (major business scenarios) is 90% or higher | H |
| ☐ | A data migration rehearsal has been conducted with at least 50% of production data volume, and the migration error rate and time required are within acceptable limits | H |
| ☐ | Performance testing in the production environment (assuming peak business load) is complete, and response times meet the passing criteria | H |
| ☐ | Key user training is complete, and it has been confirmed that all major business-side staff can operate SAP independently | H |
| ☐ | The post-Go-Live operating structure (help desk, incident response team, vendor support) is ready to operate | H |
| ☐ | Go-Live rollback (fallback) procedures are documented, and the criteria for a rollback decision (who decides, based on what) have been agreed | H |
| ☐ | Notification to external parties affected by Go-Live — regulators, customers, suppliers, etc. — is complete | M |
| ☐ | The support plan (additional staffing, extended coverage) for the two-week post-Go-Live “stabilization period” is finalized | M |
Chapter 7: Preventing Recurrence — Governance Design to Avoid Repeating the Same Crisis
7.1 Why Projects “Relapse” After Recovery
Not a few projects experience a recurrence of problems within 3–6 months after recovering from a crisis. This happens because the relief of “we recovered it” leaves the root cause only partially addressed.
| Relapse Pattern | Why It Happens | Prevention |
| Recurrence of Scope Creep | After recovery, an atmosphere of “we can’t say no to additional requests” sets in, and previously cut scope creeps back in | Keep the CCB permanently functioning, and always present the impact estimate of added scope before deciding |
| Invisible Progression of Quality Decay | The habit of skipping testing takes hold for post-Go-Live feature additions and configuration changes | Make test-completion evidence mandatory for every production change. Implement automated regression testing early |
| Retreat of Decision-Makers | Executives who were strongly engaged during recovery step back once things stabilize, and decision-making stalls | Continue quarterly project reviews with executive participation. Design a permanent authority-delegation structure |
| Continued Dependence on Consultants | Even after recovery, the organization remains unable to make decisions without consultants | Run a knowledge-transfer plan to internal staff in parallel from the very start of the recovery |
| Root Cause Left Unresolved | The symptoms of the crisis disappear, but the root cause (business design quality, decision-making structure) remains and recurs on a different project | Conduct a Root Cause Analysis and follow through to organizational countermeasures (hiring, training, procurement processes) |
7.2 Permanent Governance Design for Healthy Project Management
| The “Baseline” Checklist for a Project That Won’t Catch Fire |
|
If all of the following function as “the baseline,” the crisis can be prevented: Decision-Making Structure ✓ A DOA (delegation of authority table) is documented, making clear who decides based on the level of the decision ✓ A Change Control Board (CCB) operates at least monthly, and every scope change passes through it ✓ Escalation paths and deadlines (e.g., a response within 48 hours) are defined Scope and Plan Management ✓ The WBS/feature list is maintained as a single, current source of truth that everyone references ✓ The two-week sprint cycle of plan → execute → review runs without interruption ✓ The reason “why it changed” is recorded whenever the plan changes Quality Assurance ✓ A definition of test completion (Definition of Done) is agreed, with no “sort of done” ✓ A procedure requiring test evidence attached to every production change is strictly enforced ✓ Bug counts, resolution rate, and remaining test counts are tracked monthly Communication ✓ The weekly all-hands (decision meeting) and biweekly executive report continue ✓ A culture exists where bad news travels fast (delays are not hidden) ✓ Psychological safety is maintained so that whoever discovers a problem finds it “easy to report” |
7.3 Designing the Post-Project Post Mortem
The Post Mortem (post-implementation review) conducted after project completion (confirmed stable operation after Go-Live) is the only opportunity to leave the organization with the lessons of a troubled project. Design it not as a “hunt for the guilty,” but as “learning to keep the next project from catching fire.”
| Post Mortem Components | Content | Facilitator | Participants |
| Timeline Reconstruction | An objective, chronological account of the major events from project start to completion (no judgment) | Ideally an external, neutral third party | All key members |
| What Went Well (Keep) | Practices, throughout the crisis recovery and the project overall, that the team agrees to “do again” | — | All key members |
| Areas for Improvement (Problem) | Causes of the crisis, difficulties during recovery, points to change next time | — | All key members |
| Concrete Improvement Actions (Action) | Specific process changes, tooling changes, and talent-development plans to implement on the next project | Project Sponsor (owns execution commitment) | Executive leadership, PMO, leaders |
| Knowledge Documentation | Compile the background of design decisions and the lessons from failures into the internal wiki and the project completion report | PMO | — |
End
Have a question about this article?
Ask the author directly — no sales pitch, just an answer.