Key Concepts
- Correction of Error (CoE): A 4-6 page document analyzing a technical issue with significant customer impact.
- Five Whys: A root cause analysis technique involving asking "why" repeatedly (typically five times) to drill down to the fundamental cause of a problem.
- Action Items: Specific tasks identified to prevent recurrence of the issue, with assigned severity, due dates, and ownership.
- Incident Response: The process of detecting, diagnosing, and mitigating a technical issue.
- Change Control Process: A documented procedure for managing and tracking configuration changes to prevent unintended consequences.
- Anomaly Detection: Using metrics and alarms to automatically identify deviations from normal system behavior.
CoE Document Structure and Content
1. Title
- Briefly describes the issue and its impact (less than 20 words).
- Example: "Incorrect configuration change causes 62-minute ordering downtime and 20,000 lost orders."
- Includes the mistake, impact duration, and customer impact.
2. Summary
- Provides context for the event, how it was discovered, a brief description of the issue, and what was/is being done.
- Should give a firm grasp of the issue without going into excessive detail.
- Includes background information about the affected service (e.g., Order Handling Service - OHS).
- Details the timeline of discovery, investigation, root cause identification, mitigation, and recovery.
- Mentions the overall impact, such as downtime duration and lost customer order opportunities.
- Briefly explains the cause of the issue, such as a configuration change error.
3. Metrics/Graphs
- Visually represents the issue, its lead-up, impact, and recovery.
- Includes at least one or two graphs, focusing on relevant data.
- Uses vertical lines with short descriptions to indicate key events on the timeline.
- Example: A graph showing order volume over time, with annotations for configuration change, issue discovery, and rollback.
- Highlights the window of impact.
4. Impact
- Describes the total customer impact and the duration of the event.
- Quantifies the impact, such as the number of lost customer order opportunities (e.g., 20,000).
- Details the customer experience during the event (e.g., "item unavailable" message).
5. Incident Response
- Covers how the issue was detected and how the team responded.
- Answers pre-established questions, such as:
- How was the issue detected?
- How long did it take to detect the issue? How could we reduce the time to detect in half?
- How did you reach the point where you knew how to mitigate the impact?
- How could time to mitigation be improved?
- How did you confirm that your system was fully recovered?
- Identifies gaps in monitoring and response procedures.
- Links directly to action items to address these gaps.
6. Timeline
- A chronological list of events related to the issue.
- Requires a collaboration tool during high-severity events to document what is happening and when.
- Provides precise timing of specific events, such as configuration changes, issue discovery, team engagement, investigation, root cause identification, and rollback.
7. Five Whys
- The most important section, drilling down to identify the root cause.
- Starts with the impact and asks "why" repeatedly to uncover underlying causes.
- Can involve parallel branches of questions and answers.
- Example:
- Why did order volumes drop to zero?
- Why was the configuration change applied to all orders?
- Why did a star in the product ID cause all products to be disabled?
- Why did the evaluation logic apply the results of a single entry to all products?
- Why was there a bug in the regular expression and why was this only discovered now?
- Identifies specific bugs, logic flaws, and missing tests.
8. Action Items
- Defines specific tasks to prevent recurrence of the issue.
- Each action item is labeled with:
- Title
- Description
- Severity (High, Medium, Low)
- Status (Pending, Done)
- Due Date (and Completion Date if done)
- Action items directly address gaps identified in the Incident Response and Five Whys sections.
- Examples:
- Add anomaly detection alarms to order volume metrics.
- Add metric and anomaly detection alarm for item unavailable response codes.
- Add documented change control process for configuration updates.
- Fix evaluation logic to prevent content of one entry having adverse effects on other entries.
- Add unit tests for the evaluation logic.
CoE Process Beyond the Document
1. Team Review
- Schedule an internal team review to collect feedback on the CoE content and improve it iteratively.
2. Leadership Review
- Review the document with senior management (Senior Manager, Director, VP, or CTO).
- Ensure the document is polished, accurate, and easily understood by outsiders.
- Anticipate questions and practice responding to them.
3. Publication
- Host the CoE in a company-accessible location.
- Send an email to tech teams across the company to inform them of the CoE and any key insights.
Important Considerations
- CoEs are not punishment: They are a constructive process for root cause analysis and prevention.
- Avoid blame: Never refer to actors by name; use abstract references (e.g., "an engineer from XYZ team").
- Use third person: Avoid "I" or "we" in the document.
- Mitigation first: During an event, prioritize mitigating customer impact over writing the CoE.
- Roadmap impact: CoE action items will affect product development roadmaps and may require de-prioritization of other projects.
- Prioritize quality: The best CoE is one that is never written, meaning prioritize mechanisms to ensure high-quality software.
Benefits of Adopting CoE Culture
1. Root Cause Analysis and Future Issue Prevention
- Increased confidence that the same issue will not happen again.
- Creates a culture change focused on proactive failure analysis and edge case anticipation.
2. Record of Past Issues
- Provides a comprehensive record of past issues that can be referred to over and over again.
- Requires keeping CoEs in a publicly accessible and searchable location.
3. Company-Wide Improvement
- Leads to key insights that may not have been discovered otherwise.
- Insights are only valuable if shared, read, and responded to by other teams.
- May necessitate company-wide campaigns to address issues.
Synthesis/Conclusion
The Correction of Error (CoE) process, as practiced at companies like Amazon, is a structured approach to documenting, analyzing, and learning from technical failures. By creating a detailed CoE document, teams can identify root causes, implement preventative measures, and share valuable insights across the organization. This culture of embracing failure as a learning opportunity leads to more resilient systems, fewer repeat errors, and continuous improvement in software quality and operational practices. The key is to treat CoEs as a constructive process, focusing on prevention rather than blame, and to prioritize action items to ensure that lessons learned are translated into tangible improvements.
AI summaries can miss context or contain errors. Check important details against the original video.