An incident that three separate fixes had failed to close
Key takeaway
Three fixes failed because nobody owned the failure. The fourth one worked because someone finally did.
- Situation
- A live SaaS product with paying customers. A failure recurring under load, closed and reopened three times over several months, consuming roadmap capacity every time it returned.
- What Kendoo changed
- Ownership was assigned with the authority to act on it. Instrumentation went in before any further fix, and a decision rule was agreed for fix now, isolate, or schedule.
- Outcome
- The cause became visible, and the argument about how to fix it was settled on evidence. Recurrence stopped being the default expectation.
Read the detailed pattern
- Symptoms
- A failure recurring under load, closed and reopened three times over several months.
- Risk
- Customer-visible, and consuming roadmap capacity every time it returned.
- Constraints
- No pause in feature delivery was acceptable to the business.
- Diagnosis
- No single owner for the failing path, and no observability at the point where it broke. Each fix had addressed the most recent symptom.