Start with these five checks.
- Verify severity, incident and recurrence definitions.
- Normalize by traffic, services and changes.
- Link incidents to deployments and dependencies.
- Measure detection, mitigation and recovery separately.
- Find repeated causes and incomplete follow-up actions.
Ask for only what can change the answer.
- IncidentsStart, severity, impact and recovery
- DeploymentsChange, service and rollback
- ObservabilityAlert, detection and signal quality
- DependenciesService graph and failure propagation
- ActionsMitigation and post-incident follow-up
- ExposureTraffic, users and production changes
Test one reversible move.
Choose one recurring incident class with a supported technical path. Add one reversible control—targeted test, canary, circuit breaker or alert change—to the affected service. Compare recurrence, detection, recovery and delivery delay; remove the control if noise or lead-time cost exceeds its prevention value.
A decision your team can use.
- 01An exposure-normalized trend
- 02A change and dependency map
- 03Recurring failure priorities
- 04A reversible reliability control
Common questions.
Does a higher incident count mean reliability worsened?
Not always. Better detection, more services or more traffic can raise counts; normalize exposure and severity.
Should more approvals be added?
Only if evidence shows the approval would catch the failure. Broad gates can slow recovery without preventing incidents.
What should a first test target?
A frequent, measurable and reversible failure path with clear operational ownership.