Core SRE behavioral principles for incident response. Guides intellectual honesty, evidence-based reasoning, and communication standards.
These principles guide how to investigate and communicate findings.
Facts (observed directly):
Hypotheses (inferred):
Don't say: "The service is having problems" Do say: "The service is returning 500 errors (5% of requests per metrics)"
14:30 - Deployment v1.2.3 completed
14:32 - Error rate increased from 0.1% to 2%
14:35 - Pod restarts began
14:40 - Alert triggered
For each hypothesis, ask:
When evidence contradicts your hypothesis:
Before concluding, consider:
**Summary** (1-2 sentences)
What happened and the root cause.
**Impact**
- Users affected
- Duration
- Services impacted
**Timeline**
Chronological events with timestamps.
**Root Cause**
Specific, technical explanation.
**Evidence**
Data supporting the root cause.
**Actions Taken**
What was done to resolve.
**Recommendations**
Prevent recurrence.
Don't bury the answer. Start with:
Then provide supporting evidence.
| Depth | Example | Usefulness |
|---|---|---|
| Surface | "Service is unhealthy" | Not useful |
| Shallow | "Pods are CrashLoopBackOff" | Describes symptom |
| Adequate | "Pods OOMKilled, memory at 512MB during peak" | Actionable |
| Excellent | "Memory leak in cart serialization, commit abc123" | Root cause |
Stop when:
Don't stop just because:
**Root Cause**: Memory leak in payment-service causing OOMKilled restarts
**Evidence**:
- Memory usage: Increased 400MB/hour (metric: container_memory_working_set_bytes)
- Events: 23 OOMKilled events in last 6 hours (get_pod_events)
- Correlation: Restarts started after deploy of commit abc123 (git_log)
- Change point: Memory trend changed at 14:32 UTC (find_change_point)
**Confidence**: High
- Memory trend and OOM events are deterministic
- Direct correlation with deployment timestamp
**Hypothesis Testing**:
- Ruled out: Traffic increase (requests stable per metrics)
- Ruled out: External dependency (no correlation)
- Confirmed: Memory growth rate constant regardless of load
**Recommendation**:
1. Immediate: Rollback to previous version
2. Follow-up: Profile memory in staging
3. Prevention: Add memory alerts at 70% threshold
**Caveat**: Did not identify the specific code causing the leak