Systematic debugging techniques including error classification, root cause analysis (5 Whys), reproduction strategies, and error documentation...
Key principle: Fix root causes, not symptoms. Use 5 Whys to drill down.
Check error-log.md for similar historical errors before starting.
Check error-log.md for past occurrences:
# Search by error message
grep -i "connection timeout" error-log.md
# Search by component
grep "Component: StudentProgressService" error-log.md
# Search by error type
grep "Type: Integration" error-log.md
If similar error found:
If recurring error: Prioritize permanent fix over workaround.
Create minimal reproduction:
Example reproduction:
# API error
curl -X POST http://localhost:3000/api/endpoint \
-H "Content-Type: application/json" \
-d '{"field": "value"}'
# Frontend error
# 1. Navigate to /dashboard
# 2. Click "Load Data" button
# 3. Error appears in console
Validation: Can trigger error reliably before proceeding.
By type:
By severity:
Example classification:
Type: Integration (API call to external service fails)
Severity: High (dashboard doesn't load, blocks teachers)
Component: StudentProgressService.fetchExternalData()
Frequency: 30% of requests (intermittent)
See references/error-classification.md for detailed matrix.
Use systematic techniques:
Binary search (for large codebases):
Increase logging:
# Add debug logs around suspected area
logger.debug(f"Before API call: student_id={student_id}, params={params}")
response = api.fetch_data(student_id)
logger.debug(f"After API call: status={response.status}, data_len={len(response.data)}")
Use breakpoints (interactive debugging):
# Python
python -m pdb script.py
# Node.js
node inspect server.js
# Browser (Chrome DevTools)
# Set breakpoints, inspect variables, step through code
Check assumptions:
Apply 5 Whys (drill down to root cause):
1. Why did dashboard fail to load?
→ API call to external service timed out
2. Why did API call timeout?
→ Request took >30 seconds (timeout limit)
3. Why did request take >30 seconds?
→ External service response time was 45 seconds
4. Why was external service so slow?
→ Requesting too much data (1000 records instead of 10)
5. Why requesting 1000 records?
→ Missing pagination parameter in API call
Root cause: Missing pagination parameter causes over-fetching
Validation: Root cause identified (not just symptoms). Can explain why error occurred.
def test_fetch_external_data_with_pagination():
"""Test that API call includes pagination parameter"""
service = StudentProgressService()
result = service.fetchExternalData(student_id=123)
assert len(result) <= 10 # Max 10 records per page
assert result.has_next_page # Pagination metadata present
# Before (broken)
def fetchExternalData(self, student_id):
return api.get(f"/students/{student_id}/data")
# After (fixed)
def fetchExternalData(self, student_id, page_size=10):
return api.get(
f"/students/{student_id}/data",
params={"page_size": page_size}
)
# Run test (should pass)
pytest tests/test_student_progress_service.py
# Manually test reproduction steps (error should not occur)
Validation: Fix addresses root cause, all tests pass, error no longer reproduces.
Prevent recurrence:
# Unit test for specific bug
def test_pagination_parameter_included():
"""Regression test for ERR-0042: Missing pagination"""
service = StudentProgressService()
with patch('api.get') as mock_get:
service.fetchExternalData(student_id=123)
# Verify pagination param passed
mock_get.assert_called_with(
'/students/123/data',
params={'page_size': 10}
)
# Integration test for user flow
def test_dashboard_loads_with_large_datasets():
"""Ensure dashboard handles large datasets without timeout"""
client = TestClient()
response = client.get('/dashboard/student/123')
assert response.status_code == 200
assert response.elapsed.total_seconds() < 5 # No timeout
Validation: Tests added covering error scenario, tests pass consistently.
Create entry with ERR-XXXX ID:
## ERR-0042: Dashboard Timeout Due to Missing Pagination
**Date**: 2025-11-19
**Severity**: High
**Type**: Integration
**Component**: StudentProgressService.fetchExternalData()
**Reporter**: QA team (staging environment)
### Error Description
Dashboard fails to load student progress data, showing timeout error after 30 seconds.
### Root Cause
API call to external service was missing pagination parameter, causing over-fetching of 1000+ records instead of paginated 10 records. External service took 45 seconds to respond with large dataset, exceeding 30-second client timeout.
### 5 Whys Analysis
1. Dashboard timeout → API call timeout
2. API timeout → Request took >30s
3. > 30s request → External service slow (45s)
4. Service slow → Too much data (1000 records)
5. Too much data → Missing pagination parameter
**Root cause**: Missing pagination parameter in API call
### Fix Implemented
- Added `page_size` parameter (default 10) to `fetchExternalData()`
- Modified API call to include pagination params
- Response time reduced from 45s → 2s
### Files Changed
- `src/services/StudentProgressService.ts` (lines 45-52)
- `tests/StudentProgressService.test.ts` (added regression test)
### Prevention
- Unit test: Verify pagination parameter included
- Integration test: Dashboard loads within 5s with large datasets
- Added performance monitoring alert if response time >10s
### Related Errors
- Similar to ERR-0015 (pagination missing in different service)
- Pattern: Always include pagination for external API calls
See references/error-log-template.md for full template.
Validation: error-log.md updated with complete entry, ERR-XXXX ID assigned.
Final checks:
# Error no longer reproduces
# (run original reproduction steps)
# All tests pass
npm test # or pytest, cargo test, etc.
# Logs confirm fix
# Check application logs for success
# No regressions introduced
# Run full test suite
Update state.yaml (if debugging during a phase):
debug:
status: completed
error_id: ERR-0042
root_cause: Missing pagination parameter
tests_added: 2
Validation: Error resolved, documented, tests added, no regressions.
✅ Do: Use 5 Whys to drill down to root cause, fix underlying issue
Why: Symptom fixes lead to recurring errors, technical debt, unreliable system
Example (bad):
# Symptom fix: Hide error with try/catch
try:
data = api.fetch(student_id)
except TimeoutError:
return [] # Return empty, ignore error
Example (good):
# Root cause fix: Add pagination to prevent timeout
data = api.fetch(student_id, page_size=10) # Paginate to reduce data
✅ Do: Form hypothesis, change one thing, test, validate
Why: Can't identify what actually fixed it, might introduce new bugs
Example (bad workflow):
1. Change timeout from 30s to 60s
2. Also add retry logic
3. Also increase memory limit
4. Error gone... which change fixed it? Unknown.
Example (good workflow):
1. Hypothesis: Timeout too short
2. Change only timeout: 30s → 60s
3. Test: Still fails
4. Hypothesis: Data too large
5. Add pagination only
6. Test: Works! Pagination was the fix.
Why: Same error may recur, team can't learn from past issues
Required in error-log.md:
Why: Bug may be reintroduced later, no safety net
Example:
# After fixing ERR-0042, add:
def test_no_timeout_with_large_datasets():
"""Regression test for ERR-0042: Pagination prevents timeout"""
# This test will catch if pagination is removed later
Why: Can't validate fix if you can't trigger error consistently
Intermittent errors: Identify conditions (timing, data state, race conditions) to make reproducible
Example:
1. Why? Dashboard timeout
2. Why? API call slow
3. Why? Large dataset
4. Why? Missing pagination
5. Why? Developer didn't add parameter
Root cause: Missing pagination parameter
Fix: Add pagination to API call
Result: Root cause identified, not symptoms
Result: Fix validated, regression prevented
Result: Quickly narrow down error location in large codebase
Result: Team learns from past errors, patterns identified
Bad debugging:
Issue: 5 Whys leads to "developer mistake" Solution: Keep asking why - why did developer make mistake? (unclear docs, missing validation, no tests?)
Issue: Fix works but not sure why Solution: Add detailed logging before/after fix, compare behavior, validate hypothesis
Issue: Error recurs after fix Solution: Check if root cause was correctly identified (use 5 Whys again), verify tests actually prevent recurrence
Issue: Too many errors to debug Solution: Prioritize by severity (critical > high > medium > low), check for common root causes (same component, same pattern)