Expert methodology for identifying, assessing, and mitigating technical and operational risks including security, incidents, compliance, and disaster recovery.
This skill provides a comprehensive framework for managing technical risk and building resilient systems. Use it to conduct risk assessments, plan incident response, achieve compliance certifications, and ensure business continuity.
Trigger this skill when you need to:
Follow this systematic approach to risk and resilience:
Categorize Risks
Technical Risks:
Security Risks:
Operational Risks:
Compliance Risks:
Business Risks:
Conduct Risk Assessment
Use references/frameworks/risk-assessment-matrix.md to systematically identify and score risks.
For each risk:
Create Risk Matrix
High Impact (Critical)
|
| [Low Priority] | [Medium Priority] | [HIGH PRIORITY]
| Low Prob | Medium Prob | High Prob
| High Impact | High Impact | High Impact
(5)|___________________|_____________________|____________________
| | |
| [Low Priority] | [Medium Priority] | [High Priority]
| Low Prob | Medium Prob | High Prob
| Medium Impact | Medium Impact | Medium Impact
(3)|___________________|_____________________|____________________
| | |
| [Very Low] | [Low Priority] | [Medium Priority]
| Low Prob | Medium Prob | High Prob
| Low Impact | Low Impact | Low Impact
Low (1)|___________________|_____________________|____________________
Impact
Low (1) Medium (3) High (5)
Probability
Priority Levels
Critical (Priority 1): Address immediately
High (Priority 2): Address within 30 days
Medium (Priority 3): Address within 90 days
Low (Priority 4): Monitor and address opportunistically
Use references/templates/risk-register.md to maintain ongoing risk inventory.
Implement security controls across multiple layers:
Preventive Controls (Stop threats before they happen)
Detective Controls (Identify threats quickly)
Responsive Controls (React to incidents)
Use references/frameworks/security-controls-framework.md for comprehensive checklist.
SOC 2 Type II Certification
Timeline: 12-18 months (9 months preparation + 3-6 months audit period + 3 months report)
Use references/templates/soc2-roadmap.md for detailed plan:
Phases:
Key Areas:
ISO 27001 Certification
Timeline: 12-24 months
Use references/templates/iso27001-roadmap.md:
Phases:
Key Requirements:
GDPR / CCPA Compliance
Timeline: 6-12 months
Use references/templates/data-privacy-compliance.md:
Key Areas:
Create structured incident response capability:
1. Preparation
2. Detection
3. Triage
4. Investigation
5. Containment
6. Resolution
7. Post-Mortem
Use references/templates/incident-response-playbook.md for detailed procedures.
| Level | Definition | Response Time | Escalation |
|---|---|---|---|
| P0 - Critical | Complete service outage, data breach, security incident | Immediate | All-hands, exec team notified |
| P1 - High | Major feature broken, significant degradation | <15 minutes | On-call team, manager notified |
| P2 - Medium | Partial functionality impaired, workaround exists | <2 hours | On-call team |
| P3 - Low | Minor issue, minimal customer impact | Next business day | Normal ticket queue |
Structure:
Tools:
Health Metrics:
Use references/frameworks/on-call-framework.md for detailed guidance.
Ensure critical business functions can continue during disruptions:
Recovery Objectives:
RTO (Recovery Time Objective): How long can we be down?
RPO (Recovery Point Objective): How much data can we lose?
Disaster Scenarios:
For each scenario:
Use references/templates/disaster-recovery-plan.md for comprehensive planning.
Regular Drills:
Game Days:
Documentation:
Build systems that gracefully handle failures:
1. Circuit Breakers
2. Retry with Exponential Backoff
3. Timeout and Bulkheads
4. Graceful Degradation
5. Rate Limiting and Load Shedding
Use references/frameworks/resilience-patterns.md for implementation guidance.
Define reliability targets:
SLI (Service Level Indicator): What we measure
SLO (Service Level Objective): Our target
SLA (Service Level Agreement): Promise to customers
Error Budget:
Use references/templates/slo-definition.md for framework.
Frameworks (references/frameworks/):
risk-assessment-matrix.md - Systematic risk identification and scoringsecurity-controls-framework.md - Comprehensive security checkliston-call-framework.md - Sustainable on-call practicesresilience-patterns.md - Architecture patterns for resiliencechaos-engineering.md - Controlled failure testingTemplates (references/templates/):
risk-register.md - Ongoing risk trackingincident-response-playbook.md - Step-by-step incident proceduressoc2-roadmap.md - SOC 2 certification planiso27001-roadmap.md - ISO 27001 certification plandata-privacy-compliance.md - GDPR/CCPA compliance guidedisaster-recovery-plan.md - DR procedures and testingslo-definition.md - Service level objective frameworksecurity-audit-checklist.md - Pre-audit preparationpost-mortem-template.md - Incident analysis formatExamples (references/examples/):
Example 1: User says "We need to get SOC 2 certified for enterprise sales"
ā Load references/templates/soc2-roadmap.md
ā Conduct gap assessment against SOC 2 requirements
ā Create 12-18 month roadmap with phases
ā Identify control implementations needed
ā Estimate costs (audit fees, tools, consulting)
ā Assign ownership and timeline
ā Provide monthly checklist for evidence collection
Example 2: User says "Create incident response process for my 30-person team"
ā Load references/templates/incident-response-playbook.md
ā Define severity levels (P0-P3) with examples
ā Design on-call rotation structure
ā Create runbooks for common scenarios
ā Set up communication channels (Slack, status page)
ā Define escalation paths
ā Schedule incident response training
Example 3: User says "Conduct security risk assessment for Series B due diligence"
ā Load references/frameworks/risk-assessment-matrix.md
ā Inventory all systems and data
ā Identify risks across security, compliance, operational
ā Score by probability and impact
ā Document existing controls
ā Create risk mitigation roadmap
ā Prepare executive summary for investors
Example 4: User says "We had a major outage, help with post-mortem"
ā Load references/templates/post-mortem-template.md
ā Document incident timeline
ā Identify root cause(s)
ā Analyze what went well and poorly
ā Create blameless narrative
ā Generate action items with owners
ā Share with team and stakeholders
ā Track action item completion
Focus: Security basics, avoid catastrophic risks
Priorities:
Avoid: Over-investing in compliance certifications too early
Focus: Scalability, reliability, security hardening
Priorities:
Investment: 10-15% of engineering time on resilience
Focus: Compliance, resilience, enterprise security
Priorities:
Investment: 20-25% of engineering time on reliability/security
| Indicator | Risk | Action |
|---|---|---|
| No monitoring on production | High - can't detect issues | Immediate: Implement basic monitoring |
| No backup/DR tested in 6+ months | High - recovery may fail | Test DR procedures this quarter |
| Single person knows critical system | High - bus factor = 1 | Document and cross-train immediately |
| Increasing incident frequency | Medium-High - system degrading | Root cause analysis, resilience improvements |
| Failed security scan findings | High - vulnerable to attack | Remediate critical/high findings in 30 days |
| Compliance deadline <6 months | High - may not certify in time | Accelerate roadmap, consider consultant |
| On-call team burned out | Medium - quality and retention risk | Reduce incident load, improve tooling |
Security & Risk Update
Status: š¢ Secure and compliant
Key Metrics:
Top Risks & Mitigations:
Investment Request: $150K for penetration testing and SOC 2 audit
Incident Review - Service Outage Feb 15
What Happened: Database connection pool exhaustion caused 45-minute outage
Timeline:
What Went Well:
What We'll Improve:
Action Items: [See detailed list]
No blame - systems fail, we learn and improve.
All outputs should be:
Version: 1.0.0 Philosophy: Prevent where possible, detect quickly, respond effectively, learn continuously