Automated Incident Detection & Self-Healing Telemetry Engine
Detecting production degradations with CloudWatch Composite Alarms, X-Ray distributed tracing, and executing automated Lambda self-healing.
1. Business Problem & Context
An on-call engineering team suffered from severe alert fatigue: hundreds of individual CPU alerts woke engineers up at night for transient 10-second spikes that self-resolved. Conversely, when a complex cascading deadlock occurred between microservices, engineers spent 35 minutes manually combing through logs to locate the root cause.
2. Requirements & Constraints
- Eliminate False Alarms: Alarms must only fire if both latency AND error rates breach SLOs simultaneously.
- Distributed Trace Context: Pinpoint exact bottleneck microservices within seconds using AWS X-Ray.
- Automated Self-Healing: Automatically flush degraded cache shards and recycle unhealthy tasks before human escalation.
3. Architecture Overview & Data Flow
Interactive Architecture Diagram (Use controls to zoom & pan)
4. AWS Services Used & Rationales
AWS Services Architecture Rationale
Concrete reasons why these specific services were chosen over alternatives
| Service | Category | Architectural Rationale ("Why this service?") |
|---|---|---|
| CloudWatch Composite Alarms | Observability | Evaluates boolean logic (e.g. ALARM(HighLatency) AND ALARM(High5XX)) to eliminate single-metric false alarms. |
| AWS X-Ray | Observability | Tracks HTTP request paths across API Gateway, Lambda, ECS, and DynamoDB, isolating slow SQL queries. |
| AWS Lambda (Self-Healing) | Compute | Executes automated runbooks (e.g. cycling unhealthy container tasks) within seconds. |
5. Key Design Trade-offs
Architecture Decision & Trade-Off Matrix
Evaluating alternative approaches under real-world constraints
Individual Single-Metric CPU Alarms
- + Simple setup
- − Extreme alert fatigue
- − Pages engineers for harmless transient spikes
- − No autonomous fix
Composite Alarms + Autonomous Remediation (Chosen)
✓ Chosen Design- + Zero noise for benign spikes
- + Autonomous recovery in 45s
- + Clear trace root cause in X-Ray
- − Requires disciplined alarm rule engineering
6. Implementation Highlights
IaC Recipe Composite Alarm CloudFormation Expression
AlarmRule: >
ALARM("prod-api-p99-latency-breached") AND
ALARM("prod-api-5xx-error-rate-breached") AND NOT
ALARM("prod-scheduled-backup-in-progress") 7. Results & Key Metrics
- MTTR: Dropped from 35 minutes to 45 seconds for 92% of known incident types.
- On-Call Pages: Decreased from 140 pages/week to 6 actionable pages/week.
8. Key Architectural Takeaways
SRE Principle: Never page a human engineer for an issue that a Lambda function can diagnose and remediate automatically in 30 seconds.