High Availability
- Definition
- The system stays operational and accessible with minimal downtime, even when components fail.
- Measured as a percentage of uptime (99.9% = ~8.8 hours down per year; 99.99% = ~52 minutes).
- How AWS delivers it
- Deploy across multiple Availability Zones — the foundational pattern.
- load balancers with health checks route around failed instances.
- AWS Auto Scaling replaces unhealthy instances automatically.
- Amazon RDS Multi-AZ, Amazon S3 multi-AZ durability, Amazon DynamoDB three-AZ replication.
- Distinguish from neighbors
- Fault tolerance — the system keeps running with zero interruption during failure. Stricter and more expensive than HA, which allows brief failover.
- Disaster recovery — restoring service after a large-scale outage, usually across Regions. Measured by RTO and RPO.
- Elasticity — matching capacity to demand, not to failure.
- Agility — speed of building and changing, not uptime.