Skip to main content

High Availability

  • Definition
    • The system stays operational and accessible with minimal downtime, even when components fail.
    • Measured as a percentage of uptime (99.9% = ~8.8 hours down per year; 99.99% = ~52 minutes).
  • How AWS delivers it
  • Distinguish from neighbors
    • Fault tolerance — the system keeps running with zero interruption during failure. Stricter and more expensive than HA, which allows brief failover.
    • Disaster recovery — restoring service after a large-scale outage, usually across Regions. Measured by RTO and RPO.
    • Elasticity — matching capacity to demand, not to failure.
    • Agility — speed of building and changing, not uptime.

Linked from