All challenges

Multi-region cost blowout

expert
Scenario & brief

A team went active-active across two regions for a service whose SLA is only 99.95% availability — a level a single, well-built region hits comfortably. The result is double the infrastructure cost and a budget blown wide open.

Targets: 99.95% availability, p99 ≤ 140 ms, critical durability, under $1,200/mo.

Bring it back to a single region done right — two AZs, Multi-AZ database, a cache, Graviton on a Savings Plan — and only pay the multi-region premium if the SLA actually demanded it (here it doesn't).

6,000 rps peakp99 ≤ 140ms99.95% availdurability: criticalbudget $1,200/mo

CloudFront

Networking

System health

Erupting · SLA breach

28

/ 100

Score

SLA not met yet

Monthly cost

$4,100

Budget $1,200/mo · $2,900 over

Metrics

Capacity57
Availability100
Durability25
Cost efficiency0

Requirements

  • Peak capacity 21600 rps compute · 3400 rps db (need ≥ 6000 rps)
  • p99 latency ~114 ms (need ≤ 140 ms)
  • Availability 100.00% (need ≥ 99.95%)
  • Durability at risk (need redundancy + backups)
  • Budget $4100/mo (need ≤ $1200/mo)

Advisor

  • The database is saturated at peak — add a cache to shed read load, or scale it up.
  • The database has no Multi-AZ standby or replica — a failure risks data loss. Enable Multi-AZ and backups.

Discussion

Sign in to join the discussion.

No comments yet. Be the first to start the discussion.

For learning purposes only. Costs and capacities are illustrative, not live AWS prices. Not affiliated with or endorsed by Amazon Web Services.