All challenges

Handle a flash-sale spike

advanced
Scenario & brief

Most of the day this store idles, but during the flash sale it spikes to 10,000 rps for about an hour, then collapses again. The current design is a fixed fleet sized for the spike, running 24×7 — you pay peak prices around the clock, and it's still single-AZ with no CDN or cache.

Targets: p99 ≤ 130 ms, 99.95% availability, critical durability, under $1,400/mo.

The trick is to follow the curve: an autoscaling group with a low floor and a high ceiling so you only pay for the spike while it lasts, a CDN and cache to blunt the surge before it reaches the fleet and database, spread across two AZs, and Multi-AZ RDS with backups.

10,000 rps peakp99 ≤ 130ms99.95% availdurability: criticalbudget $1,400/mo

CloudFront

Networking

System health

Erupting · SLA breach

22

/ 100

Score

SLA not met yet

Monthly cost

$4,871

Budget $1,400/mo · $3,471 over

Metrics

Capacity34
Availability30
Durability25
Cost efficiency0

Requirements

  • Peak capacity 52000 rps compute · 3400 rps db (need ≥ 10000 rps)
  • p99 latency ~134 ms (need ≤ 130 ms)
  • Availability 99.00% (need ≥ 99.95%)
  • Durability at risk (need redundancy + backups)
  • Budget $4871/mo (need ≤ $1400/mo)

Advisor

  • The database is saturated at peak — add a cache to shed read load, or scale it up.
  • Compute runs in a single AZ — spread across ≥2 AZs (with ≥2 instances) to meet the availability target.
  • The database has no Multi-AZ standby or replica — a failure risks data loss. Enable Multi-AZ and backups.

Discussion

Sign in to join the discussion.

No comments yet. Be the first to start the discussion.

For learning purposes only. Costs and capacities are illustrative, not live AWS prices. Not affiliated with or endorsed by Amazon Web Services.