All challenges

Tune the data tier

intermediate
Scenario & brief

A service peaking at 5,000 rps is bottlenecked entirely at the data tier: it runs a burstable db.t3.medium, no cache, single-AZ, no backups. Queries are queueing and any failure loses data.

Targets: p99 ≤ 140 ms, 99.95% availability, critical durability, under $1,000/mo.

Fix the data tier: put an ElastiCache layer in front to absorb hot reads, pick a right-sized database class, make it Multi-AZ with backups, and keep the compute tier sensibly sized across two AZs.

5,000 rps peakp99 ≤ 140ms99.95% availdurability: criticalbudget $1,000/mo

App Servers

Compute

3instances

System health

Erupting · SLA breach

28

/ 100

Score

SLA not met yet

Monthly cost

$492

Budget $1,000/mo · within budget

Metrics

Capacity7
Availability30
Durability25
Cost efficiency80

Requirements

  • Peak capacity 5400 rps compute · 350 rps db (need ≥ 5000 rps)
  • p99 latency ~174 ms (need ≤ 140 ms)
  • Availability 99.00% (need ≥ 99.95%)
  • Durability at risk (need redundancy + backups)
  • Budget $492/mo (need ≤ $1000/mo)

Advisor

  • The database is saturated at peak — add a cache to shed read load, or scale it up.
  • Compute runs in a single AZ — spread across ≥2 AZs (with ≥2 instances) to meet the availability target.
  • The database has no Multi-AZ standby or replica — a failure risks data loss. Enable Multi-AZ and backups.

Discussion

Sign in to join the discussion.

No comments yet. Be the first to start the discussion.

For learning purposes only. Costs and capacities are illustrative, not live AWS prices. Not affiliated with or endorsed by Amazon Web Services.