All challenges

Run batch jobs on Spot

intermediate
Scenario & brief

A nightly batch/analytics job runs on 4× m5.4xlarge On-Demand, 24×7, even though the work only runs off-hours and can safely restart if an instance disappears.

Targets (relaxed): p99 ≤ 300 ms, 99.0% availability, standard durability, under $300/mo.

Interruptible batch work is the textbook case for Spot instances, and there's no reason to run 24×7 — set a nights/weekends schedule, downsize to a right-sized Graviton class, and keep backups on the results database.

300 rps peakp99 ≤ 300ms99% availdurability: standardbudget $300/mo

Batch Workers

Compute

4instances

System health

Erupting · SLA breach

28

/ 100

Score

SLA not met yet

Monthly cost

$2,593

Budget $300/mo · $2,293 over

Metrics

Capacity100
Availability100
Durability25
Cost efficiency0

Requirements

  • Peak capacity 26000 rps compute · 3400 rps db (need ≥ 300 rps)
  • p99 latency ~74 ms (need ≤ 300 ms)
  • Availability 99.00% (need ≥ 99.00%)
  • Durability at risk (need backups)
  • Budget $2593/mo (need ≤ $300/mo)

Advisor

  • The database has no Multi-AZ standby or replica — a failure risks data loss. Enable Multi-AZ and backups.

Discussion

Sign in to join the discussion.

No comments yet. Be the first to start the discussion.

For learning purposes only. Costs and capacities are illustrative, not live AWS prices. Not affiliated with or endorsed by Amazon Web Services.