
Reliability engineering
GitHub’s August 17 outage: how retries delayed recovery for hours
One unobserved limit prevented autoscaling, while retries turned latency into fresh load. The incident shows why spare servers are not enough without retry controls and a tested way to work through a platform failure.