
The August 17 outage, and the work ahead
GitHub's August 17 outage lasted 7 hours and 47 minutes after a Central US infrastructure component failed to scale with record traffic. The failure disrupted authentication, Actions, APIs, pull requests, issues, and Copilot; retrying Copilot clients then added load during recovery. Monthly commits had grown from 1.4 billion in April to 2.9 billion, exposing capacity and operational practices that had not kept pace. GitHub is introducing consistent retry limits, retry budgets, variable timeouts, stronger capacity alerts, and more isolation between critical systems. The incident shows why recovery behavior must be designed and load-tested as part of normal capacity planning.


