GitHub has published its account of the August 17 outage that knocked out github.com and services including Issues, Pull Requests, Actions, authentication and Copilot for nearly eight hours, attributing the failure to a capacity shortfall that its automated scaling systems failed to anticipate as traffic hit a new peak.
According to the postmortem GitHub published on its blog, the outage ran from 13:28 to 21:15 UTC — 7 hours and 47 minutes — with web and API error rates reaching roughly 20% at peak and archive and raw-content downloads spiking to about 50% failure. GitHub said the root cause was not a code change or misconfiguration but "fundamentally due to insufficient storage capacity" in a critical infrastructure component in its Central US data center that failed to scale as demand surged.
A Retry Loop Made It Worse
GitHub's account describes a cascading failure pattern that will be familiar to engineers who have dealt with large-scale outages: as the initial capacity shortfall caused errors, a client-side retry loop in Copilot authentication amplified the problem rather than allowing the system to recover. Requests to the service that processes Copilot authentication tokens swelled from a typical 7,000-9,000 per second to between 70,000 and 100,000 per second, according to the company, as clients — including, GitHub said, a potential bug in Visual Studio Code — retried failed authentication attempts far more aggressively than intended.
The underlying growth pressure was severe. GitHub said its monthly commit volume nearly doubled in about four months, climbing from 1.4 billion in April 2026 to 2.9 billion by mid-August, a pace of adoption the company linked to the broader surge in AI-assisted coding tools and agentic workflows now running against its platform.
Scale is not our only challenge.
— GitHub, in its August 17 outage postmortem
GitHub's Second Major Incident This Month
GitHub acknowledged in the postmortem that this was its second significant incident in August, without detailing the earlier one, and framed the response as addressing both infrastructure headroom and operational practice rather than capacity alone. The company said its status page will continue to be the authoritative real-time source for any future incidents.
To prevent a repeat, GitHub said it has added more than 3 million CPU cores and 120 petabytes of high-speed storage capacity, and has shifted a majority of platform load — now 58% — onto Microsoft Azure infrastructure. The company also said it is redesigning systems handling large repositories so that read capacity scales linearly with the number of readers, rather than hitting the kind of hard ceiling that contributed to last Monday's collapse.
What's Next
GitHub said it is also tightening operational practices independent of the infrastructure upgrades: stronger pre-deployment testing, safer phased rollouts, improved observability and alerting, removal of shared dependencies between critical systems so a failure in one can't cascade into another, and consistent retry limits and budgets across its services to stop exactly the kind of amplifying retry storm that prolonged last week's outage. The company did not give a timeline for completing the broader architectural changes, saying only that the work is underway.