GitHub blames 8-hour outage on autoscaling fail and VS Code retry storm
Load balancers buckled after a monitoring blind spot allowed traffic to spiral
GitHub has published its account of this week's nearly eight-hour outage, tracing the developer pain to saturated load balancers, a faulty autoscaling policy, and a "latent retry bug in Visual Studio Code."
According to GitHub, problems began at 1328 UTC on August 17 and weren't fully resolved until 2115 UTC – a 7-hour, 47-minute incident that produced elevated errors across Issues, Pull Requests, APIs, Actions, and Copilot.
The immediate cause was network saturation on load balancers in the company's Central US facility, triggered when an Istio sidecar reached its concurrency limit.
Surely autoscaling would add capacity as those limits were reached? Alas, no. A misconfigured policy monitored the host service but not the sidecar's concurrency limit, allowing a cascading failure to develop. "The problem," according to GitHub, "was worsened by optimistic retry logic which overloaded internal load...
Copyright of this story solely belongs to theregister.com. To see the full text click HERE