Google Cloud outage shows it’s still hard to understand hyperscalers’ real resilience regimes
off prem
Single datacenter and just three services taken down by ‘upstream’ power problem, while the rest of a zone and region kept humming
Google Cloud last week experienced an outage that analysts say demonstrates that not all promises of cloudy resilience are created equal.
Google’s incident report explained that three services – the VMware Engine (GCVE), NetApp Volumes, and Bare Metal Solutions (BMS) – experienced a 15-hour outage due to a cooling failure in its europe-west4-a zone.
The report includes the following detail: “The datacenter serving europe-west4-a for GCVE, BMS, and NetApp has experienced a power failure, which subsequently caused a cooling failure.”
The important detail there is that Google uses a discrete datacenter for those three services.
Another notable element of the incident report is the admission that “An electrical fault occurred on the utility grid upstream of the datacenter, disrupting the electrical distribution gear and cooling equipment.”
...
Copyright of this story solely belongs to theregister.com. To see the full text click HERE