Archived Content

The following content is from an older version of this website, and may not display correctly.

The data center that supports East Coast availability zones of Amazon's cloud computing services experienced a number of power issues last week, causing three power outages to multiple equipment racks and downtime that for some customers lasted up to seven hours on Saturday.

Amazon Web Services said the outage on Saturday was not related to the two incidents that occurred on Tuesday.

AWS said on its Web site that it would implement changes to power distribution architecture in all of its data centers to reduce the number of virtual-machine instances such incidents affect. The company did not specify how many customers were affected.

The May 8 incident took out several equipment racks in a northern Virginia data center, making a portion of Elastic Compute Cloud instances in that availability zone unreachable beginning shortly after midnight PDT.

AWS restored power to most of the affected instances by 7:20 a.m. Some of the provider's Elastic Block Storage volumes were also affected, resulting in loss of data by some users.

The cause was an electrical ground fault and short circuit in a power distribution panel. In its explanation of the incident, Amazon implied that the restoration of power was delayed by the process of identifying and correcting the ground fault and testing switch gear and power distribution systems by the data center's facility engineers.

"Restoring power without having taken this precaution would have put personnel at risk and run the risk of impacting the other hosts in this availability zone," AWS staff wrote in a public report of the incident.

In the same report, AWS plugged features of its service it says can be implemented to mitigate such failures.

They include ability to build applications across multiple availability zones and using rapid instance provisioning to replace failed instances.

"No matter what measures a provider takes, instances within a single data center will always be exposed to some probability of correlated failure, whether due to power, networking, fire or flooding."

The two outages on May 4 were caused by a problem with a UPS system in the Virginia data center.

The first outage occurred during a transfer of utility power to a new electrical substation around 2:20 a.m. PDT.

The UPS unit failed to switch to generator power during the transfer, cutting power to the equipment racks that were connected to it.

Facility engineers bypassed the faulty UPS and connected the affected racks to a generator directly.

According to Amazon, most affected instances were back up by 3:40 a.m. and almost all were recovered by around 6:30 a.m.

Less than 12 hours later, the back-up generator that was still feeding the racks lost power, which Amazon attributed to a human error.

The racks were again bumped off line, starting around 5:15 p.m. and stayed down until the generator was reset. According to Amazon, almost all instances were recovered by 7:40 p.m.

The faulty UPS was replaced on the same day.