Archived Content

The following content is from an older version of this website, and may not display correctly.

When clicking on a button on Facebook, it is easy to take for granted that a series of complex infrastructure decisions had to be made for that click to provide the expected result. As the features the company adds to the site get increasingly sophisticated, infrastructure demands grow in complexity as well.

 

For the popular “Look Back” video feature Facebook added to mark 10th anniversary of the social network's existence to work, the company's engineers had to look at everything from network bandwidth, storage and compute capacity to data center power consumption scheme at the facilities level.

 

The feature, which combined photos and other posts that generated the most activity on each user's account over the time span of the account's existence into a short video with background music, ended up being a big hit. Hundreds of millions of videos were created and shared by users in the first few days after it was launched in early February.

 

In a blog post describing the infrastructure challenge behind Look Back, Facebook engineers Alexey Spiridonov and Krish Bandary explain that the data center power concern stemmed primarily from the amount of interconnect bandwidth the feature would need.

 

Look Back and the infrastructure to support it had to be built in less than one month. The infrastructure challenge was exacerbated by the looming launch of Facebook's Paper app – the launch was scheduled five days before the anniversary – and the 2014 Winter Olympics was going to start one week later, which would flood the site with Olympics-related content.

 

As a rough starting point, the team assumed people would create and share 25m videos on the day of the anniversary. They calculated that supporting this volume would require additional bandwidth of about 62Gbps on average and 187Gbps at peak usage.

 

“That’s more than 20,000 average US high-speed internet connections, fully saturated all the time,” Spiridonov and Bandary wrote. “Our network engineers had some work to do.”

 

Storing of the videos would require about 25 petabytes. “We couldn’t make the videos without having this space, so the disks had to be procured within a week or two at the most.”

 

To prevent the new feature from disrupting operations of the main site, the team chose to use dedicated storage servers to store the videos. They placed these dedicated servers on both US coasts, treading them as different types of video than regular video files users usually upload.

 

Simply getting the storage capacity and stuffing as much disk into each data center rack as possible would not be enough, however. The traffic coming in and out of the drives would quickly saturate top-of-rack switches.

 

The team had to deploy more hardware and provision more racks to spread the read/write workload.

 

The necessary amount of compute muscle was still undetermined at this point. They had access to lots of servers that were already provisioned but Facebook did not run all servers at full capacity all the time to keep energy consumption in check.

 

The solution was to track power usage in real time as the jobs were running and build software capabilities to allow the infrastructure to “slow down” when necessary. This would ensure the rest of Facebook would not be disrupted by Look Back gobbling up too much capacity.

 

By 3 February (one day before launch) all videos were ready, driving used storage capacity from 0 bytes to 11 petabytes over the previous six days. At its fastest, the infrastructure cranked out 9m videos per hour, in the end delivering more than 720 rendered videos.

 

More than 200m people watched their videos in the first two days and more than half shared them with others. “And all of this came without a single power breaker tripping, which made our data center teams very happy,” Spiridonov and Bandary wrote.