Archived Content

The following content is from an older version of this website, and may not display correctly.

Have you ever wondered what happens between the moment you click the “Home” button on your Facebook page and the moment you see your friends’ latest posts – the rants, the news links, the baby pictures, the cat videos – with some not-too-finely personalized ads to the right of the newsfeed? If you have, you’ll finally get your questions answered. If you have not, what Jason Taylor has to say will make you wonder why that is.

Once you hit the “Home” button, a ping goes out to a “Web-tier” server, which passes it on to a cache to check whether the hit comes from a logged-in user. Once it gets the green light, the cache pings the Web-tier server, which in turn sends a request to a 40-server rack called the “Multi-feed rack”, which contains all Facebook user activity for the past two days. An aggregator program on one of the servers here (it runs on all of them) sends a request to all 40 servers in parallel, asking them to hand over the latest and greatest of what each of your friends has posted. The aggregator then takes that data, ranks it based on who Facebook thinks your closest friends are and passes the top 40 posts back to the Web-tier server. The Web-tier sends another request to cache to get the material it needs to show you those stories: the text, the comments, the number of “Likes”, links to photos, names of the authors, etc. Chances are, all of this stuff is in cache (Facebook has a 96% cache/hit ratio). If something is missing, a request goes out to a database server to find it there. Finally, the software grabs some ads from an “Ad” server, makes a few more cache hits and ships the product to your screen.

A second has gone by.

Jason Taylor is Facebook’s director of capacity engineering and analysis, which means his job is to make sure you see your friends’ stupid sandwich photos that quickly. His job is to make sure Facebook never runs out of capacity. His other job is to make sure Facebook does not pay more for equipment than it absolutely needs to, and to make sure the company uses every penny invested in infrastructure and every kilowatt burned as efficiently as unhumanly possible. Taylor talked about his job at January’s Open Compute Summit in Santa Clara, California.

A heavy and constant load

His is an around-the-clock occupation. Most Facebook data centers are in the US, while 82% of its users are in other countries. This means “our servers are hot for probably about 14 to 16 hours of the day.” Most European users get on Facebook around 8am Pacific, when traffic from Europe starts to pick up, reaching its peak by 9am and staying high until about 4pm. Traffic from the US East Coast is at its peak between about 10am and 4pm. The peak is lower in the middle of the US, where there are fewer computers, so there is usually a dip after 4pm Pacific, which lasts only until about 7pm, when Asia comes online.

The load is heavy and constant. The social network has more than one billion users who add 320m photos every day, post, comment and “Like” posts and comments 12.2bn times a day, make 140m friend connections and check in at 17bn places. As of mid-January, Facebook’s data centers were storing 212bn photos.

The company spent US$1bn in capital on data centers and equipment inside them in the first nine months of 2012. Not surprisingly, “we spend a lot of time thinking about efficiency and cost,” Taylor says. To sum up Taylor’s perfect world in a sentence: A perfect world is where every Facebook service uses 100% of the hardware it is getting. They’re not there, but they are fighting hard to get there.

Five Facebook servers

Right now, the company uses five types of servers. “Facebook engineers can have any server they want, as long as it’s one of these five types,” Taylor says. Facebook designed all five of them in-house, each optimized for a certain task or multiple tasks, “vanity-free”, without bells or whistles. Web-tier servers have a lot of CPU but very little everything else; database servers are heavy on Flash-storage capacity; Feed servers have a lot of CPU and RAM; Hadoop servers have a lot of disk and CPU; Haystack (the photo server) is optimized for as much disk per dollar as Taylor can possibly get.

All except for Feed do exactly the task they are designed to do. Life is more complicated for Feed servers, which do a lot more than just generate the newsfeed.

Why limit your engineers to five types of servers? Well, the price per server is a lot cheaper when you are buying a few thousand units of the same SKU at once. Vendors will bend over backwards to outbid each other when you’re buying at Facebook’s volumes. Another reason is repurposing. It is very easy to move a server from one service to another if a service has been overprovisioned. A homogenous server fleet also makes life easy when it comes to repairs, drivers, debugging, etc. Facebook only needs a few data center techs to service a huge amount of boxes. Finally, there is plenty of servers on hand for unexpected needs.

So what’s the rub? The rub is Facebook is made of more than five services. “We have 40 major services and about 200 minor ones,” Taylor says. “The problem there is that not all of the services fit perfectly within that footprint.”

Does this mean Facebook has hit the efficiency brick wall?

Enter the Disaggregated Rack

Not only do all services not fit the existing footprint perfectly, services also change all the time. As they change, their hardware requirements may change as well. This is why Facebook’s open-source-hardware community the Open Compute Project is now hard at work on the Disaggregated Rack – a rack where each server resource, be it RAM, CPU, disk IOPS, disk space, Flash IOPS or Flash space, can be scaled independently.

This design effort is focused mainly on the Multi-feed, or Feed, servers. The other ones, like the photo server or the Web-tier server, each consume only one “bottleneck,” as Taylor calls the crucial resources. To scale its photo capacity, the company can just buy more photo servers without worrying that it will end up with a lot of stranded CPU capacity. CPU is already fairly low in each of these disk-heavy boxes.

The reason it is a disaggregated rack instead of a disaggregated server is that Facebook’s software, when thinking in terms of hardware resources, thinks about resources within a rack. That is the social network’s unit of capacity.

Thus, a rack of Multi-feed servers has 80 processors (640 cores), 5.8TB of RAM, 80TB of disk storage and 30TB of Flash. In a disaggregated rack, each of these resources would be treated as a node. A compute node, for example, would have two processors, 16 DIMM slots, no hard drive (network boot) and a big NIC. Rather than depreciate this server over three or four years (the typical hardware refresh rate), Facebook will be able to depreciate it faster – more in line with Moore’s law – over about two years. Because over three years, CPU capacity triples, while other bottlenecks retain their diameters. “When you look at the power consumed by these servers, the TCO starts to turn in the direction of replacing a server at around two [or] two-and-a-half years, so we’d like to replace this independently of all other resources in the rack,” Taylor says.

But the same idea applies to all other resources in the rack, which would also have RAM sleds, Flash sleds, hard-drive sleds, etc. These memory and storage sleds should all last four to six years, Taylor says. “There’s nothing in particular that would cause any of those to fail – outside of the disk – as fast as the two- or three-year hardware refresh, so we want to keep those longer.”

With friends like these

At the Open Compute Summit, Intel and Facebook announced they were working together to make the disaggregated rack a reality. Taiwanese hardware manufacturer Quanta Computer showed a mechanical prototype of the architecture. The prototype supported Intel’s Xeon processors and its next-generation Atom System-on-Chip (SoC) nodes. It featured Intel’s silicon-photonics technology and the chipmaker’s Ethernet silicon for I/O.

Right now, the disaggregated rack is a concept and a prototype, but with Intel’s seemingly limitless engineering resources and Facebook’s open-mindedness, a shipment of these bad boys from somewhere in Taiwan to a data center in Oregon, North Carolina or Sweden may arrive very soon.