Microsoft is one of the world’s most innovative data center designers and operators. The company has a massive global network of data centers, much of which it designed itself.
Earlier this week, the company’s Global Foundation Services team released a whitepaper where it listed 10 best practices it follows to make its data center infrastructure more resilient and cost- and energy-efficient.
Here’s a summary:
Microsoft’s first advice is to provide incentives that support your goals. The company ahs a charge0back mechanism for online services it provides to its employees, which is based on the amount of energy used to provide those services.
The company also provides incentives to data center managers for both uptime and for energy efficiency improvements measured by Power Usage Effectiveness (PUE).
Microsoft also charges its business units fees for carbon emissions.
The second piece of advice is to focus on effective resource utilization. An underutilized data center equals to expensive stranded capacity, stranded in UPS, generators, chillers, etc.
A regular 12MW data center that is half-used can mean US$4-8m a year in unused capital expenditure.
The next word of wisdom is “virtualization”, which ties into the previous advice. Virtualization can be used to improve server utilization and increase operational efficiency.
“By migrating applications from physical to virtual machines and consolidating these applications onto shared physical hardware, Microsoft data centers are increasing utilization of server resources such as central processing unit, memory and disk input/output,” the whitepaper’s authors write.
Next on the list is driving quality up through a comprehensive compliance program. This means tying operations technologies and processes into a security program and control framework that is regularly evaluated by outside parties.
Also important is a standardized change-management process. “Poorly planned changes to the production environment can have unexpected and sometimes disastrous results, which can spill over into the planet’s environment when the impacts involve lower energy utilization and other inefficient use of resources,” the authors write.
“Standardized procedures for the request, approval, coordination, and execution of changes can greatly reduce the number and severity of unplanned outages.”
The sixth advice is investing in understanding application workload and behavior. The better you understand applications and particulars of the traffic on your network, the better positioned you are to improve performance.
Right-sizing the server and network platforms is also key. Microsoft works with sever manufacturers to optimize their designs for its applications.
If your scale is too small to incentivize the OEM, develop exact specifications of the servers you need and do not buy servers that exceed those specifications. “It’s often tempting to buy the latest and greatest technology, but you should only do so after you have evaluated and quantified whether the promised gains provide an acceptable return on investment.”
Of course, evaluating and testing servers for performance, power and total cost of ownership is paramount. Microsoft’s procurement philosophy is built around testing. “Our hardware teams run power and performance tests on all “short list” candidate servers, then calculate the total cost of ownership, including energy costs.”
Another piece of advice is to standardize as much as possible. The more varied your infrastructure is, the more expensive it is.
Standardizing on a small set of servers, network gear and data center technologies drives economies of scale and reduces support costs.
Finally, Microsoft suggests that you take advantage of competitive bids from multiple manufacturers. This fosters innovation and reduces cost.
“Competition between manufacturers is a good thing, which Microsoft encourages through ongoing analysis of proposals from multiple companies that puts most of the weight on price, power, and performance.”
Download the actual whitepaper from the Global foundation Services team’s blog.