Archived Content

The following content is from an older version of this website, and may not display correctly.

Earlier this month we posted a feature from the current issue of DatacenterDynamics FOCUS on Lawrence Berkeley National Laboratory’s real-life tests of a variety of demand-response strategies in multiple California data centers. Here’s a piece from the same magazine issue that describes those tests in detail:

As nice as it would be to have some real-world mission-critical data centers to run tests on, LBNL researchers had to work with what they had access to. These were the lab’s own data center and a small NetApp facility in the San Francisco Bay Area, as well as the San Diego Super Computer Center in South California. Here’s what they did at each of the facilities:

From 100% to 50% load in Berkeley

The first test was done in LBNL’s own data center in October 2011. Facility staff shut down some servers and turned off compressors on two of the CRAC units. As a result, total power went from about 490kW to about 280kW. This shaved 245kW, or about 41%, off the data center’s total load.

The second test at this facility was simply rescheduling processing jobs and idling servers. Using a scheduler program, researchers put off all processing on 30% of a compute cluster’s capability multiple times but the idling servers’ power load did not drop by any “appreciable” amount, since these were older machines whose power consumption does not drop much when idle.

The next test at LBNL’s data center was to adjust the temperature set-point. Operators raised CRAC and CRAH set points by 2F at a time, reducing 1.6% of total facility demand on average – a small savings.

The fourth and final test at this facility involved shutting down IT and CRAC units completely. They did this to study the sequence of operations and to determine demand savings, which were 50% of total load: from 555kW to 280kW.

Shedding storage load at NetApp

At NetApp’s Java-1 data center, the first test was shutting down and idling storage clusters. The operators shut down four filers and their storage hard-disk drives. Some hard drives were idled. This resulted in a drop in outlet air temperature. CRAC units responded automatically by reducing chilled-water flow. The shut-down reduced IT load by about 25kW (14.5%) and cooling and UPS load by 16.5kW (about 10%) – a 25% overall load shed.

Next was adjusting temperature set point at  NetApp. They increased inlet-air temperature by 2F from 72F. One zone’s temperature went up 8F within 30 minutes, and the operators rolled the set-point back. Since the building’s operating PUE was relatively low (1.4), there was no major drop in overall data center load, except a 17kW  load reduction detected at the chiller plant.

The third test was similar to the first one, the only difference being that the operators raised CRAC inlet temperature manually after reducing IT load. Shutting down multiple rows of storage equipment, while at the same time raising inlet temperature by 6F in 2F increments resulted in savings of 24kW of IT load, 23kW of cooling load and about 5kW of UPS load. Overall load drop happens faster in this scenario than it does in the first one.

Shutdown of six filer heads and disk drives at this facility without adjusting CRAC temperatures manually led to an automatic CRAC adjustment, resulting in a 12.1kW cooling-load shed. Including the load shed of IT and site infrastructure, the total load shed was 51kW, or 31%.

Little savings in job migration

The next series of tests was done between the University of California – Berkeley data center and the San Diego Super Computer Center. They first idled 85 of the 268 nodes of the Mako HPC cluster in Berkeley by migrating their processing jobs to the Threshser cluster in San Diego. While shedding Mako’s load by 8.7kW (14%), the migration did little do decrease demand at the building level, which is 1MW. Shutting Mako nodes down instead of idling them led to better savings, but still insignificant for the building as a whole.

The researchers calculated that if they wanted to shed 5% of the building’s load, such migration test would have to be done over 58 racks of the cluster if they were idled, or 30 racks if they were to be shut down completely.

A six-hour response

In the previous test, the researchers moved processing jobs between homogeneous compute clusters. They did another load-migration test moving jobs among heterogeneous systems at the San Diego Super Computer Center and at the LBNL data center. To shed load in San Diego, they let the cluster finish the processing jobs it was doing, while redirecting any new ones to Berkeley.

If this approach was used as part of a demand-response strategy, it would have to account for the time it would take for ongoing jobs to finish, as the response time in this case is directly tied to the size and duration of a job. In one such test, for example, it took six hours for the CPU utilization rate of a compute cluster to reach 0%.

Another thing to take into consideration in any job migration is whether the systems processing load is moved to have enough spare capacity to absorb them safely.

The researchers observed a 9kW demand savings during such a test for the San Diego cluster. While amounting to 17% of the cluster’s total load, this was too little to be noticeable at the building level, whose total demand is 2.3MW.

 A version of this article first appeared in DatacenterDynamics FOCUS magazine. Visit the FOCUS registration page for a free subscription.