Data centers, once few and far between, have been brought to the forefront in many ways, from the need to build new ones to the extraordinary power demands that affect many people beyond the construction and operation of the facility itself.

As the complexity of data centers continues to grow, several technologies can be standardized, resulting in more efficient processes. This raises the question: Can a data center be ordered as a single part number? Some possibilities exist.

IT components of a data center

The core component of a data center is the server, whether for generalized enterprise compute or a higher-powered server designed for AI training or inferencing. Starting at the center of a data center are the CPUs and GPUs.

A modern AI-optimized server will contain one or more CPUs, typically supplied by AMD, Intel, or one of several ARM-based suppliers, such as Nvidia. The GPUs, which usually do the heavy lifting, are significant power consumers, requiring forward-looking planning.

In addition to servers, cooling systems (in-rack or in-row), optimal-length tubing (based on the server specs in the rack) and external-to-the-data-center cooling towers are needed for the next generation of data centers. Networking, cabling, battery backup units, and power shelves round out the first-level requirements for a modern data center.

Experience shows that 90 percent of the racks in a data center are common, allowing some “standardization” for the data center builder/operator. This also allows for common network topologies and connectivity across a data center.

Standardized components

While the combination of servers in a typical pre-AI data center is nearly limitless, an AI-optimized data center will generally feature clusters of identical servers in each rack, in many cases up to 90 percent of the racks. This leads to the idea that “the rack is the new server.”

Therefore, each rack can be considered a standardized building block available as a single SKU (Stock Keeping Unit). Inside the rack, there would be a specific number of servers with designated CPU models, memory, GPU hardware, networking, and direct-attached storage.

In a liquid-cooled data center, the in-rack CDU can easily be included in the rack-level SKU. This single part number, which provides for other parts, simplifies the ordering process.

Rising up the hardware stack from the rack is the idea of a compute cluster. Just as a rack is a single unit, so is a cluster. Usually, a cluster consists of eight or more racks and can be assigned to a specific workload or dedicated to a particular organization within an enterprise.

SKu
SKU = Stock Keeping Unit: A unique alphanumeric code that a business assigns to a product to internally track components – Supermicro

A multi-rack cluster will most likely have higher-speed networking connections between servers and racks than between clusters. For liquid-cooled data centers seeking to avoid in-rack CDU units, in-row CDUs are a viable option at the cluster level. Therefore, a single cluster SKU would include the necessary number of racks and switches, and cabling components.

A modern data center hosts multiple system clusters connected by high-speed networks or additional switch layers. It typically features a wider variety of servers than a single cluster, including racks with storage servers, network switches, or servers dedicated to enterprise applications, scheduling, accounting, and control.

Although these racks may vary from one data center to another, they share common features based on AI-focused workloads and control systems.

A data center as a SKU?

From a specific vendor perspective, a data center may have many, many racks and associated IT equipment available, which could be ordered as a single SKU. This would simplify ordering but limit customization.

Conversely, the argument against standardized SKUs for a data center is that each data center will be different, in terms of compute performance, storage type, and capacity. Overall, yes, a data center could be a single SKU, but this might only work for smaller, repeatable instances where all networking is standardized.

Servers, performance, or power-based?

The definition of a data center is changing, so too must the concept of a data center as a SKU. In previous time frames, a data center may have been defined and/or measured by the number of servers or flops (floating-point operations per second).

Then, the size of a data center was measured by the number of servers that could fit within its physical footprint. Now, with the latest generations of systems that are drawing unprecedented amounts of power, data centers are specified in terms of power, sometimes in the Gigawatt range. Is the new data center SKU power-based?

Perhaps in the future, instead of ordering a fixed number of servers, racks, switches, etc., the ordering process will be specified by workload and power delivered. The identified workload can then be translated into the number of servers with specific CPUs/GPUs/memory/storage/networking, and the same process can be repeated for the electricity delivered. We will see.

Hear more from Supermicro on the DCD>Broadcast, 'Engineering GPU performance: Design choices for training & inference'.