In preparing for this article, I was looking back through some old emails and found a 2017 conversation I’d had with Alexis Richardson, the then CEO of Weaveworks. At the time, he and I had roles at the Cloud Native Computing Foundation. Alexis was the chairperson for the Technical Oversight Committee; I was the marketing chairperson, and we were reviewing the slide deck that would later introduce the idea of GitOps to the world, a reconciliation loop between a desired state in Git and an actual state in Kubernetes.
Alexis and the Weaveworks team had come up with GitOps to tame the complexity of managing Kubernetes clusters. Born of the foundational work Weaveworks had been doing in the container ecosystem since before the term “cloud native” had ever been uttered, and a hair-raising moment when an engineer accidentally deleted prod, the idea was clever and well-timed, but it was not new.
Reconcile, reconcile, reconcile
Reconciliation loops, systems that continuously compare a desired state against an actual state and reconcile the differences, are not new. Physical regulation systems have been doing this since Watt's centrifugal governor in 1788. Configuration management tools have been doing this since CFEngine in the 90s. Puppet, Chef, Ansible, and Terraform all implemented versions of the same idea: define what you want, let the tool figure out how to get there, and keep checking that you stay there.
The idea of GitOps was planted in fertile ground. Kubernetes provided a useful interface on which reconciliation loops could act: Declarative APIs, resource definitions that could be version-controlled, controllers that understood drift detection, and the CNCF was corralling a huge Kubernetes user community into existence, with the majority of them struggling to sanely scale their deployments while maintaining quick release cycles.
And so it was that GitOps found its place as a member of the ever-growing reconciliation loop hall of fame.
Virtual meets Physical
These days, I spend most of my time interacting with people who are architecting and building massive data centers at breakneck speeds to serve armies of GPUs to AI-hungry users. In the cloud native world, we tend to assume that the underlying physical infrastructure and how it is made available to us are somebody else’s problem. In my current role, that is the problem, and there are reconciliation loops at play here, too, with unique challenges.
High-level and low-level designs
The scale of infrastructure being deployed in AI data centers is exploding, expected deployment timelines are shrinking, and resilience requirements keep increasing. Infrastructure teams typically approach designing for these constraints at two levels:
- High-level designs (HLDs) map out the overall topology, how many racks, which networking, compute, and storage devices will be used, cabling plans, what the network fabric looks like, IP addressing boundaries and constraints, and more.
- Low-level designs (LLDs) are what a high-level design becomes when they are actually deployed - where exactly are the racks, specific device variants and their positions in racks, specific OS and firmware versions, specific interface connections, every cable, every IP assignment & ASN, routing configuration, circuits, and more.
It was not long ago that, for the majority of the market, the practices used to capture and express high-level designs and low-level designs would look decidedly old school to most cloud native folks and software engineers. High-Level designs would often be expressed in tools like LucidChart, and Low-Level designs would often be expressed in spreadsheets. Lucidchart and spreadsheets are no longer going to cut it.
Reconciliation loop 1: High-level designs to low-level designs
Companies that need to move fast are adopting, en masse, new ways to model high-level designs. These are typically YAML files that are declarative, versioned, composable, and customizable on the fly, so that infrastructure architects can create known-good high-level designs that can be reasoned about programmatically and reused to deploy new infrastructure building blocks. These designs vary in scope, from a single device variant with its associated modules, to a full rack of equipment with cabling and management IPs, to a pod of racks, or often these days, a “superpod” containing up to 80 racks.
Once the architects have created their high-level designs, which capture the “what,” the reconciliation tooling takes over and provides the “how.” High-level designs are combined with deploy time variables, rendered and applied idempotently to the low-level design, which more often than not is a Source of Truth system, which is then the reference that is used by infrastructure operators, network operations centers, data center engineers, and anyone else in the company who wants to know “what do we think we actually have running in our data center?”
As is common in reconciliation loop systems, the value doesn’t stop there. Moving to a declarative approach for high-level designs also allows these companies to ask questions like “Has the day-to-day reality in the low-level design drifted from my high-level design intent?” At scale, automation is essential, and automation benefits from homogeneity of infrastructure. Being able to quickly see when the low-level design is drifting helps to prevent snowflakes.
Reconciliation loop 2: Procurement and installation
Of course, designing infrastructure doesn’t magically make your equipment–servers, switches, routers, racks, cables, etc.- magically show up in the data center.
The second advantage of adopting programmatic high-level designs is that they can be used to rationalize and speed up procurement and installation processes, which is especially important in the current market where lead times for equipment can often be the longest pole in the tent.
A programmatic high-level design can be used to create a Bill of Materials (BOM), then one or more Purchase Orders (POs), which lead to multiple Shipments of equipment, which can then be racked and stacked in the data center.
Automatic reconciliation of the full cycle is still difficult (although people are already looking at automating data center tasks with robots), but starting with a queryable high-level design allows all these processes to be tied together with greater ease, shaving off valuable weeks of lead time on deployment dates.
Reconciliation loop 3: Configuring the equipment
Once we have used the high-level design to both populate the low-level design in the Source of Truth for day-to-day use and to power the procurement processes that get the equipment installed, there is one step left: configuring the equipment.
In this stage, we take the infrastructure specifics, stored in the Source of Truth as the low-level design, and talk to the equipment in the infrastructure to make sure it is ready for use. This should be a familiar concept for most readers, as it’s analogous to exactly how you’d use Terraform/OpenTofu to apply changes in the cloud. The Source of Truth serves as the input to the process, and tooling like Ansible, Terraform, or other workflow runners act to idempotently apply the changes to the assets in the network.
The world of physical infrastructure is in the middle of a shift like the DevOps movement we saw in the 2010s. In many ways, the approaches being taken are identical: automation is paramount, homogeneity is a requirement for scale. There are, however, unique challenges in the physical infrastructure space that previous generations of automators were mainly able to avoid. A great number of smart people are working on the problem, and it looks like Design-Driven Automation is one of the leading patterns that we’re going to be hearing much more about.
Comments