Skip to main content

The buildings and physical infrastructure for a new hyperscale data-center region are complete. Power is live and redundant, the cooling loops are balanced, the network fabric is up, and tens of thousands of machines are racked, cabled, and reporting healthy. Every milestone on the construction schedule is closed. The region will not carry production traffic for another six months, and whether it turns out to be six or closer to twelve is decided almost entirely by software.

There is a large literature on how hyperscale data centers get financed, powered, and cooled. The standard reference on warehouse-scale machines, Barroso and Hölzle’s book, covers the design of these computing facilities in depth. Most of that literature describes data centers that are already operating. The transition from completed facilities to production readiness receives much less attention, even though it carries a significant share of the schedule risk. Getting that transition wrong is expensive.

Once the facilities and hardware are ready, platform teams still have to bring hundreds of interdependent services online. Many services require other services to be running first. If you draw each service as a box and each startup requirement as an arrow, the result is a dependency graph. The graph shows the order in which services can be brought online and identifies the dependencies that must be resolved before the region can serve production traffic.

Getting a region into production means making several different things true at once. The network fabric has to route traffic within each data center, and the links between facilities have to carry traffic reliably. Compute platforms have to schedule workloads, while storage platforms have to persist their data. Identity has to work too: certificate authorities have to issue certificates, services need access to secrets, and engineers need permission to finish the build. Engineer access sounds obvious until the controls protecting the new region are the same controls slowing down the people trying to turn it on. Stateful systems that depend on existing production data have to be populated and validated, because a database cluster with no data in it is furniture. The installed capacity has to be assigned to foundational services and production workloads, and teams have to test how the new region behaves when networks, services, or dependencies fail. Applications are then deployed and validated before traffic moves over gradually, with health checks and a tested rollback at each step, ideally without users noticing.

Each platform or service has an owner, a plan, milestones, and its own definition of done. What is often missing is ownership of the complete dependency graph. The team coordinating the region launch has to map the dependencies across teams, determine the order in which services must come online, track what is blocking that sequence, and keep the graph current as plans change. Without that end-to-end view, every team can report that its own work is on track while the region as a whole remains blocked.

Mapping the graph begins with a simple question for every service: What must already be available before this service can start in a new region? The goal is to identify true startup requirements, not every system the service communicates with during normal operation. Repeat that exercise across a large platform’s control plane, and you can uncover hundreds of dependencies spanning dozens of systems. Some dependencies surface only when another service identifies them as a prerequisite. Hidden dependencies are usually ordinary services that have been quietly reliable for so long that the teams relying on them no longer think about what would happen without them.

Once the startup dependencies are mapped, the next step is to look for loops: cases where one service needs another service to be running, but that second service eventually depends on the first. In a large platform, a surprising share of the control plane can be tied together by these loops. The result is that there is no valid order in which to start the services. Every possible starting point eventually leads back to a service that is still waiting. The loops themselves are usually mundane. DNS may depend on the inventory system that tracks what hardware exists, while the inventory system relies on DNS to resolve names. The artifact repository holding every installable package may depend on configuration management, which is itself installed from a package in that repository.

Nobody designed any of this. Each dependency was a locally sensible decision made by a competent team, often years apart from the decisions that completed the loop. In a running region, the required services are already available, so the loop remains silent. The database is up when the alerting store starts, and nobody learns whether either could recover without the other. Starting a region from scratch is often the only event that reveals whether those services can start independently. Meta’s outage in October 2021 shows how a large failure can expose dependencies that normal operation keeps hidden. When the backbone network connecting Meta’s data centers went down, its DNS servers withdrew their routes as designed to keep traffic away from unhealthy connections. That safeguard made DNS and many internal tools unreachable. With remote access unavailable as well, engineers had to go onsite, slowing recovery. A locally sensible safeguard had made system-wide recovery harder.

Traffic-drain tests can reveal some of these dependencies. Meta’s Maelstrom encodes service dependencies and resource constraints to shift traffic safely from a failing data center to healthy ones, and its drain tests can uncover missing dependencies. But the receiving data centers are already running. A drain tests whether live infrastructure can absorb traffic; it does not test what an empty region needs in order to start. The drain graph is a useful input to the startup graph, not a substitute for it.

Status reports for a new region can be misleading even when every team is reporting honestly. A service may be deployed, configured, monitored, and passing its health checks, yet still be blocked by a startup dependency. Readiness therefore has to propagate through the graph: a service is ready only when its own checks have passed and every service it needs at startup is also ready.

Once the dependency graph exists, five practices turn it from a diagram into a launch plan. The first is finding what actually determines the launch date. The graph shows which services must wait for others, but the order alone does not reveal how long the work will take. Ten services that can start in parallel may finish before three services that must start one after another. Estimate the bring-up and validation time for every service. If several services remain tied together in a loop, treat them as a single planning block and include the time required to break the loop. The chain with the greatest total time becomes the critical path. Then count how many other services each foundational service can hold back, including dependencies several steps away. DNS, relational databases, secrets stores, and configuration management often rise to the top. Staff those teams early, because a week lost in one of them can become a week lost across the entire region.

The second practice is to break the loops. One way to do that is to temporarily borrow a service from a region that is already running. Designate a small bootstrap tier, the minimum set of services required to deploy other services, and configure those bootstrap services to use working dependencies in an existing region until the local versions are ready. Suppose configuration management needs a package from the artifact repository, while the artifact repository needs configuration management before it can start. For the first installation, configuration management can fetch its package from another region. It can then bring up the local artifact repository and switch to using it. The bootstrap tier might also include identity, inventory, and package distribution. This shortcut works only when cross-region access is permitted and reliable enough. It also creates another cutover that must be planned, tested, and completed later.

Borrowing from another region helps the new region get started, but it should not become permanent. Over time, each bootstrap service should be able to start without all of its usual dependencies. Suppose a service normally waits for the configuration system before it can start. Package the minimum settings it needs with the service itself. The service can start with those settings and fetch the latest configuration once the configuration system is running. Test this by turning off the dependency and starting the service from a clean state. If the service still cannot start, record the problem with an owner and a target date for fixing it.

The third practice is to bring the region up in explicit tiers. A typical sequence begins with foundational configuration such as network ranges and routes, hardware inventory, identity configuration, access policies, service endpoints, and deployment settings. Bootstrap services such as DNS, certificate issuance, secrets, software package distribution, and configuration management follow. Next come the control planes that provision resources, schedule workloads, manage storage, and support service discovery. Stateful systems and applications come after the platforms they depend on. The exact tiers will vary by architecture, but the dependency graph should determine the sequence. Tier gates prevent visible application progress from hiding unfinished foundations.

Stateful systems that depend on existing production data need special treatment because creating a cluster is often quick, while filling it with data is not. A storage system is ready only after the required data has arrived and been validated. Estimate that work using data volume, available bandwidth, validation time, and enough headroom for retries. Give each system a time box based on those measurements. If the estimate changes, require updated measurements that explain why. This keeps the plan honest without pretending that every delay is avoidable.

The fourth practice is to make the bring-up repeatable, which does not mean automating every step. Automation helps only when it is maintained and tested. A script written for one region and left untouched for years may be more dangerous than a clear manual procedure. Automate steps that use the same tools as regular deployments or can be exercised frequently. For rare steps, maintain a runbook with validation checks, a named owner, and a schedule for testing it. Both automation and runbooks should clearly identify the required inputs, the evidence that a step succeeded, and how to continue after a partial failure. Google’s SRE book raises the same concern: turn-up automation maintained separately from the systems it supports can become outdated. At every manual step, ask whether another engineer could repeat it during a recovery without relying on the people who completed the original build. The goal is to leave behind a procedure that still works after the original team has moved on.

The fifth practice is validation. It’s natural to test only the highest-traffic paths, but that approach misses structural problems. You also want the flows with the widest fan-out, the ones touching the most systems on the way through, even if few people use them, because those flows traverse more of the dependency graph. A high-volume request may prove that the region can handle load. A wide-reaching request can uncover an unready identity, storage, messaging, or data service. Meta’s Kraken shows another useful validation method by shifting live user traffic into a data center while monitoring latency, errors, and system health. Live traffic also shows where capacity goes as caches warm, retries appear, and background jobs compete with user requests, a combination synthetic tests struggle to reproduce.

When the traffic ramp begins, send a small percentage of representative production traffic to the new region, then increase it in stages. The size of each step should reflect the scale and risk of the platform, because even 1% can represent a substantial workload. This allows the entire application path to experience load together instead of testing each service in isolation. Hold at each step long enough for caches to warm, queues to stabilize, and relevant periodic jobs to run. Decide in advance which health signals allow the ramp to continue and which ones require it to stop. Also define and test how traffic will return to the existing region if health degrades. Confirm that the existing region has enough capacity and that data written in the new region will remain safe. Otherwise, the health signals may tell you something is wrong without giving you a reliable way to recover.

These practices make no promise of a fast launch. They make the build understandable and leave behind a process the next region can use. The dependency map will begin aging as soon as systems change, but the ownership model, readiness rules, tier gates, repeatable procedures, and validation process can remain. The same tools are also useful for a disaster-recovery rebuild. Planning around dependencies will matter more as organizations rethink where their workloads should run, moving some systems from public clouds to private clouds, colocation facilities, or data centers they operate themselves. Each new environment brings another dependency graph that must be understood before it can carry production traffic.

Finishing the facilities remains a genuine milestone. Power, cooling, networking, and hardware create the place where production can run. A working region emerges when its services can start in a valid order, the required data is ready, and the complete system has been tested under traffic. Finishing construction gives the organization a data center. Satisfying every required startup dependency in the graph turns it into an operational region.

Post topics: AI & ML•Data•Infrastructure