Manufacturing & Supply Chain

Component Procurement or Full-Rack Delivery: Five Links a Rack-Level Liquid-Cooled AI Cluster Must Close, and the Spares Chain After Handover

2026-08-146 min readturnkey delivery / rack-level liquid cooling / CDU

For the same batch of GPU compute, the line between component procurement and turnkey delivery is who carries the integration risk and integration time. This post breaks down the five links a rack-level liquid-cooled AI cluster must close in sequence — cabinet, CDU, water loop, coolant supply/return monitoring and rack commissioning — how this differs from the delivery boundary of an air-cooled single node, and how the delivery process connects to spare-parts and expansion services afterwards.

One batch of compute, two ways to take delivery

When building an inference or training cluster, procurement usually offers two routes. One is component procurement: servers, GPUs, memory and drives are ordered separately, and after arrival your own engineering team handles assembly, cabling, installation and stress testing. The other is turnkey system or full-rack delivery: the equipment is configured, assembled and verified before leaving the factory, and on site it only needs power, network and water connections before entering production deployment. JH Semiconductor Semiconductor works both routes: a supply-chain service for servers, GPUs and memory on one side, and turnkey delivery covering air-cooled, liquid-cooled and rack-level liquid-cooled clusters on the other. Neither route is absolutely better — the dividing line is who carries the integration risk and integration time. The component route leaves both with your own team; the full-rack route moves them upstream of the factory gate.

Air-cooled single nodes have a short delivery boundary; a liquid-cooled rack has a long one

The delivery boundary of an air-cooled platform is relatively clean. Take the 8-GPU air-cooled platform: eight dual-width or quad-width blower cards, twelve PCIe 5.0 slots and twelve LFF bays, aimed at rapid deployment in standard machine rooms — its whole positioning is to lower the barrier of liquid-cooling retrofits. As long as the room meets power and airflow requirements, delivery converges to three steps: rack, power up, verify. A liquid-cooled full rack is a different matter. The cold-plate platform provides full CPU/GPU cold-plate coverage, so cooling capacity is no longer determined by the chassis alone — it forms one system together with the cabinet, the CDU and the facility water loop. The delivery object is then no longer "a machine" but "a closed loop": if any segment is left open, the rack cannot reach its design operating point.

The five links a rack-level liquid-cooled delivery must close, in order

Broken down by delivery sequence, a rack-level liquid-cooled cluster has five links to close:

  • Cabinet: confirm load bearing, power distribution and the placement of coolant manifolds; where each node sits in the cabinet directly determines pipe routing and service access;
  • CDU: the heat-exchange and circulation hub between the secondary coolant loop and the facility's primary water loop — its interface and redundancy scheme must be aligned with room conditions before the equipment ships;
  • Water loop: after the piping is connected, run pressure-hold and leak checks first and confirm the loop is tight before applying power — the order cannot be reversed;
  • Supply/return monitoring: the liquid-cooled platform supports cluster-level coolant supply and return monitoring; at delivery, temperature and flow data must be wired into the facility management system so anomalies are seen before workloads are affected;
  • Rack commissioning: verify sustained stable operation of the full rack under design supply/return conditions with a full-load stress run — the acceptance criterion is that curve, not the equipment lighting up.

A substantial share of this work happens before the equipment arrives on site. The earlier the facility water conditions, CDU interfaces and monitoring protocols are confirmed, the shorter the on-site dwell time — which is exactly where full-rack delivery compresses the time to production.

Handover is not the end of the process: spares and expansion ride the same chain

Half the value of full-rack delivery shows up after handover. Once a cluster is in operation, the most frequent actions are part replacement and expansion, and serviceability is designed in for both: the 10-GPU Thor T5000 platform carries four CRPS power supplies in 2+2 or 3+1 redundancy, so a single supply failure does not interrupt service; the PC nodes of the PC-Farm platforms support hot plugging and fully independent power on/off, so replacing a failed unit never touches the rest of the rack. And because spare parts come from the same supply chain as the system BOM, replenishment of GPUs, memory and drives — and later expansion nodes — can continue in the original configuration, avoiding the maintenance burden of two divergent configurations inside one rack two years in. Connecting the delivery process to the ongoing supply-chain service into a single line is what separates full-rack delivery from a one-off sale.

Keywords
turnkey deliveryrack-level liquid coolingCDUcoolant supply and return monitoringrack commissioningcomponent procurementAI compute clusterspares and expansionPC-FarmCRPS redundant power