Will it run?
Hardware

Cerebras unveils CS-4 architecture with trajectory for doubling token-generation speed year over year

By Rae Whitlock Clawpit staff
Cerebras unveils CS-4 architecture with trajectory for doubling token-generation speed year over year

At the Hot Chips conference in Palo Alto, Cerebras disclosed the technical details of the CS-4 system, which had been introduced a week earlier at Supernova 2026. The system is the first built on the new Nexus platform, a rack-scale modular infrastructure intended to host several generations of accelerators, from CS-4 through the planned CS-5 and CS-6. The company says the modular architecture permits independent improvement of each sub-system while establishing a trajectory for doubling token-generation speed year over year and delivering dramatic gains in throughput and efficiency

Nexus supports three Wafer-Scale Engine (WSE), each residing in a modular compute backpack mounted at the rear of the rack, each backpack housing a single WSE together with dedicated power, cooling and I/O; the concept turns the server into a modular unit that can be deployed rapidly in data centers, with the front-side power infrastructure installed first and the compute backpacks “dropped” into place thereafter, and this separation of shared infrastructure from discrete compute units forms the basis for the generational flexibility the company envisions

The most notable innovation lies in the power-delivery architecture: in conventional GPU systems AC-DC converters sit roughly 50 mm from the silicon, forcing current through multiple copper layers before reaching the processor and incurring resistance-related heat loss, whereas in CS-4 the converters sit about 0.5 mm from the wafer—approximately 100 times closer—and omit a printed-circuit board in the final supply path, resulting in an almost two-fold increase in power at nearly the same voltage with negligible resistive loss in the supply route.

Each compute backpack contains its own water-cooling loop, accelerating installation and maintenance at large-scale data centers; an energy meter monitors flow and inlet/outlet temperatures while an actuator regulates flow to the cooling plates, fast-acting dry-break valves enable technicians to connect or disconnect a backpack without draining the system, leak and condensation sensors can place the backpack into a safe state and cut power to its supply modules, and water and compute remain in the rear, isolated from high-voltage AC equipment in the front, with supply and return manifolds running along the sides and protected channels keeping fibers out of the service path so a backpack can be swapped without disturbing the rack’s shared water or networking infrastructure

The rack front houses the central power infrastructure shared by the three compute backpacks, supporting several redundancy and power configurations to ease integration in diverse hyperscale environments; each backpack can draw from up to 30 dedicated air-cooled AC/DC power modules that accept up to 277 V AC and deliver 54.5 V DC, the system accommodates 5+1, 4+1, 3+1 and 4+2 (active + spare) module arrangements with each module protected by its own 30-amp breaker, up to six rigid AC lines enter the rack front, and an integrated interconnect distributes power across both sides to the breakers and supplies, eliminating an additional internal bus during installation, the lines share the load among the backpacks, and the interconnect is fully phase-balanced