Cerebras Systems (NASDAQ:CBRS) has launched the CS-4, its latest rack-scale artificial intelligence accelerator, combining three newly introduced Wafer Scale Engine 3 Turbo (WSE-3T) processors in a single system. According to the company, initial shipments are expected to begin during the current quarter.
The CS-4 delivers 750 petaflops of AI computing performance, alongside 129.6 petabytes per second of memory bandwidth and 7.2 terabits per second of system I/O bandwidth.
Cerebras says the platform can generate more than 4,400 tokens per second per user when running the GPT-OSS-120B model. The company claims that performance is as much as 30 times faster than GPU-based alternatives and approximately twice the speed of its previous-generation CS-3.
CS-4 Targets Higher Performance With Improved Efficiency
Beyond raw computing power, Cerebras is positioning the CS-4 as a significant efficiency upgrade over its predecessor.
The company claims the new system can deliver as much as 10 times greater throughput per watt than the CS-3. I/O latency has also been reduced from five microseconds on the previous platform to as little as two microseconds.
Those improvements are designed to support much larger AI workloads, with Cerebras saying clusters built around the CS-4 can accommodate models containing more than 50 trillion parameters.
WSE-3T Packs Four Trillion Transistors Onto a Single Wafer
At the heart of the CS-4 is the new WSE-3T processor, which incorporates four trillion transistors and 900,000 computing cores across 46,225 square millimetres of silicon.
Each processor includes 44GB of on-wafer SRAM and provides 250 petaflops of AI compute performance together with 43.2 petabytes per second of memory bandwidth. Those figures represent roughly twice the compute and memory bandwidth offered by the previous WSE-3 generation.
By installing three WSE-3T processors within the CS-4, Cerebras is seeking to deliver the performance required for increasingly demanding AI inference and large-model workloads without relying on conventional clusters containing large numbers of individual GPUs.
Nexus Architecture Designed to Simplify AI Infrastructure Deployment
The CS-4 also introduces Cerebras’ new modular Nexus Platform Architecture, designed to simplify installation and management of its wafer-scale systems.
A rear-mounted “backpack” brings together power conversion, liquid cooling, high-speed I/O and control electronics around each wafer.
Cerebras says the redesigned architecture can reduce deployment times from days to hours. It also uses 50% fewer components than the previous generation while incorporating 60% more automated manufacturing.
The changes could make large-scale Cerebras installations easier to deploy while reducing some of the infrastructure complexity associated with high-performance AI computing systems.
Cerebras Expands Connectivity for Large AI Clusters
The CS-4 supports RoCE v2 RDMA over Ethernet and introduces Direct Wafer Links, allowing wafers to communicate directly without requiring conventional network switches.
Cerebras says the lower-latency architecture enables the construction of large clusters capable of handling models exceeding 50 trillion parameters.
The system’s I/O infrastructure is also designed to work with disaggregated inference configurations involving technology from partners including AMD Helios and AWS Trainium.
With the CS-4, Cerebras is strengthening its challenge to conventional GPU-based AI infrastructure by combining greater wafer-scale computing performance, faster memory access, lower latency and a more modular deployment architecture.
Cerebras Systems stock price