AMD Stops Selling Chips, Starts Selling Racks
For most of the GPU era, AMD's pitch to customers was a chip: buy the accelerator, then build and wire your own rack around it. Helios changes that pitch entirely. The system pairs 72 Instinct MI455X GPUs with 18 sixth-generation EPYC 'Venice' CPUs and AMD's own Pensando networking silicon inside a single rack, delivering more than 31 TB of aggregate HBM4 memory as one pre-integrated unit. Pareekh Jain, CEO of EIIRTrend & Pareekh Consulting, frames the shift bluntly: Helios is 'AMD's first complete AI rack system with GPUs, CPUs, and networking built together, instead of selling separate chips' [1]. That distinction matters more than the spec sheet suggests - a rack-level product lets AMD compete on the same terms Nvidia has used for years with its NVL72 systems, rather than asking customers to assemble a system themselves and hope the parts play nicely together.
AMD is also pushing the integration idea past the rack itself. Its new partnership with Cerebras splits an inference request into two stages across two different chip architectures: Helios handles the prompt-processing and long-context portion of a request, while Cerebras's wafer-scale engine takes over for rapid token generation - a disaggregated design the two companies say can push tokens-per-second-per-watt up to 5x higher than a single-architecture approach [2]. Cerebras CEO Andrew Feldman called it 'an incredible opportunity' for ultra-low-latency inference [2]- a notable admission, coming from a would-be rival chipmaker, that no single architecture currently owns every stage of the inference pipeline.




