Building the Standardized AI Factory Stack
Nvidia's rack-scale push arrived in a pair of announcements dated August 25, 2026: Cisco is expanding its Secure AI Factory architecture with Nvidia into rack-scale systems, bringing in Supermicro to supply the compute layer[1]. The joint architecture is described as the first NVIDIA Cloud Partner-compliant reference design built on partner-developed networking spanning both Cisco Silicon One and Nvidia Spectrum-X switch silicon, with rack-to-fabric liquid cooling exceeding 200 kilowatts per rack and integrated Supermicro systems slated for October 2026[1][2]. Cisco's Will Eatherton described the technical split cleanly: Cisco layers its NX-OS or SONiC software on top of Nvidia's Spectrum silicon for the scale-out GPU fabric, while Nvidia's Marc Hamilton frames the deal as reaching beyond a single rack to a full AI-factory-wide reference architecture that partners adopt wholesale rather than assemble piecemeal[1]. That framing lines up with how Jensen Huang has described the AI factory in a recent fireside chat with Cisco CEO Chuck Robbins: as five stacked layers - energy, chips, infrastructure, models, and applications - with Huang arguing enterprises should build on-premises or hybrid rather than rent pure cloud capacity, since the most valuable intellectual property a company holds is not the answers a model produces but the proprietary questions and context used to prompt it.
Underneath the rack sits Nvidia's newest networking layer: BlueField-4 DPUs introduce what Nvidia calls 'Scale-In,' a fifth pillar of AI factory networking. The 64-core Grace-based DPU runs at up to 800 Gb/s and offloads security, storage, and tenant-isolation processing independent of the host CPU, with Nvidia citing up to 1.45x the storage throughput of off-the-shelf Ethernet[3]. Nvidia has also pointed to its own internal AI factory operations as evidence that capacity like this cannot be conjured overnight - the company has said its internal AI factory serves several trillion tokens per month at very high availability, and that standing up new capacity requires many months of procurement and power planning, which is part of why pre-validated, off-the-shelf reference designs matter even as physical buildout timelines stay long. Together, the rack-scale reference architecture and the DPU-level networking layer are meant to do the same job at two different altitudes: turn what used to be a bespoke, vendor-by-vendor data-center build into a standardized, validated design that any enterprise, neocloud, or sovereign cloud can order off the shelf.



