Home
Zigong and Tsinghua University unveil ZCube networking architecture, boosting large model inference throughput by 15% while cutting network costs by a third
Large model inference is reshaping AI infrastructure, and innovations in network architecture are now key to unlocking hardware potential. In September 2025, Zhipu, Yuchen Network, and Tsinghua University presented their research on the ZCube network architecture at ACM SIGCOMM 2025, a leading conference in networking.
On May 21, 2026, Zhipu announced the successful deployment of the ZCube architecture in the GLM-5.1 coding production environment, resulting in substantial performance gains. Benchmark tests, conducted with unchanged GPU hardware, software stack, and applications, showed that ZCube reduced capital expenditure on switches and optical modules by 33%, increased average GPU inference throughput by 15%, and cut first token latency (TTFT P99) by 40.6%. This system-level breakthrough strikes a balance between high economic efficiency and high performance.

Currently, with long-context inference and Prefill-Decode (PD) separation deployment becoming industry standards, the cross-node transmission of KV Cache exhibits significant asymmetry. Traditional ROFT (Rail-Optimized Fat-Tree) architectures, which rely on multi-layer switch stacking, are constrained by static topology, making them prone to local hotspots and PFC backpressure. This creates a structural bottleneck where total bandwidth is sufficient but local congestion is frequent.

To address this pain point, the ZCube architecture moves away from the hierarchical stacking of traditional Clos designs. It eliminates Spine layer switches and instead uses two groups of fully flat switches in a bipartite graph interconnection, combined with a dual-port NIC's single/multi-track hybrid access mechanism. Through its unique routing strategy, ZCube guarantees a dedicated optimal path for any GPU pair, achieving perfect traffic load balancing at the structural level and supporting ultra-large-scale expansion to tens of thousands or even hundreds of thousands of GPUs.
During the production environment transformation, the Yuchen Network team successfully tackled the challenges of cabling and route strategy reconstruction using automated control and verification tools, ensuring a fast and stable cluster upgrade. The current thousand-card cluster has been running stably for over two weeks. The successful deployment of ZCube marks a shift in intelligent computing infrastructure from general interconnection to system collaboration driven by model traffic. In the future, the deep integration of network topology, communication libraries, and scheduling strategies will become the core driver for further improving token production efficiency and reducing overall MaaS costs.
Related article
Swedish AI Startup Lovable Eyes $13.2 Billion Valuation After Major Funding Round
As AI-driven coding tools gain traction, Swedish startup Lovable has secured a major funding round. The company aims to raise $3 billion, potentially boosting its valuation to $13.2 billion—double the $6.6 billion recorded last December. Menlo Ventur
Google Tests Remy AI Agent for Gemini as Focus Shifts to User Control
According to Business Insider, Google is testing Remy, a new AI personal agent for Gemini. This tool aims to execute tasks on behalf of users, streamlining both professional workflows and daily routines.Currently, Remy is undergoing testing in an int
How to fix Core Web Vitals for better SEO rankings
Streamline Report Card Comments with AI ToolsIntroductionAI Tools for Generating Report Card CommentsMagic SchoolAlmanac AIChat GPTUsing Magic School to Generate Report Card CommentsLogging into Magic SchoolSelecting the Report Card Comments ToolCust
Related Special Topic Recommendations
Comments (0)
0/500
Large model inference is reshaping AI infrastructure, and innovations in network architecture are now key to unlocking hardware potential. In September 2025, Zhipu, Yuchen Network, and Tsinghua University presented their research on the ZCube network architecture at ACM SIGCOMM 2025, a leading conference in networking.
On May 21, 2026, Zhipu announced the successful deployment of the ZCube architecture in the GLM-5.1 coding production environment, resulting in substantial performance gains. Benchmark tests, conducted with unchanged GPU hardware, software stack, and applications, showed that ZCube reduced capital expenditure on switches and optical modules by 33%, increased average GPU inference throughput by 15%, and cut first token latency (TTFT P99) by 40.6%. This system-level breakthrough strikes a balance between high economic efficiency and high performance.

Currently, with long-context inference and Prefill-Decode (PD) separation deployment becoming industry standards, the cross-node transmission of KV Cache exhibits significant asymmetry. Traditional ROFT (Rail-Optimized Fat-Tree) architectures, which rely on multi-layer switch stacking, are constrained by static topology, making them prone to local hotspots and PFC backpressure. This creates a structural bottleneck where total bandwidth is sufficient but local congestion is frequent.

To address this pain point, the ZCube architecture moves away from the hierarchical stacking of traditional Clos designs. It eliminates Spine layer switches and instead uses two groups of fully flat switches in a bipartite graph interconnection, combined with a dual-port NIC's single/multi-track hybrid access mechanism. Through its unique routing strategy, ZCube guarantees a dedicated optimal path for any GPU pair, achieving perfect traffic load balancing at the structural level and supporting ultra-large-scale expansion to tens of thousands or even hundreds of thousands of GPUs.
During the production environment transformation, the Yuchen Network team successfully tackled the challenges of cabling and route strategy reconstruction using automated control and verification tools, ensuring a fast and stable cluster upgrade. The current thousand-card cluster has been running stably for over two weeks. The successful deployment of ZCube marks a shift in intelligent computing infrastructure from general interconnection to system collaboration driven by model traffic. In the future, the deep integration of network topology, communication libraries, and scheduling strategies will become the core driver for further improving token production efficiency and reducing overall MaaS costs.
Swedish AI Startup Lovable Eyes $13.2 Billion Valuation After Major Funding Round
As AI-driven coding tools gain traction, Swedish startup Lovable has secured a major funding round. The company aims to raise $3 billion, potentially boosting its valuation to $13.2 billion—double the $6.6 billion recorded last December. Menlo Ventur
How to fix Core Web Vitals for better SEO rankings
Streamline Report Card Comments with AI ToolsIntroductionAI Tools for Generating Report Card CommentsMagic SchoolAlmanac AIChat GPTUsing Magic School to Generate Report Card CommentsLogging into Magic SchoolSelecting the Report Card Comments ToolCust











