OpenAI, NVIDIA Forge MRC Protocol to Transform AI Training Networks
OpenAI has officially announced a collaboration with five industry leaders—AMD, Broadcom, Intel, Microsoft, and NVIDIA—to launch the Multipath Reliable Connection (MRC) protocol. This open-source protocol, released through the Open Compute Project (OCP), is designed to tackle the network latency and failures commonly encountered in large-scale AI training.

Eliminating the "Single Point of Failure": From Three-Tier to Two-Tier Architecture
In traditional AI model training, network congestion or a minor failure on a single link can cascade like dominoes, forcing tens of thousands of GPUs into idle states and leading to significant computational waste.
To fundamentally improve system resilience, the MRC protocol introduces a multi-plane network design. It intelligently divides a single 800Gb/s interface into multiple smaller links. This structural optimization enables the system to support massive clusters of up to approximately 131,000 GPUs using just two switch layers. Compared to traditional two- or four-tier architectures, this shift not only drastically reduces the number of physical components and energy consumption but also significantly cuts construction costs.
Advanced Traffic Management: Packet "Spraying" and Microsecond-Level Recovery
Beyond architectural simplification, MRC introduces a novel approach to traffic distribution. It employs adaptive packet spraying technology, moving away from traditional single-path transmission. This method breaks down task packets and distributes them across hundreds of parallel paths. Even if packets arrive out of order, the receiver can accurately reassemble them, effectively preventing localized congestion in the core network.
For network control, MRC replaces complex dynamic routing protocols (like BGP) with SRv6 source routing technology. This allows the sender to directly specify the path, while switches perform only simple static forwarding. This design slashes network fault recovery time from seconds to microseconds, enabling the system to achieve near "seamless self-healing" in the face of link instability.
Real-World Validation: The Supercomputer "Stabilizer"
The MRC protocol is already deployed in NVIDIA's GB200 supercomputer and Oracle's cloud infrastructure. Test data confirms that even during active training scenarios, MRC can automatically reroute around disruptions—such as sudden link jitter or switch reboots—ensuring complex training tasks continue without interruption.
Related article
DeepMind CEO Hassabis: I sleep six hours a day, usually feel energetic around 1 a.m.
Fortune recently featured an interview with Demis Hassabis, CEO of Google DeepMind, revealing his unconventional approach to rest and productivity. Hassabis disclosed that he sleeps very little, structuring his waking hours into two distinct work blo
OpenAI, Anthropic Vie for Market Share Despite Revenue Shortfalls
Despite recent reports suggesting OpenAI missed revenue targets, creating pressure on tech stocks this Tuesday, private AI lab investors remain resilient. Seasoned backers have confirmed they will not reduce investment despite negative media coverage
California AV Compliance: A New Era of Tickets, Geofences, and 1M Miles
Guident operates an AuveTech shuttle in South Florida, managing a four-mile route in West Palm Beach and a one-mile route in Boca Raton using its remote monitoring technology. | Credit: GuidentCalifornia is redefining the regulatory landscape for dri
Related Special Topic Recommendations
Comments (0)
0/500
OpenAI has officially announced a collaboration with five industry leaders—AMD, Broadcom, Intel, Microsoft, and NVIDIA—to launch the Multipath Reliable Connection (MRC) protocol. This open-source protocol, released through the Open Compute Project (OCP), is designed to tackle the network latency and failures commonly encountered in large-scale AI training.

Eliminating the "Single Point of Failure": From Three-Tier to Two-Tier Architecture
In traditional AI model training, network congestion or a minor failure on a single link can cascade like dominoes, forcing tens of thousands of GPUs into idle states and leading to significant computational waste.
To fundamentally improve system resilience, the MRC protocol introduces a multi-plane network design. It intelligently divides a single 800Gb/s interface into multiple smaller links. This structural optimization enables the system to support massive clusters of up to approximately 131,000 GPUs using just two switch layers. Compared to traditional two- or four-tier architectures, this shift not only drastically reduces the number of physical components and energy consumption but also significantly cuts construction costs.
Advanced Traffic Management: Packet "Spraying" and Microsecond-Level Recovery
Beyond architectural simplification, MRC introduces a novel approach to traffic distribution. It employs adaptive packet spraying technology, moving away from traditional single-path transmission. This method breaks down task packets and distributes them across hundreds of parallel paths. Even if packets arrive out of order, the receiver can accurately reassemble them, effectively preventing localized congestion in the core network.
For network control, MRC replaces complex dynamic routing protocols (like BGP) with SRv6 source routing technology. This allows the sender to directly specify the path, while switches perform only simple static forwarding. This design slashes network fault recovery time from seconds to microseconds, enabling the system to achieve near "seamless self-healing" in the face of link instability.
Real-World Validation: The Supercomputer "Stabilizer"
The MRC protocol is already deployed in NVIDIA's GB200 supercomputer and Oracle's cloud infrastructure. Test data confirms that even during active training scenarios, MRC can automatically reroute around disruptions—such as sudden link jitter or switch reboots—ensuring complex training tasks continue without interruption.
DeepMind CEO Hassabis: I sleep six hours a day, usually feel energetic around 1 a.m.
Fortune recently featured an interview with Demis Hassabis, CEO of Google DeepMind, revealing his unconventional approach to rest and productivity. Hassabis disclosed that he sleeps very little, structuring his waking hours into two distinct work blo
OpenAI, Anthropic Vie for Market Share Despite Revenue Shortfalls
Despite recent reports suggesting OpenAI missed revenue targets, creating pressure on tech stocks this Tuesday, private AI lab investors remain resilient. Seasoned backers have confirmed they will not reduce investment despite negative media coverage





Home






