Why AI needs a new kind of supercomputer network
Training frontier models is not just a matter of adding GPUs; one small fault can break the whole coordinated dance. Mark Handley and Greg Steinbrecher explain the new network design behind OpenAI's latest training runs and Multipath Reliable Connection, a protocol developed with AMD, Broadcom, Inte
Training frontier models is not just a matter of adding GPUs; one small fault can break the whole coordinated dance. Mark Handley and Greg Steinbrecher explain the new network design behind OpenAI's latest training runs and Multipath Reliable Connection, a protocol developed with AMD, Broadcom, Intel, Microsoft and Nvidia and opened up to the industry.