ARTIFICIAL INTELLIGENCE

Stanford Researcher Pitches Homa as Successor to TCP for AI Data Centers

In a conference talk, Stanford professor emeritus John Ousterhout argued that AI workloads are shifting toward small, latency-sensitive messages that TCP and RDMA handle poorly — and presented Homa, a clean-slate transport protocol he says cuts tail latency by an order of magnitude.

Data center technician inspects server racks and network switches in a cooled server aisleARTIFICIAL INTELLIGENCE

Image: IntraGoals Media · Uploaded by IntraGoals — usage rights confirmed

AI networking has long been judged by one metric: throughput, the raw volume of data that can be pushed between machines. But according to John Ousterhout, professor emeritus at Stanford University, that measure is losing its relevance as AI workloads evolve — and the protocols built for the old world, chiefly TCP and RDMA, are struggling to keep up.

Speaking at a recent industry conference, Ousterhout laid out a case for rethinking data center networking from the ground up. Historically, he explained, AI workloads were dominated by enormous transfers — gigabytes of data such as weight gradients moving between machines during training. For these jobs, throughput was the only metric that mattered, and legacy protocols like TCP and RDMA (specifically RoCE, RDMA over Converged Ethernet) handled them well.

That is changing. Inference and so-called agentic workloads, Ousterhout said, increasingly rely on small, frequent exchanges of metadata and coordination signals — checking whether an entry exists in a distributed key-value cache, for instance, or synchronizing nodes at the end of a compute cycle. For these exchanges, what matters is not throughput but latency, and specifically tail latency: the worst-case delay experienced by a small fraction of messages.

The stakes are higher than they might appear. When a workload is split across multiple nodes that each complete a burst of GPU computation and then synchronize before proceeding, the entire process stalls until every exchange finishes. If computation phases take only milliseconds — as is increasingly common in agentic, token-generating workloads — then a synchronization step that also takes milliseconds can waste a significant share of GPU capacity. During the talk, a show of hands revealed that many in the audience had already observed this effect in their own systems.

Ousterhout traced the problem to a phenomenon called incast, in which multiple nodes send data to the same destination simultaneously, overwhelming the last link into that node and causing packets to queue — or be dropped — at the top-of-rack switch. Short messages caught behind long ones in that queue suffer disproportionate delays. Compounding the issue, he said, congestion control in TCP and RDMA is handled by the sender, which only learns about congestion indirectly, often several round trips after it begins, making the system prone to oscillating between sending too much and too little. A further structural issue, he noted, is that TCP treats data as an undifferentiated byte stream, with no awareness of message boundaries — making it impossible to prioritize short messages or know how much more data is coming.

To address these shortcomings, Ousterhout and a former PhD student, Behnam Montazeri, developed Homa, a transport protocol designed from scratch for data center conditions. Unlike TCP, Homa is message-based rather than stream-based, built around the remote procedure call as its fundamental unit. Because Homa tracks message lengths throughout, a receiver knows from the very first packet exactly how much data is coming, allowing it to prioritize shorter messages using a scheduling approach called shortest-remaining-processing-time-first.

Critically, Homa shifts congestion control to the receiver, which has far more visibility into incoming traffic than a sender does. Receivers request additional data from senders via explicit "grant" packets, timing and pacing those requests to avoid overloading the network and to favor shorter messages. Homa also makes use of the priority queues built into modern data center switches — typically eight per port — routing short messages into higher-priority queues so they can bypass long messages queued behind them.

In benchmark testing described by Ousterhout, Homa's 99th-percentile tail latency for short messages ran below 100 microseconds, compared with more than a millisecond for TCP — roughly a 13-fold improvement. Notably, long messages did not suffer as a tradeoff: Homa was nearly twice as fast as TCP even at the largest message sizes, which Ousterhout attributed to its run-to-completion scheduling approach.

Ousterhout has built a Linux kernel module implementing Homa, available on GitHub, and is working to get it upstreamed into the mainline kernel. He said he has gone semi-retired from Stanford specifically to focus on the project full time, and invited engineers facing latency bottlenecks in their own AI infrastructure to reach out for help adopting the protocol.

Sources and further reading- YouTube ↗
ABOUT THE DESK

Harsh DV

IntraGoals reports on important changes in technology and work. We check each story for clear writing, trusted sources and useful information before it is published.

KEEP READING

Latest from IntraGoals.

All latest stories ↗