John Ousterhout, professor emeritus at Stanford University, is again proposing Homa, a datacenter transport protocol, as an option for reducing communication latency in AI clusters.
In a talk video published by AI Engineer on September 17, 2026, he argued that what matters is not only moving large volumes of data quickly, but also keeping small messages from waiting behind bulk transfers.
Homa, which has been under research since 2018, assigns priorities per RPC message and lets the receiver control how much the sender transmits. In 2026, changes were also made to limit interference when it coexists with TCP and to alter how long messages begin transmission.
However, using Homa in a real AI cluster involves more than network performance. Integrating it into existing applications, handling encryption, and preparing the operating environment all remain open challenges.
Small RPC delays can stall an entire GPU group
Ousterhout's central point in the talk was that AI cluster traffic is no longer just bulk transfers.
In large-scale training, moving huge volumes of data such as gradients between nodes is critical. In inference and agent-style workloads, on the other hand, small messages are exchanged frequently, such as KV cache queries and synchronization.
Being able to keep sending large amounts of data efficiently and being able to answer small requests immediately cannot be measured by the same performance metric.
When multiple GPUs or compute nodes wait for communication to finish before moving to the next step, the slowest response can determine the progress of the whole job.
Even if most communications finish quickly, if a few are held up for a long time, expensive GPUs cannot proceed in the meantime.
This is where tail latency matters.
P99 is the boundary within which 99% of measured operations complete. Even if the average is small, a large P99 means a few slow communications can stall the entire distributed computation.
One typical cause of such delays is incast.
When data converges on a single receiver from multiple senders at once and the combined sending rate exceeds the capacity of the link to that receiver, packet queues build up at the switch just before it.
If a small request arrives there, it can end up waiting behind the bulk data already in transit.
A 2021 paper comparing TCP and Homa also cites head-of-line blocking, which occurs when long and short messages share the same TCP connection.
TCP treats a connection as a byte stream and delivers it to the application in order. If a large chunk of data sent earlier is stuck, it is difficult to process only the small data that arrived later.
Applications can work around this, for example by using multiple TCP connections.
Homa aims to handle this kind of prioritization in the transport protocol rather than in the application.
Prioritizing messages with less data remaining
Homa treats an RPC (Remote Procedure Call), a request followed by a response, as its basic unit.
Because the total size of each message is known, it can prioritize RPCs that still have less data to send.
The underlying idea is scheduling close to SRPT (Shortest Remaining Processing Time).
By finishing jobs with less remaining work first, it prevents small RPCs from waiting a long time behind bulk transfers.
The receiver uses GRANT packets to tell the sender how much it may send and what priority to assign that data.
Even when data flows into one receiver from multiple senders, the receiver can look at the messages it expects to arrive and adjust which communication proceeds first.
However, even if the receiver decides on priorities, it cannot easily overtake packets that are already queued in large numbers inside a switch.
Homa therefore also uses the multiple priority queues that datacenter switches provide.
Short messages are given high priority and handled in queues separate from bulk transfers.
On the sender side as well, when several messages are ready to send, those with less remaining data are prioritized.
Furthermore, if a large number of packets are pushed ahead into the NIC, short messages that arise later get trapped in the NIC's queue.
To prevent this, Homa has a mechanism called the pacer, which regulates how much data is fed into the NIC.
By combining receiver-side scheduling, sender-side control, and switch priority queues, Homa completes short RPCs as early as possible.
Homa is not a protocol that gives up reliability, either.
The protocol description explains how the receiver detects missing data and requests retransmission.
On the other hand, it does not guarantee that multiple RPCs complete in the same order in which they were started.
For applications where ordering matters, such as reading after a write has finished, the application must manage those dependencies itself.
Because continually prioritizing short messages could leave long messages unable to send indefinitely, Homa also includes a mechanism to allocate a certain amount of bandwidth to older communications.
The 2021 performance gap cannot simply be applied to AI clusters
In a peer-reviewed paper presented at USENIX ATC 2021, Ousterhout evaluated the Linux implementation of Homa on 40 physical machines.
Each node was connected to a single switch over a 25Gbps network and ran Linux 5.4.80.
Using message-size distributions observed at companies such as Google and Facebook, each node sent requests to randomly chosen peers and received responses of the same size.
Section 5.2 of the paper reports that, for short messages under high load, Homa's P99 latency was 19 to 72 times lower than TCP and 7 to 83 times lower than DCTCP.
This comparison has conditions, however.
Network load was set to roughly 80 to 90% of the maximum rate each protocol could sustain, so it was not a comparison in which all protocols were given the same absolute amount of traffic.
In workloads with many short messages, Linux software processing sometimes hit its limit before network bandwidth did.
It therefore cannot be read as saying that Homa is dozens of times faster than TCP on any network.
In simple tests closer to low load, the differences are smaller still.
Table 2 of the paper shows that when a 100-byte request and a 100-byte response were sent one pair at a time in sequence, the round-trip time was 15.1 microseconds for Homa and 23.4 microseconds for TCP.
On the other hand, when 500KB requests and responses were processed sequentially, throughput was 10.0Gbps for Homa and 20.3Gbps for TCP, higher for TCP.
| Test | Homa | TCP |
|---|---|---|
| Round-trip time, 100B request + response | 15.1µs | 23.4µs |
| Sequential transfer, 500KB request + response | 10.0Gbps | 20.3Gbps |
These are results from a single client thread, using the best average from five runs of a five-second test.
Homa's strength is not that it always speeds up simple bulk transfers.
It lies in reducing the chance that small communications get left behind large ones in high-load environments where many RPCs of different sizes coexist.
In tests running multiple RPCs at once, Homa's bulk-transfer throughput rose to about the same level as TCP.
The comparison methodology has also been debated
There has also been ongoing debate about how Homa's performance was evaluated.
Network engineer Ivan Pepelnjak, in a 2023 critique, pointed out that tests sending multiple RPCs over the same TCP connection invite head-of-line blocking and may put TCP at a disadvantage.
He also questioned the lack of comparison with existing technologies such as RDMA and RoCE.
In a published rebuttal, Ousterhout explained that Homa also performed well in the 2018 experiments that used multiple TCP connections.
Those results also appear in Figure 8 of the SIGCOMM 2018 paper.
Ousterhout himself also acknowledges a limitation: because he could not obtain sufficient data on how communication partners are distributed in real datacenters, destinations were chosen uniformly at random for the evaluation.
In addition, the main 2021 performance comparison targeted TCP and DCTCP, and was not a direct comparison under the same conditions as the RDMA-based communication used in today's AI clusters.
There are measurements showing the advantages of Homa's design, but how much it would improve GPU utilization or inference performance for a particular AI service has to be verified using that service's own communication patterns.
A tenfold improvement in P99 latency does not mean the overall AI workload becomes ten times faster.
In 2026, TCP coexistence and the GRANT mechanism were refined
Homa's development has continued since the 2021 paper.
The HomaModule change log shows that in 2026, changes were made concerning coexistence with TCP, supported operating systems, and how message transmission begins.
| Date | Main change | Relevance at adoption |
|---|---|---|
| January 2026 | Introduced homa_qdisc |
Reduces interference when TCP and Homa are used on the same host |
| March 2026 | Backported to RHEL 8 and RHEL 9.5 | Widens supported environments |
| September 2026 | Messages requiring GRANTs are now fully scheduled, and START_MSG was added |
Even long messages can be controlled by the receiver before actual data is sent |
This table summarizes the Homa project's own change log and does not imply commercial adoption or standardization.
The existence of RHEL branches likewise does not mean Red Hat officially supports Homa as a product feature.
The September change in particular calls for care when reading earlier descriptions of Homa.
The 2021 paper and the currently published protocol.md describe a mechanism in which even long messages send an initial portion without the receiver's permission, with the rest controlled by GRANT.
The September 2026 change log, however, says that for messages that need GRANTs, actual data is no longer sent at first; instead the receiver is first notified of the message's existence through START_MSG.
In other words, for long messages the receiver now takes part in scheduling from the very start of data transmission.
Explaining the current implementation based only on the old paper or protocol.md could miss this change.
Using it alongside TCP requires dedicated queue control
When migrating an existing datacenter to Homa step by step, coexisting with TCP is hard to avoid.
But being able to run Homa and TCP on the same host is not the same as both achieving high performance without special configuration.
According to a January 2026 record from the Homa developers, when TCP and Homa were run simultaneously without homa_qdisc in a 100Gbps CloudLab c6620 environment, Homa's P99 for short messages increased about fourfold.
With homa_qdisc enabled, performance reportedly returned to near what Homa achieves alone, even with TCP running alongside.
The installation instructions explain that homa_qdisc has two roles.
One is to prevent long transmit queues from forming inside the NIC, so that small Homa messages are not stuck waiting behind them.
The other is to coordinate transmit queues with other protocols such as TCP, so that their traffic does not interfere with each other excessively.
When evaluating an actual deployment, it is necessary to measure not only Homa-only benchmarks but also with existing TCP traffic running on the same hosts and network.
Using Homa requires more than the kernel
Homa is published as a Linux kernel module.
The current installation instructions state that the main branch has been verified to work on Linux 6.17.8.
However, simply loading the kernel module does not deliver the same performance as in the paper.
To get the most from Homa, you need to tune NIC interrupt settings, Receive Packet Steering (RPS), NIC transmit queues, jumbo frames, and more.
On the top-of-rack switch side, multiple priority queues must also be configured, for example based on DSCP.
To send large messages efficiently, segmentation offloads such as TSO and GSO are also important.
In other words, installing Homa and operating it while maintaining low tail latency are separate tasks.
The state of standardization also deserves attention.
Homa has been assigned IANA IP protocol number 146.
However, obtaining an IP protocol number does not mean Homa has been established as an IETF standard.
The RFC draft published on GitHub also primarily targets datacenter environments with round-trip times of tens of microseconds or less, and is not designed as a protocol for WANs.
Understanding it as a technology to replace TCP across the entire internet would overstate its scope.
gRPC integration and encryption are major adoption hurdles today
On the application side, migrating to Homa would be easier if RPC frameworks could absorb the differences between transport methods.
However, the README of the published gRPC integration, grpc_homa, says that active development stopped in late 2023.
At the time development stopped, the C++ version worked, while Java support was only partial.
In addition, it supports only unencrypted channels.
The Linux kernel implementation of Homa itself was updated in 2026, but the surrounding software has not reached a point where it can be safely integrated into existing gRPC services and used as-is in production.
When evaluating Homa in an AI cluster, benchmarks of the transport protocol alone are not enough.
With the actual application, you need to check the following together:
- P99 latency of small RPCs
- CPU time consumed by network processing
- Time GPUs sit idle waiting for communication to complete
- Performance when used alongside TCP
- How to implement it in production, including encryption
If its advantage holds up under those checks and the RPC framework in use can be maintained over time, Homa could be an option for reducing the time that small RPCs get buried under bulk transfers and GPUs are left waiting on communication.
