Stanford professor pushes Homa to replace TCP in AI clusters

Homa: The end of TCP for AI clusters [video]

A Hacker News discussion of a talk arguing that TCP is the wrong fit for AI clusters and that Homa, a receiver-driven transport protocol, should take its place. The thread links the original USENIX ATC paper by John Ousterhout and related coverage, including a Register piece on the Stanford professor's campaign for a new protocol.

Homa: The end of TCP for AI clusters
  1. Animats

    Homa has been around for a while. Here's the 2018 paper.[1]

    The core idea: When a message arrives at the sender’s transport module, Homa

    divides the message into two parts: an initial unscheduled portion (the first RTTbytes bytes), followed by a scheduled portion.

    The sender transmits the unscheduled bytes immediately, using

    one or more DATA packets. The scheduled bytes are not transmitted until requested explicitly by the receiver using GRANT packets.

    So it sends blind for short requests, then needs a go-ahead from the receiver.

    That's reasonable when the main application is a remote procedure call. It's reminiscent of QNX's networking protocol, which is also single packet message request/response but can also handle arbitrarily long messages.

    What makes this work today is that per-packet processing overhead in hardware switches is low vs. per-byte overhead. In early software driven switches, per-packet overhead tended to dominate, and sending small packets was very inefficient. In modern hardware switches, where FPGAs are doing the processing, the per-packet overhead is low enough that small packets are not inefficient.

    It's amusing that web stuff is so bloated today that any transaction under 1MB is considered "small".

    So this is not a suitable protocol for open web use.

    [1] https://people.csail.mit.edu/alizadeh/papers/homa-sigcomm18....

  2. Veserv

    Homa is not a good design. [1]

    1. No way to detect whole RPC loss. Since there is no outer connection state, if every packet in the send-side of a RPC is lost then there is no way for a server to detect that it should issue a resend. The RPC is just lost to the ether. This affects small messages, like messages that fit in a single packet, more since there are fewer packets in the send-side.

    2. Related to the above, there is no builtin encryption support. So, if you want encryption then you need to layer it either above or below.

    3. Benchmarked performance is awful. The 60 kB average message case in [2] Table 4 takes 5(!) hyperthreads to average 20 Gbit/s. That is just 4 Gbit/s per hyperthread. Even a totally naive one-packet per system call network protocol design and implementation should get to ~8 Gbit/s per hyperthread. 30 Gbit/s per hyperthread is easy with just a little focus on performance.

    4. Despite all the performance design problems in QUIC (though still faster than Homa) it already solves basically every problem Homa is trying to solve in a much cleaner way. Stream IDs correspond to RPC IDs. Stream Max corresponds to Grants. Multiple streams under single Client allows prioritization.

    Except you do not randomly lose entire messages. You can compact small messages into packets. You get more precise RTT time allowing more accurate pacing/congestion calculations. You get builtin encryption. It survives ossified middleboxs. It has multiple ack frames/packets reducing ac […]

More from this day

2026-10-04