Nova Notes

Tuning TCP buffers for long fat networks

· Clara Lehto · 6 min read

Last month we moved a nightly backup job between two data centers roughly 40 ms apart. Both servers had 10 Gbit uplinks, iperf between them showed plenty of headroom with parallel streams, yet a single rsync never went above 300 Mbit/s. The culprit was not the disks or the network — it was the TCP window.

The bandwidth-delay product

TCP can only have as much unacknowledged data in flight as the receive window allows. To keep a link full, the window has to be at least as large as the bandwidth-delay product (BDP):

BDP = bandwidth × round-trip time
    = 10 Gbit/s × 0.040 s
    = 400 Mbit ≈ 50 MB

Most distributions ship with a maximum receive buffer of 6 MB. With a 40 ms RTT that caps a single flow at about 1.2 Gbit/s in theory, and noticeably less in practice because the kernel uses part of the buffer for its own bookkeeping.

The three settings that matter

You do not need to touch dozens of knobs. On modern kernels autotuning works well once it is allowed to grow far enough:

# /etc/sysctl.d/90-tcp-buffers.conf
net.core.rmem_max = 67108864
net.core.wmem_max = 67108864
net.ipv4.tcp_rmem = 4096 131072 67108864
net.ipv4.tcp_wmem = 4096 65536 67108864

The three numbers in tcp_rmem and tcp_wmem are the minimum, default and maximum buffer size per socket. Only the maximum really needs to change; leave the default alone so that thousands of idle connections do not eat your memory.

Congestion control

Bigger buffers expose another problem: classic loss-based congestion control (CUBIC) reacts badly to the occasional dropped packet on a long path. Switching to BBR made the transfer far more stable for us:

net.core.default_qdisc = fq
net.ipv4.tcp_congestion_control = bbr

Measuring the result

Apply the settings with sysctl --system and test with a single stream, because parallel streams hide exactly this problem:

iperf3 -c backup-dst -t 30 -P 1

Our single-stream throughput went from 290 Mbit/s to a little over 3.1 Gbit/s, at which point the disks became the bottleneck — which is where you want to be.

Rule of thumb: if a single TCP flow is slow but parallel flows are fast, look at window sizes before you look at anything else.

Further reading

  • The ip-sysctl.txt documentation in the kernel source tree
  • ss -tin — shows the current congestion window and RTT per connection