-
Type:
Task
-
Resolution: Unresolved
-
Priority:
Major - P3
-
None
-
Affects Version/s: None
-
Component/s: None
-
None
-
Networking & Observability
-
None
-
None
-
None
-
None
-
None
-
None
-
None
Summary
Add a transport-layer capability to set SO_SNDBUF / SO_RCVBUF on a specific connection's socket, usable by replication for both the client side (the secondary's oplog-fetch egress connection) and the server side (the sync source's ingress connection serving that fetcher). This is the networking primitive underneath SERVER-131936 (replication startup parameters) and unblocks fixing cross-region oplog-fetch replication lag, which is bandwidth-delay-product (BDP) limited on high-RTT links.
Background
On a multi-region replica set, a far-region secondary's single oplog-fetch TCP stream is throttled by the receive window (≈ socket buffer ÷ RTT). At the default net.ipv4.tcp_rmem cap (~6 MiB) and ~60–120 ms RTT the stream tops out well below the primary's oplog production rate, so the secondary falls behind linearly. Enlarging the socket buffers to the BDP removes the ceiling. The transport layer currently has no supported way to set these options on a chosen connection after/while it is established.
Scope / Acceptance criteria
- A transport::Session method to set SO_SNDBUF/SO_RCVBUF on the session's socket (no-op for session types without a settable TCP socket), implemented on the ASIO session.
- Egress (client / oplog fetcher): the receive buffer must be applied before the TCP handshake, because the receive-window scale factor is negotiated in the SYN — setting SO_RCVBUF post-connect does not raise the effective window. Thread an optional receive-buffer size through the egress connect path (TransportLayer::connect → AsioTransportLayer::_doSyncConnect, applied after socket.open() and before connect()), and carry it on DBClientConnection so autoreconnect re-applies it.
- Ingress (server / serving thread): the send buffer is set on the live serving socket. Note the socket is moved into the SSL stream on TLS handshake, so the raw _socket member is closed; the setter must use the accessor that returns the live fd (getSocket() → _sslSocket->lowest_layer() under TLS), guarded by the ssl-socket lock. (SO_SNDBUF does not need to be pre-handshake.)
- Both directions safe under requireTLS.
Technical notes
- SO_RCVBUF/SO_SNDBUF are clamped by net.core.rmem_max/net.core.wmem_max (raising those is a deployment prerequisite). setsockopt(SO_RCVBUF) also disables receive autotuning for the socket's lifetime, so reverting to autotuning requires a fresh connection.
- TCP window scaling (RFC 7323, on by default) is required for windows > 64 KiB.
- related to
-
SERVER-131936 Startup parameters to size the TCP window of replication connections
-
- Backlog
-