-
Type:
Bug
-
Resolution: Unresolved
-
Priority:
Major - P3
-
None
-
Affects Version/s: 8.0.30
-
Component/s: None
-
None
-
ALL
-
None
-
None
-
None
-
None
-
None
-
None
-
None
Environment
- MongoDB Community 8.0.30 (Debian 12 based image, `mongod --version` gitVersion 360770fbc9b6a893748150d2ab2799e119b9fac9, tcmalloc-google)
- Sharded cluster on Kubernetes (StatefulSets, hosts addressed through a headless Service DNS name)
- CSRS: 3 members
- 2 shards, each PSA (primary + secondary + arbiter)
- 2 mongos
- HDD storage, resource-constrained dev cluster (arbiter memory limit 1Gi, CSRS 2Gi, data 8Gi)
- Same image/topology in a second (production) cluster shows the same pattern at a smaller scale
Symptom
Arbiters were OOMKilled hundreds of times (shard0 arbiter: 520 restarts, shard1 arbiter: 214), CSRS members are OOMKilled repeatedly (22-30 restarts each), and one shard data node was OOMKilled.
All of the excess connections are server-internal awaitable hello streams from `NetworkInterfaceTL-ReplicaSetMonitor-TaskExecutor` (not application drivers). A healthy ReplicaSetMonitor should hold ~1 streaming hello per (monitoring process, target). We see hundreds to thousands per pair.
Evidence
Grouping `db.currentOp({$all: true})` hello ops on the target by client IP + `clientMetadata.driver.name`:
const ops = db.currentOp({"$all": true}).inprog.filter(o => o.command && o.command.hello);
// group by o.client IP and o.clientMetadata.driver.name
Observed (dev cluster), all `driver.name = NetworkInterfaceTL-ReplicaSetMonitor-TaskExecutor/8.0.30`, all `maxAwaitTimeMS: 10000`:
| monitoring process -> target | concurrent hello ops |
|---|---|
| CSRS primary -> shard0 data node A | 6,420 (later 5,827) |
| CSRS primary -> shard1 data node B | 2,885 |
| CSRS primary -> each other CSRS member | ~297 |
| mongos #1 -> shard0 arbiter | 3,142 (arbiter total ~3,360 connections) |
| shard1 data node -> shard0 data node A | 253 |
- These are live streams, not stale sockets: on a target receiving 297 of them, every op had `secs_running` between 5.6 and 6.4 s, i.e. they are being re-armed in lockstep. `serverStatus().connections.exhaustHello` on the target matches the count.
- Restarting the monitoring process clears its streams (mongos restart: arbiter connections 3,360 -> ~220, arbiter RSS 728 MiB -> 535 MiB). When the CSRS member holding ~5.8k streams to data node A was OOMKilled, its streams disappeared (now 3), and another CSRS member is now accumulating instead (595, then a single step of +51 inside one 30 s sample window; flat in the four samples before).
- Growth is step-wise, not linear. Steps coincide with topology events: member OOM/restart, elections (CSRS `term` is currently 5859), a Kubernetes node going NotReady for ~1 min (headless DNS returned "Could not find address", ~800 "RSM monitoring host in expedited mode" log lines within 10 min).
- During a burst on the new CSRS primary, within 8 minutes we logged 763 `Host failed in replica set` (id 4712102) and 762 `RSM received error response` (id 4333222), 761 of which are `NetworkInterfaceExceededTimeLimit` on the streaming hello itself, e.g.:
NetworkInterfaceExceededTimeLimit: Request 758848 timed out, deadline was 2026-10-10T11:56:50.152+00:00,
op was RemoteCommand 758848 -- target:[<shard0-arbiter>:27017] db:admin expDate:2026-10-10T11:48:19.423+00:00
cmd:{ hello: 1, maxAwaitTimeMS: 10000, topologyVersion: { processId: ObjectId('6ac90c1a458c23d74edecda4'), counter: 2 },
internalClient: { minWireVersion: 25, maxWireVersion: 25 } }
plus ~580 `Ending connection due to bad connection status` (CONNPOOL 22566) in the same window.
- Production cluster (same image, far fewer topology events, no OOMs yet): shard0 arbiter has 1,163 inbound connections: 499 from one CSRS member, 419 from another, 136 from a shard data node — same driver name, same pattern.
Hypothesis
When an awaitable/exhaust hello from the streamable RSM fails client-side with NetworkInterfaceExceededTimeLimit (or the host is marked failed / expedited mode kicks in), the RSM starts a new monitoring stream, but the previous exhaust stream is not cancelled and keeps being re-armed by the target. Each topology event therefore leaks a batch of live streams, which accumulate until the monitoring process restarts. The memory cost lands on both sides (each stream = one ingress connection + thread on the target), which is what kills small arbiters first.
We have not found an existing ticket for this. Possibly related changes in the same code path: SERVER-128517 (8.0.28), SERVER-132246 / SERVER-132650 (8.0.30). `minWaitForStreamingHelloMillis` is at its default (500). SERVER-76621 (exhaust command leak in ThreadPoolTaskExecutor, fixed in 7.0) looks similar in shape.
Questions
- Is it expected that one process holds more than one streaming hello per target?
- Is there a parameter to bound/cancel these streams short of restarting the process?
- Would comparing against 8.0.27 (before
SERVER-128517) help? We can test that and upload FTDC / logs on request.
- is related to
-
SERVER-132246 minWaitForStreamingHelloMillis default breaks spec-compliant drivers
-
- Closed
-
-
SERVER-128517 Add configurable minimum timeout for pre-auth streamable hello
-
- Closed
-
-
SERVER-132650 Make `mongos` use the configurable minimum timeout for pre-auth streamable hello
-
- Closed
-
-
SERVER-76621 Thread pool task executor can cause memory leak when handling exhaust command.
-
- Closed
-