• Type: Bug
    • Resolution: Unresolved
    • Priority: Major - P3
    • None
    • Affects Version/s: 8.0.30
    • Component/s: None
    • None
    • ALL
    • None
    • None
    • None
    • None
    • None
    • None
    • None

      Environment

      • MongoDB Community 8.0.30 (Debian 12 based image, `mongod --version` gitVersion 360770fbc9b6a893748150d2ab2799e119b9fac9, tcmalloc-google)
      • Sharded cluster on Kubernetes (StatefulSets, hosts addressed through a headless Service DNS name)
        • CSRS: 3 members
        • 2 shards, each PSA (primary + secondary + arbiter)
        • 2 mongos
      • HDD storage, resource-constrained dev cluster (arbiter memory limit 1Gi, CSRS 2Gi, data 8Gi)
      • Same image/topology in a second (production) cluster shows the same pattern at a smaller scale

      Symptom

      Arbiters were OOMKilled hundreds of times (shard0 arbiter: 520 restarts, shard1 arbiter: 214), CSRS members are OOMKilled repeatedly (22-30 restarts each), and one shard data node was OOMKilled.

      All of the excess connections are server-internal awaitable hello streams from `NetworkInterfaceTL-ReplicaSetMonitor-TaskExecutor` (not application drivers). A healthy ReplicaSetMonitor should hold ~1 streaming hello per (monitoring process, target). We see hundreds to thousands per pair.

      Evidence

      Grouping `db.currentOp({$all: true})` hello ops on the target by client IP + `clientMetadata.driver.name`:

      const ops = db.currentOp({"$all": true}).inprog.filter(o => o.command && o.command.hello);
      // group by o.client IP and o.clientMetadata.driver.name
      

      Observed (dev cluster), all `driver.name = NetworkInterfaceTL-ReplicaSetMonitor-TaskExecutor/8.0.30`, all `maxAwaitTimeMS: 10000`:

      monitoring process -> target concurrent hello ops
      CSRS primary -> shard0 data node A 6,420 (later 5,827)
      CSRS primary -> shard1 data node B 2,885
      CSRS primary -> each other CSRS member ~297
      mongos #1 -> shard0 arbiter 3,142 (arbiter total ~3,360 connections)
      shard1 data node -> shard0 data node A 253
      • These are live streams, not stale sockets: on a target receiving 297 of them, every op had `secs_running` between 5.6 and 6.4 s, i.e. they are being re-armed in lockstep. `serverStatus().connections.exhaustHello` on the target matches the count.
      • Restarting the monitoring process clears its streams (mongos restart: arbiter connections 3,360 -> ~220, arbiter RSS 728 MiB -> 535 MiB). When the CSRS member holding ~5.8k streams to data node A was OOMKilled, its streams disappeared (now 3), and another CSRS member is now accumulating instead (595, then a single step of +51 inside one 30 s sample window; flat in the four samples before).
      • Growth is step-wise, not linear. Steps coincide with topology events: member OOM/restart, elections (CSRS `term` is currently 5859), a Kubernetes node going NotReady for ~1 min (headless DNS returned "Could not find address", ~800 "RSM monitoring host in expedited mode" log lines within 10 min).
      • During a burst on the new CSRS primary, within 8 minutes we logged 763 `Host failed in replica set` (id 4712102) and 762 `RSM received error response` (id 4333222), 761 of which are `NetworkInterfaceExceededTimeLimit` on the streaming hello itself, e.g.:
      NetworkInterfaceExceededTimeLimit: Request 758848 timed out, deadline was 2026-10-10T11:56:50.152+00:00,
      op was RemoteCommand 758848 -- target:[<shard0-arbiter>:27017] db:admin expDate:2026-10-10T11:48:19.423+00:00
      cmd:{ hello: 1, maxAwaitTimeMS: 10000, topologyVersion: { processId: ObjectId('6ac90c1a458c23d74edecda4'), counter: 2 },
      internalClient: { minWireVersion: 25, maxWireVersion: 25 } }
      

      plus ~580 `Ending connection due to bad connection status` (CONNPOOL 22566) in the same window.

      • Production cluster (same image, far fewer topology events, no OOMs yet): shard0 arbiter has 1,163 inbound connections: 499 from one CSRS member, 419 from another, 136 from a shard data node — same driver name, same pattern.

      Hypothesis

      When an awaitable/exhaust hello from the streamable RSM fails client-side with NetworkInterfaceExceededTimeLimit (or the host is marked failed / expedited mode kicks in), the RSM starts a new monitoring stream, but the previous exhaust stream is not cancelled and keeps being re-armed by the target. Each topology event therefore leaks a batch of live streams, which accumulate until the monitoring process restarts. The memory cost lands on both sides (each stream = one ingress connection + thread on the target), which is what kills small arbiters first.

      We have not found an existing ticket for this. Possibly related changes in the same code path: SERVER-128517 (8.0.28), SERVER-132246 / SERVER-132650 (8.0.30). `minWaitForStreamingHelloMillis` is at its default (500). SERVER-76621 (exhaust command leak in ThreadPoolTaskExecutor, fixed in 7.0) looks similar in shape.

      Questions

      1. Is it expected that one process holds more than one streaming hello per target?
      2. Is there a parameter to bound/cancel these streams short of restarting the process?
      3. Would comparing against 8.0.27 (before SERVER-128517) help? We can test that and upload FTDC / logs on request.

            Assignee:
            Unassigned
            Reporter:
            JAEMOO HAN (EXT)
            Votes:
            0 Vote for this issue
            Watchers:
            1 Start watching this issue

              Created:
              Updated: