-
Type:
Bug
-
Resolution: Won't Do
-
Priority:
Major - P3
-
None
-
Affects Version/s: None
-
Component/s: None
-
DB Integration & Observability
-
ALL
-
200
-
None
-
None
-
None
-
None
-
None
-
None
-
None
Origin
Adjacent-risk hunt from the BF-46466 postmortem — SUSPECTED, NOT YET CONFIRMED. Confidence: high. Related: [BF-46466].
Summary
advanceUntilReadyRequest in src/mongo/executor/async_rpc_test_fixture.h:173-187 holds NetworkInterfaceMock::InNetworkGuard across the entire {{while (!net->hasReadyRequests())
{ net->advanceTime(net->now()+10ms); sleep_for(100us); }}} loop. The retry re-dispatch (startCommand) must run on the executor thread, which needs the same network mutex the main thread is still holding -> hasReadyRequests() cannot become true from inside the loop. The only exit is a request already enqueued in the pre-guard sleep_for(1ms) window — timing-dependent. For the retry-dispatch case this is a genuine near-deadlock, not a benign race.
The sibling fixture already documents the fix: ShardingTestFixtureCommon::advanceUntilReadyRequest (src/mongo/db/sharding_environment/sharding_test_fixture_common.cpp:164-188) releases the guard each iteration with the comment "We must release the InNetworkGuard on each iteration to allow the executor thread to run callbacks."
Scope
The tight form in async_rpc_test_fixture.h:173-187 (used by every retry/backoff test in async_rpc_test.cpp: TargeterDeprioritizedServer, RetryOnSuccessfulHelloAdditionalAttempts, DynamicDelayBetweenRetries, BaseBackoffMSExtractedFromRemoteError) and the identical skeleton in async_multicaster_test.cpp:88-100. Plus a broader family one edit away (see Investigation).
Acceptance Criteria
- advanceUntilReadyRequest releases the InNetworkGuard on each loop iteration (port the sharding_test_fixture_common.cpp:164-188 pattern), so retry re-dispatch is not blocked by the held guard.
- The retry/backoff tests in async_rpc_test.cpp and async_multicaster_test.cpp pass under the released-guard form (confirming the fix does not regress them).
- A scan/lint flags InNetworkGuard alive across a while(...)/advanceTime/advance( loop.
Verification (confirm/deny)
Port the sibling's release-per-iteration pattern into the async-rpc fixture and run the listed retry tests. Hang or failure confirms S3; unchanged-pass leaves it latent (with the same reading — the guard is still the hazard, just not currently triggered by timing).
Investigation
- Sweep for the same "wrap-in-guard / advance-in-loop / never release" skeleton in the broader family: server_ping_monitor_test.cpp, remote_command_retry_scheduler_test.cpp, server_discovery_monitor_test.cpp, replication_coordinator_impl_test.cpp, replication_coordinator_disagg_storage_test.cpp, transaction_coordinator_{{futures_util,service_test.cpp}}, transaction_coordinator_test.cpp, transaction_router_test.cpp. Fold any hits into this fix or file per-team.
- Confirm whether any of these rely on the guard being held for correctness (vs. just convenience).