ExportXMLWordPrintableJSON

    • Type: Bug
    • Resolution: Won't Do
    • Priority: Major - P3
    • None
    • Affects Version/s: None
    • Component/s: None
    • DB Integration & Observability
    • ALL
    • 200
    • None
    • None
    • None
    • None
    • None
    • None
    • None

      Origin

      Adjacent-risk hunt from the BF-46466 postmortem — SUSPECTED, NOT YET CONFIRMED. Confidence: high. Related: [BF-46466].

      Summary

      advanceUntilReadyRequest in src/mongo/executor/async_rpc_test_fixture.h:173-187 holds NetworkInterfaceMock::InNetworkGuard across the entire {{while (!net->hasReadyRequests())

      { net->advanceTime(net->now()+10ms); sleep_for(100us); }

      }} loop. The retry re-dispatch (startCommand) must run on the executor thread, which needs the same network mutex the main thread is still holding -> hasReadyRequests() cannot become true from inside the loop. The only exit is a request already enqueued in the pre-guard sleep_for(1ms) window — timing-dependent. For the retry-dispatch case this is a genuine near-deadlock, not a benign race.

      The sibling fixture already documents the fix: ShardingTestFixtureCommon::advanceUntilReadyRequest (src/mongo/db/sharding_environment/sharding_test_fixture_common.cpp:164-188) releases the guard each iteration with the comment "We must release the InNetworkGuard on each iteration to allow the executor thread to run callbacks."

      Scope

      The tight form in async_rpc_test_fixture.h:173-187 (used by every retry/backoff test in async_rpc_test.cpp: TargeterDeprioritizedServer, RetryOnSuccessfulHelloAdditionalAttempts, DynamicDelayBetweenRetries, BaseBackoffMSExtractedFromRemoteError) and the identical skeleton in async_multicaster_test.cpp:88-100. Plus a broader family one edit away (see Investigation).

      Acceptance Criteria

      • advanceUntilReadyRequest releases the InNetworkGuard on each loop iteration (port the sharding_test_fixture_common.cpp:164-188 pattern), so retry re-dispatch is not blocked by the held guard.
      • The retry/backoff tests in async_rpc_test.cpp and async_multicaster_test.cpp pass under the released-guard form (confirming the fix does not regress them).
      • A scan/lint flags InNetworkGuard alive across a while(...)/advanceTime/advance( loop.

      Verification (confirm/deny)

      Port the sibling's release-per-iteration pattern into the async-rpc fixture and run the listed retry tests. Hang or failure confirms S3; unchanged-pass leaves it latent (with the same reading — the guard is still the hazard, just not currently triggered by timing).

      Investigation

      • Sweep for the same "wrap-in-guard / advance-in-loop / never release" skeleton in the broader family: server_ping_monitor_test.cpp, remote_command_retry_scheduler_test.cpp, server_discovery_monitor_test.cpp, replication_coordinator_impl_test.cpp, replication_coordinator_disagg_storage_test.cpp, transaction_coordinator_{{futures_util,service_test.cpp}}, transaction_coordinator_test.cpp, transaction_router_test.cpp. Fold any hits into this fix or file per-team.
      • Confirm whether any of these rely on the guard being held for correctness (vs. just convenience).

            Assignee:
            Charlie Swanson
            Reporter:
            Charlie Swanson
            Votes:
            0 Vote for this issue
            Watchers:
            1 Start watching this issue

              Created:
              Updated:
              Resolved: