ExportXMLWordPrintableJSON

    • Type: Bug
    • Resolution: Duplicate
    • Priority: Major - P3
    • None
    • Affects Version/s: None
    • Component/s: None
    • DB Integration & Observability
    • ALL
    • 200
    • None
    • None
    • None
    • None
    • None
    • None
    • None

      Origin

      Adjacent-risk hunt from the BF-46466 postmortem — SUSPECTED, NOT YET CONFIRMED. Confidence: highest of the sibling risks. Related: [BF-46466], SERVER-133148 (which established the correct relative-count pattern).

      Summary

      In src/mongo/db/repl/replication_coordinator_impl_reconfig_test.cpp, the test StepUpReconfigConcurrentWithForceHeartbeatReconfig (line 2478) calls fpb->waitForTimesEntered(1) at line 2554 to assert that a heartbeat reconfig is "in progress, stuck at the failpoint" (per the comment at 2552-2553). Because the blockHeartbeatReconfigFinish failpoint's _hitCount persists across tests in the binary and is never reset, and waitForTimesEntered is an absolute comparison, the wait returns immediately — the barrier is a no-op. The drain/clock steps at 2556-2567 can therefore run before the current reconfig actually reaches the failpoint.

      This is the exact bug class SERVER-133148 hardened the BRM tests against by using the relative form waitForTimesEntered(initialTimesEntered() + 1) (blocking_results_merger_test.cpp:796-802, 877-879).

      Scope

      The single absolute waitForTimesEntered(1) at replication_coordinator_impl_reconfig_test.cpp:2554 (and a sweep for any other absolute waitForTimesEntered(<literal>) across the repo). Failpoint hit counts are process-global and cumulative (fail_point.cpp:88-124; shouldFail() increments _hitCount as a side effect at fail_point.h:506-508; FailPointEnableBlock does not reset it).

      Acceptance Criteria

      • The barrier at line 2554 actually waits for the current test's reconfig to reach blockHeartbeatReconfigFinish, not for some earlier entry.
      • No remaining absolute waitForTimesEntered(<literal>) barrier in the repo whose count can be pre-satisfied by an earlier test.
      • The converted test is verified not to hang under the new barrier (see Verification).

      Verification (confirm/deny — cheap, do this first)

      1. Confirm the no-op (~1 min): at the test's start (~line 2545) capture and print fpb->initialTimesEntered(). On a normal binary run it will be >0 (the earlier test NodeReturnsConfigurationInProgressWhenReceivingAReconfigWhileInTheMidstOfAHeartbeatReconfig at lines 674-735 drives the failpoint alwaysOn at 690 / off at 728, leaving the count >=1; the target test's own net->runReadyNetworkOperations() at 2548-2550 is another candidate). Then show waitForTimesEntered(1) at 2554 returns instantly. (The denial direction — asserting initialTimesEntered() == 0 — will fail, proving the bug.)
      2. Stronger diagnostic: add a chunk-level assert that the count increased during this test's window (between its setMode and the drain steps at 2556-2567); it will flake/fail.

      Fix direction (with a caveat)

      Convert to the relative form: fpb->waitForTimesEntered(fpb.initialTimesEntered() + 1).

      Caveat (do not land as a drive-by one-liner): the target test's reconfig can be scheduled asynchronously on the executor (reschedule path at replication_coordinator_impl_heartbeat.cpp:917-922). With the relative form the barrier now genuinely waits, and whether it hangs depends on whether the test's runReadyNetworkOperations synchronously drives the reconfig through the failpoint. Run the converted test N times verifying it does not block; if it blocks, the test needs the worker-driven shape, not just a counter change.

      Investigation

      • Confirm which earlier test(s) in the binary enter blockHeartbeatReconfigFinish and whether the target test's own pump can satisfy the absolute count on its own (affects the fix's hang risk).
      • Decide whether a per-test failpoint-count reset (or a fixture-level reset) is wanted in addition to the relative-count conversion (postmortem Open Question 5).
      {info:title=Team note}Filed to Query Integration as a placeholder owner per the postmortem; the owning team is Replication. Please reassign Assigned Teams on triage.{info}

            Assignee:
            Charlie Swanson
            Reporter:
            Charlie Swanson
            Votes:
            0 Vote for this issue
            Watchers:
            1 Start watching this issue

              Created:
              Updated:
              Resolved: