-
Type:
Bug
-
Resolution: Duplicate
-
Priority:
Major - P3
-
None
-
Affects Version/s: None
-
Component/s: None
-
DB Integration & Observability
-
ALL
-
200
-
None
-
None
-
None
-
None
-
None
-
None
-
None
Origin
Adjacent-risk hunt from the BF-46466 postmortem — SUSPECTED, NOT YET CONFIRMED. Confidence: highest of the sibling risks. Related: [BF-46466], SERVER-133148 (which established the correct relative-count pattern).
Summary
In src/mongo/db/repl/replication_coordinator_impl_reconfig_test.cpp, the test StepUpReconfigConcurrentWithForceHeartbeatReconfig (line 2478) calls fpb->waitForTimesEntered(1) at line 2554 to assert that a heartbeat reconfig is "in progress, stuck at the failpoint" (per the comment at 2552-2553). Because the blockHeartbeatReconfigFinish failpoint's _hitCount persists across tests in the binary and is never reset, and waitForTimesEntered is an absolute comparison, the wait returns immediately — the barrier is a no-op. The drain/clock steps at 2556-2567 can therefore run before the current reconfig actually reaches the failpoint.
This is the exact bug class SERVER-133148 hardened the BRM tests against by using the relative form waitForTimesEntered(initialTimesEntered() + 1) (blocking_results_merger_test.cpp:796-802, 877-879).
Scope
The single absolute waitForTimesEntered(1) at replication_coordinator_impl_reconfig_test.cpp:2554 (and a sweep for any other absolute waitForTimesEntered(<literal>) across the repo). Failpoint hit counts are process-global and cumulative (fail_point.cpp:88-124; shouldFail() increments _hitCount as a side effect at fail_point.h:506-508; FailPointEnableBlock does not reset it).
Acceptance Criteria
- The barrier at line 2554 actually waits for the current test's reconfig to reach blockHeartbeatReconfigFinish, not for some earlier entry.
- No remaining absolute waitForTimesEntered(<literal>) barrier in the repo whose count can be pre-satisfied by an earlier test.
- The converted test is verified not to hang under the new barrier (see Verification).
Verification (confirm/deny — cheap, do this first)
- Confirm the no-op (~1 min): at the test's start (~line 2545) capture and print fpb->initialTimesEntered(). On a normal binary run it will be >0 (the earlier test NodeReturnsConfigurationInProgressWhenReceivingAReconfigWhileInTheMidstOfAHeartbeatReconfig at lines 674-735 drives the failpoint alwaysOn at 690 / off at 728, leaving the count >=1; the target test's own net->runReadyNetworkOperations() at 2548-2550 is another candidate). Then show waitForTimesEntered(1) at 2554 returns instantly. (The denial direction — asserting initialTimesEntered() == 0 — will fail, proving the bug.)
- Stronger diagnostic: add a chunk-level assert that the count increased during this test's window (between its setMode and the drain steps at 2556-2567); it will flake/fail.
Fix direction (with a caveat)
Convert to the relative form: fpb->waitForTimesEntered(fpb.initialTimesEntered() + 1).
Caveat (do not land as a drive-by one-liner): the target test's reconfig can be scheduled asynchronously on the executor (reschedule path at replication_coordinator_impl_heartbeat.cpp:917-922). With the relative form the barrier now genuinely waits, and whether it hangs depends on whether the test's runReadyNetworkOperations synchronously drives the reconfig through the failpoint. Run the converted test N times verifying it does not block; if it blocks, the test needs the worker-driven shape, not just a counter change.
Investigation
- Confirm which earlier test(s) in the binary enter blockHeartbeatReconfigFinish and whether the target test's own pump can satisfy the absolute count on its own (affects the fix's hang risk).
- Decide whether a per-test failpoint-count reset (or a fixture-level reset) is wanted in addition to the relative-count conversion (postmortem Open Question 5).
- duplicates
-
SERVER-135704 FP Safety 2/5: Fix no-op failpoint wait in reconfig test
-
- In Code Review
-
- is related to
-
SERVER-135704 FP Safety 2/5: Fix no-op failpoint wait in reconfig test
-
- In Code Review
-
-
SERVER-135707 FP Safety 5/5: Restrict absolute failpoint wait API
-
- In Progress
-
-
SERVER-133148 Avoid using semi-initialized opCtx when retrying getMores
-
- Closed
-
-
SERVER-135703 FP Safety 1/5: Add relative-wait API to FailPointEnableBlock
-
- Closed
-
- related to
-
SERVER-133148 Avoid using semi-initialized opCtx when retrying getMores
-
- Closed
-