-
Type:
Bug
-
Resolution: Unresolved
-
Priority:
Major - P3
-
None
-
Affects Version/s: None
-
Component/s: None
-
DB Integration & Observability
-
ALL
-
200
-
None
-
None
-
None
-
None
-
None
-
None
-
None
Background
A failpoint's hit counter is process-wide and is never reset between tests in the same unit-test binary. So fpb->waitForTimesEntered(1) returns immediately if an earlier test already hit it, and the wait silently does nothing.
What to do
In replication_coordinator_impl_reconfig_test.cpp, test StepUpReconfigConcurrentWithForceHeartbeatReconfig has exactly this bug. An earlier test in the same binary hits blockHeartbeatReconfigFinish, so the wait passes immediately. The test passes today only because of timing. Change the wait to fpb.waitForOneNewEntry() (from T1 / FP Safety 1/5).
Done when
- The wait uses the relative API.
- An Evergreen patch that runs this test many times is green, which shows that the now-real wait doesn't hang. Put the patch link in the PR description so reviewers can check it.
Details
See attached pr2_reconfig_test_fix.md: why the wait is a no-op today, the hang risk once it's real, and what to do if the patch hangs.
- depends on
-
SERVER-135703 FP Safety 1/5: Add relative-wait API to FailPointEnableBlock
-
- Closed
-
- is depended on by
-
SERVER-135705 FP Safety 3/5: Sweep absolute waitForTimesEntered literal call sites
-
- In Code Review
-
- is duplicated by
-
SERVER-135699 Absolute waitForTimesEntered(1) barrier is a no-op in StepUpReconfigConcurrentWithForceHeartbeatReconfig reconfig test
-
- Closed
-
- related to
-
SERVER-135699 Absolute waitForTimesEntered(1) barrier is a no-op in StepUpReconfigConcurrentWithForceHeartbeatReconfig reconfig test
-
- Closed
-