-
Type:
Task
-
Resolution: Fixed
-
Priority:
Major - P3
-
Affects Version/s: None
-
Component/s: None
-
DB Integration & Observability
-
Fully Compatible
-
200
-
None
-
None
-
None
-
None
-
None
-
None
-
None
Background
A failpoint's hit counter is process-wide and is never reset between tests in the same unit-test binary. So fpb->waitForTimesEntered(1) ("wait until the failpoint has been hit at least once") returns immediately if an earlier test already hit it, and the wait silently does nothing.
What to do
Add methods to FailPointEnableBlock that count only hits since the block enabled the failpoint:
- waitForOneNewEntry() and waitForNNewEntries(n)
- timesEnteredSinceEnabled(), a non-blocking getter
Today, tests get the correct behavior by writing fpb->waitForTimesEntered(fpb.initialTimesEntered() + 1) by hand. With the new API, the correct pattern becomes the easy one.
Done when
- The new API exists in src/mongo/util/fail_point.h and has unit tests. One test must show that hits from before the block was enabled do not satisfy the wait.
- bazel run +fail_point_test passes.
Details
See attached pr1_relative_wait_api.md: API signatures, why the waits return void, and the 6 unit tests.
- is depended on by
-
SERVER-135704 FP Safety 2/5: Fix no-op failpoint wait in reconfig test
-
- In Code Review
-
- related to
-
SERVER-135699 Absolute waitForTimesEntered(1) barrier is a no-op in StepUpReconfigConcurrentWithForceHeartbeatReconfig reconfig test
-
- Closed
-
-
SERVER-136240 Revisit new FailPointEnableBlock functions (and use of tassert)
-
- Needs Scheduling
-
-
SERVER-136078 Suppress tassert stack trace logging for tasserts expected via ASSERT_TASSERT_CODE
-
- Closed
-