-
Type:
Task
-
Resolution: Fixed
-
Priority:
Major - P3
-
Affects Version/s: None
-
Component/s: None
-
DB Integration & Observability
-
Fully Compatible
-
200
-
None
-
None
-
None
-
None
-
None
-
None
-
None
Background
About 76 waits in 47 test files already avoid the counter bug by adding the starting count by hand: fpb->waitForTimesEntered(fpb.initialTimesEntered() + 1). These waits are correct but wordy. Copying the pattern makes it easy to drop the initialTimesEntered() part, and that brings the bug back.
What to do
- Rewrite these waits with the T1 API (waitForOneNewEntry() / waitForNNewEntries(N)). This is a mechanical rewrite: test behavior does not change.
- Update the failpoint docs so the new API is the documented default.
Done when
- No initialTimesEntered() + N waits remain in src/.
- Every touched test builds. A sample of suites passes locally, and an Evergreen patch covers the rest.
- The failpoint docs are updated.
- The PR has break-glass approval. The PR touches test files owned by many teams, so one approver signs off for all of them instead of dozens of separate code-owner reviews.
Details
See attached pr4_pattern_migration.md: commands to find every call site, the cases that need hand edits, known hotspots, and the break-glass justification.
- depends on
-
SERVER-135705 FP Safety 3/5: Sweep absolute waitForTimesEntered literal call sites
-
- In Code Review
-
- is depended on by
-
SERVER-135707 FP Safety 5/5: Restrict absolute failpoint wait API
-
- In Progress
-