-
Type:
Improvement
-
Resolution: Unresolved
-
Priority:
Major - P3
-
None
-
Affects Version/s: None
-
Component/s: None
-
Replication
-
None
-
None
-
None
-
None
-
None
-
None
-
None
Investigate oplog truncation logic that incurs a fixed per-marker overhead. This can become a performance issue in cases where a large marker queue is present. In some cases the truncation thread is truncating in a continuous loop, as seen in some large-scale loads such as HELP-99461 so any logic we run per-marker contributes directly to truncation latency.
Generally right now we re-evaluate whether each truncation marker can be truncated by recomputing the truncation bounds. If we can assume that the truncation point never goes backwards, we should be able to cache it and only recompute when the next marker that we're checking is behind it.
Findings from claude:
- O(n²) drain loop. _hasExcessMarkers sums marker.bytes over the entire deque every call (oplog_truncate_markers.cpp:316-320), called once per truncate iteration. Draining N markers walks the deque N times under _markersMutex. Fix: maintain a running total on push/pop.
- Redundant getPinnedOplog(). Computed once per drain pass (oplog_truncation.cpp:206) then recomputed inside _hasExcessMarkers on every iteration (oplog_truncate_markers.cpp:330-331), which also takes _oplogPinnedByBackupMutex.
- _readEarliestTimestamp on every truncate (wiredtiger_record_store.cpp:1527) — an extra cursor open+read solely to refresh a cache; could be derived from the truncate bound instead of re-fetched.
- checkOplogTruncationBounds re-derives the first record and next-after-marker record every iteration (oplog_truncation.cpp:48,76-77), even though the prior iteration already truncated up to that exact point.
- Redundant re-entry in the drain loop. _deleteExcessDocuments releases/reacquires the global lock and calls awaitHasExcessMarkersOrDead again on success (oplog_cap_maintainer_thread.cpp:212-230,436-438), repeating the O( n ) scan from #2 each time.
Worth flagging, not clearly a bug: the truncate's WUOW being open blocks OplogProvider from shipping past that timestamp (oplog visibility hole, oplog_provider.cpp:618-621) — ties truncate commit latency directly to replication-shipping progress. Not itself fixable cheaply, but it means fixes #2-#6 (which shrink per-truncate latency) also reduce this stall's exposure.
- is related to
-
SERVER-131905 Update the default for oplogSizeMB now that it's not exposed to end users
-
- Closed
-
- related to
-
SERVER-134612 Oplog truncation re-walks every unreclaimed deleted page because each pass starts from a null lower bound
-
- In Code Review
-
-
WT-18638 Audit MongoDB use cases of range truncate under disaggregated storage and document usage expectations
-
- In Code Review
-
-
WT-18651 Resolve the open HELP, AF, and SERVER tickets reporting disaggregated oplog truncation problems
-
- Closed
-