-
Type:
Task
-
Resolution: Unresolved
-
Priority:
Major - P3
-
None
-
Affects Version/s: None
-
Component/s: None
-
None
-
Networking & Observability
-
None
-
None
-
None
-
None
-
None
-
None
-
None
Context
As part of the lock-free serverStatus effort (SPM-4755), slow serverStatus sections contribute to FTDC gaps when collection exceeds 1 second. Analysis in SERVER-124926 and the slow serverStatus spreadsheet shows that the oplogTruncation section is a significant contributor. The spreadsheet data was collected from live Atlas clusters.
Owning team
Suggested owner (from CODEOWNERS): Oplog (@10gen/server-oplog)
Impact (from slow serverStatus logs)
- Tab: union (mongod + mongos)
- Total duration share: 10.2% of slow-run section time
- Avg duration when slow: 367 ms
- Rows where section exceeded 1s: 0.4% of slow runs
Code location
- Section registration: ServerStatusSectionBuilder<OplogTruncateMarkersServerStatusSection>("oplogTruncation")
- Primary file: src/mongo/db/storage/oplog_cap_maintainer_thread.cpp
Task
Please audit the oplogTruncation serverStatus section and remove any unnecessary blocking work (locks, synchronous I/O, contended atomics, etc.).
Use SKUNK-40 branch for reference implementations and SKUNK-40-nonblocking for nonblocking annotations.
Acceptance criteria
- oplogTruncation section generation does not acquire blocking locks on the hot path
- No regression in section output for FTDC/default serverStatus collection
- Add or update tests if behavior changes
Related work
- Parent analysis: SERVER-124926
- Similar tickets: SERVER-127312 (oplog)
- is related to
-
SERVER-127312 Remove blocking work on the oplog server status section
-
- Needs Scheduling
-