-
Type:
Task
-
Resolution: Unresolved
-
Priority:
Major - P3
-
None
-
Affects Version/s: None
-
Component/s: Statistics
-
Storage Engines - Transactions
-
158.752
-
SE Transactions - 2026-10-09
-
3
Motivation
The HELP-99633 PIT restore investigation showed that existing dirty-cache gauges cannot explain how much cache a range truncate dirties on a specific table. One no-fix replay batch containing 61 truncateRange commands dirtied 959,671,703 bytes of leaf-page footprint while allocating only 24,448 bytes of slow-path tombstone updates. Recording cumulative statistics at both the WT connection and data-source levels will let us diagnose this cost on the WT table underlying MongoDB's local.oplog.rs, without losing the connection-wide picture. These statistics should also apply to other range-truncated WT tables.
Scope: ten connection and data-source statistics
Add these ten statistics to BOTH the connection and data-source statistic sets; update both at the same action point:
- truncate_leaf_pages_dirtied and truncate_leaf_bytes_dirtied: leaf pages and memory footprint newly charged to dirty cache by range truncate.
- truncate_slow_path_leaf_pages and truncate_slow_path_update_bytes: distinct slow-path leaf visits and tombstone/update bytes allocated by range truncate.
- truncate_fast_deleted_leaf_pages: successfully fast-deleted leaf pages specifically from range truncate. Existing rec_page_delete_fast already counts a broader set of fast page deletions at both levels; document the subset relationship and avoid redundant counting if the two prove equivalent for the intended scope.
- truncate_fast_delete_fallback_pages and truncate_fast_delete_fallback_in_memory: candidates that cannot be fast-deleted, and the subset rejected because their ref is in memory or busy.
- truncate_boundary_leaf_pages_read: leaf pages read while positioning range-truncate boundaries.
- truncate_internal_pages_dirtied and truncate_internal_bytes_dirtied: internal pages and memory footprint newly charged to dirty cache by range truncate.
Use the existing WT_STAT_CONN_DSRC_INCR pattern (as for rec_page_delete_fast and cache_eviction_blocked_uncommitted_truncate) where appropriate. The per-table statistics must be readable using the underlying WT table URI for local.oplog.rs; do not assume that the URI is literally table:local.oplog.rs. This task does not add session statistics, MongoDB Applied op logging or the six first/second slow-path leaf fields from the HELP-99633 diagnostic build.
These are cumulative activity counters, not current dirty-cache gauges. Newly dirtied page bytes are counted on clean-to-dirty transition and are not later subtracted when the page is cleaned or a truncate rolls back. A page may appear in more than one role counter, and in-memory fallback is a subset of total fallback. With concurrent operations, connection and table deltas do not isolate one individual truncate.
Acceptance criteria
- Generate the ten named connection and data-source statistics through the normal WT statistics generator, with descriptions and byte-size flags as applicable; resolve any true duplication with existing fast-delete counters rather than publishing redundant metrics.
- Increment both levels at each relevant range-truncate action, including clean-to-dirty accounting, successful fast deletion, and in-memory fallback; avoid attributing unrelated modifications to range truncate.
- Cover slow-path, fast-delete, fallback and boundary-positioning scenarios. Test connection and per-table deltas on two tables so a truncate on one does not change the other table's counters.
- Document how to read the oplog WT table's statistics alongside connection statistics and dirty-cache gauges, including the counter-overlap and concurrency limitations.
Related: HELP-99633 is the investigation motivating these counters; WT-18637 covers fast-truncate performance and cache-pressure benchmarking.
- is related to
-
WT-18637 Measure fast truncate deletion performance and cache pressure impact
-
- In Code Review
-