-
Type:
Task
-
Resolution: Unresolved
-
Priority:
Major - P3
-
None
-
Affects Version/s: None
-
Component/s: Cache and Eviction
-
Storage Engines - Transactions
-
26.878
-
None
-
None
Summary
Under update-heavy workloads (TSBS steady state, majority-lag locust) the dirty cache ratio overshoots the 20% eviction trigger to 25–35% at the tail of every checkpoint. Almost all of the excess is dirty history store content that eviction is not allowed to touch until the checkpoint finishes. This trips the sys-perf ftdc_wt_dirty_ratio_check added by PERF-9081 on master and 8.0; see BF-46531.
Evidence
From the FTDC of BFG-3661641 (tsbs_steady_state_high_short, sys-perf master, 38 GB cache):
- Checkpoint generation 49 ran for about 62 s. In its final 15 s, history store dirty bytes went from 0.02 GB to 7.79 GB (0.1% to 73% of all dirty content) and the total dirty ratio went from 19% to 32%. It fell back to 7% within 5 s of the checkpoint completing.
- "checkpoint: number of history store pages reconciled" was flat at 0/s for the first 47 s of the checkpoint, then ran at 5,000–7,900 pages/s for the duration of the spike.
- "eviction server skips dirty pages during a running checkpoint" rose 10–25x (from roughly 200K/s to 2M/s) in the same window.
- Application inserts into the history store ran at a steady 10K–190K/s throughout. The workload did not change; the spike is the backlog that accumulated before the history store pass.
- The pattern recurs on every checkpoint in the run (56 cycles sampled: peak dirty ratio 14.6–29.7%, history store share of dirty content 7–82%). The failing cycle was simply the largest.
Mechanism
Checkpoint writes the history store in a dedicated pass after the data trees (_checkpoint_hs). That pass uses the normal per-tree sync, which marks the history store tree as syncing for its whole duration. The eviction server skips every dirty page in a syncing tree (_evict_try_queue_page), so no dirty history store page can be evicted until the pass completes.
The existing safeguards do not close the gap:
- Urgent history store eviction (__wti_evict_hs_dirty) is a threshold check on history store dirty bytes. It engages only after the overshoot has started, and once the history store pass begins it hits the same syncing wall.
- Throttling non-history-store eviction while history store content dominates limits how fast new content is generated but does nothing about the existing backlog.
The overshoot is bounded by the backlog accumulated between the start of the checkpoint and the start of the history store pass, plus the length of that pass. It is not a regression; it is inherent in the current ordering.
Ruled out
Checkpointing the history store incrementally, or interleaved with the data trees, is not an option.
Directions to evaluate
- Engage urgent history store eviction earlier or more aggressively during the data-tree phase of a checkpoint, so less backlog exists when the history store pass starts.
- Narrow what "syncing" blocks for the history store: whether eviction must be excluded from the tree for the entire pass, or only around the final write and block-list stabilisation.
- Any other way to let eviction make progress on dirty history store pages while the pass runs.
Definition of done
Includes revisiting the ftdc_wt_dirty_ratio_check threshold with the Product Performance team (currently trigger + 3%, chosen on the assumption that WiredTiger holds near the trigger).
Related
BF-46531 (the failures), WT-17964 (identified the same skip-dirty-during-checkpoint backstop as a contributing factor), WT-18469 (eviction versus a running checkpoint in disaggregated storage).
- is related to
-
WT-18469 Investigate the trade-off of letting eviction reconcile ahead of a running checkpoint in disaggregated storage
-
- Open
-
-
WT-17964 Understand and fix change in dirty cache management during ycsb 95/5 workload
-
- Closed
-
-
WT-18768 perf-atlas-M30-real 2 GiB cache too small for YCSB 95/5 working set, causing severe write-latency stalls
-
- Needs Scheduling
-
- related to
-
WT-18768 perf-atlas-M30-real 2 GiB cache too small for YCSB 95/5 working set, causing severe write-latency stalls
-
- Needs Scheduling
-