• Storage Engines - Transactions
    • 26.878
    • None
    • None

      Summary

      Under update-heavy workloads (TSBS steady state, majority-lag locust) the dirty cache ratio overshoots the 20% eviction trigger to 25–35% at the tail of every checkpoint. Almost all of the excess is dirty history store content that eviction is not allowed to touch until the checkpoint finishes. This trips the sys-perf ftdc_wt_dirty_ratio_check added by PERF-9081 on master and 8.0; see BF-46531.

      Evidence

      From the FTDC of BFG-3661641 (tsbs_steady_state_high_short, sys-perf master, 38 GB cache):

      • Checkpoint generation 49 ran for about 62 s. In its final 15 s, history store dirty bytes went from 0.02 GB to 7.79 GB (0.1% to 73% of all dirty content) and the total dirty ratio went from 19% to 32%. It fell back to 7% within 5 s of the checkpoint completing.
      • "checkpoint: number of history store pages reconciled" was flat at 0/s for the first 47 s of the checkpoint, then ran at 5,000–7,900 pages/s for the duration of the spike.
      • "eviction server skips dirty pages during a running checkpoint" rose 10–25x (from roughly 200K/s to 2M/s) in the same window.
      • Application inserts into the history store ran at a steady 10K–190K/s throughout. The workload did not change; the spike is the backlog that accumulated before the history store pass.
      • The pattern recurs on every checkpoint in the run (56 cycles sampled: peak dirty ratio 14.6–29.7%, history store share of dirty content 7–82%). The failing cycle was simply the largest.

      Mechanism

      Checkpoint writes the history store in a dedicated pass after the data trees (_checkpoint_hs). That pass uses the normal per-tree sync, which marks the history store tree as syncing for its whole duration. The eviction server skips every dirty page in a syncing tree (_evict_try_queue_page), so no dirty history store page can be evicted until the pass completes.

      The existing safeguards do not close the gap:

      • Urgent history store eviction (__wti_evict_hs_dirty) is a threshold check on history store dirty bytes. It engages only after the overshoot has started, and once the history store pass begins it hits the same syncing wall.
      • Throttling non-history-store eviction while history store content dominates limits how fast new content is generated but does nothing about the existing backlog.

      The overshoot is bounded by the backlog accumulated between the start of the checkpoint and the start of the history store pass, plus the length of that pass. It is not a regression; it is inherent in the current ordering.

      Ruled out

      Checkpointing the history store incrementally, or interleaved with the data trees, is not an option.

      Directions to evaluate

      • Engage urgent history store eviction earlier or more aggressively during the data-tree phase of a checkpoint, so less backlog exists when the history store pass starts.
      • Narrow what "syncing" blocks for the history store: whether eviction must be excluded from the tree for the entire pass, or only around the final write and block-list stabilisation.
      • Any other way to let eviction make progress on dirty history store pages while the pass runs.

      Definition of done

      Includes revisiting the ftdc_wt_dirty_ratio_check threshold with the Product Performance team (currently trigger + 3%, chosen on the assumption that WiredTiger holds near the trigger).

      Related

      BF-46531 (the failures), WT-17964 (identified the same skip-dirty-during-checkpoint backstop as a contributing factor), WT-18469 (eviction versus a running checkpoint in disaggregated storage).

            Assignee:
            [DO NOT USE] Backlog - Storage Engines Team
            Reporter:
            Chenhao Qu
            Votes:
            0 Vote for this issue
            Watchers:
            2 Start watching this issue

              Created:
              Updated: