-
Type:
Improvement
-
Resolution: Unresolved
-
Priority:
Major - P3
-
None
-
Affects Version/s: None
-
Component/s: Schema Management
-
Storage Engines - Foundations
-
560.828
-
SE Foundations - 2026-09-15
-
5
Problem:
After WT-18428, WT_SESSION::drop still performs the file unlink and the history store truncate inside the call and the truncate still runs under the schema lock. The caller pays for both and the truncate holds the lock for the duration of a history store range truncate.
Deferring the truncate exposes a pre-existing crash window that must be closed as part of this change. A drop makes its metadata removal durable before it truncates the dropped btree's history store content. A crash between the two leaves that content in the history store; rollback-to-stable does not remove it (RTS truncates history only for live non-timestamped tables and partial-backup remove ids). Recovery resets next_file_id to the metadata maximum, so the next table created reuses the dropped btree id and inherits its history. Reproduced with a crash point after the metadata checkpoint: table id 16 with 2000 history store records, drop crashes, restart, create gets id 16 and writes at timestamp 40; reads at timestamps 10 and 20 return 1000 rows each of the dropped table's values, timestamp 30 returns 0 rows, timestamp 40 returns the new values. Today the window is microseconds; deferring the truncate widens it to the sweep latency.
Solution:
- Durable record. __drop_file inserts, in the same tracked metadata transaction as the file: removal, system:drop.<original>.<btree_id> with value original=,renamed=[,btree_id=] whenever a rename or a history store truncate is scheduled (btree_id present only when a truncate was decided). The truncate is tracked as part of WT_ST_DROP_COMMIT; WT_ST_HS_TRUNCATE is removed. One pending entry per drop carries the renamed file, the btree id and the record key.
- Sweep drain. __sweep_server drains the pending list at the top of every wake, before the WT_CONN_CKPT_GATHER skip; __wti_sweep_destroy drains after the thread is joined. Per entry: unlink, history store truncate, then delete the record under the schema lock (metadata walkers re-search keys they iterate). __session_drop does not drain; the internal startup callers still drain synchronously.
- Startup cleanup in __conn_startup_cleanup_and_verify (after recovery, before any table can be created): for each system:drop. record remove the renamed file if present, remove the original only if a rename was scheduled and no file: entry owns the name, truncate the history store for the recorded btree id, delete the record. Read-only connections skip it. A malformed record fails the open naming the key. Backups keep the records so a restored backup truncates the dropped id's history at open.
- RTS exclusion. Rollback-to-stable's history store pass edits pages assuming it is the only writer and its pre-check ignores internal sessions. A new connection spinlock rts_hs_lock is held by __wt_rollback_to_stable around its whole run, outside the checkpoint and schema locks and by the sweep drain and startup cleanup around the truncate only. Lock order: RTS rts_hs_lock -> checkpoint -> schema; sweep rts_hs_lock alone. A sweep pass reaching a truncate waits for a running RTS.
- timing_stress_for_test flags: drop_deferred_skip (the sweep skips deferred drops while set) and drop_hs_truncate_delay (3 s delay inside the deferred truncate while holding the lock). rts_hs_lock is a tracked spinlock with lock_rts_hs_* statistics.
- Observable change: disk and history store space are reclaimed by the sweep shortly after drop returns. Tests that check file absence right after drop poll instead.