-
Type:
Task
-
Resolution: Unresolved
-
Priority:
Major - P3
-
None
-
Affects Version/s: None
-
Component/s: None
-
None
-
Storage Engines - Transactions
-
34.463
-
None
-
None
Summary
A disaggregated table created while a stable schema epoch is set starts out "awaiting publication". Until a checkpoint publishes it, none of its pages may be written to shared storage. Reconciliation and the in-memory split code have special handling for this state, treating these tables like in-memory tables. That handling can't run: eviction is switched off for a table for as long as it awaits publication, so the table is never reconciled in that state. The dead branches make the code harder to reason about. They should be replaced with an assertion.
Why the table is never reconciled while awaiting publication
- Eviction is disabled when the table opens in this state. Eviction could otherwise leave pages clean without a durable address, and the publishing checkpoint would then skip them.
- The eviction server doesn't walk the table while eviction is disabled.
- Forced eviction is skipped: reading a page that has grown too large doesn't force-evict it when eviction on the table is disabled.
- Checkpoint skips the table until it is published.
- Close discards its pages rather than writing them.
- Publication is ordered safely: it clears the awaiting-publication state first and only then re-enables eviction, so eviction never sees a table still awaiting publication.
Uses of __wt_btree_stays_in_memory and WT_BTREE_AWAITS_PUBLISH
The helper returns true for an in-memory table (WT_BTREE_IN_MEMORY) or a table awaiting publication (WT_BTREE_AWAITS_PUBLISH). The call sites, as of the current develop branch, fall into three groups.
Reconciliation, including the eviction flag choice: unreachable for a table awaiting publication, so only the in-memory case matters
- __rec_need_save_upd: returns early so no update chain is saved, neither for the history store nor for the disaggregated progress check. The progress check relies on durable flags that only a page write sets.
- __wti_rec_upd_select: chooses in-memory update selection and never writes prepared updates.
- __rec_upd_select_inmem: a direct WT_BTREE_AWAITS_PUBLISH check under precise checkpoint keeps updates newer than the pinned stable timestamp out of the in-memory image.
- __wti_rec_row_leaf: exempts the table from the "leaked prepared update" assertion.
- __wti_rec_time_window_clear_obsolete: decides whether globally visible time points can be cleared.
- __rec_is_checkpoint: a root page write without a split is never a checkpoint.
- __rec_write: diagnostic assertion that the page is never written to disk.
- __evict_reconcile: picks in-memory update restore with an always-saved image.
- __split_multi_inmem: skips instantiating updates from the page image. A direct WT_BTREE_AWAITS_PUBLISH check in the same function keeps the rebuilt page dirty.
Eviction outside reconciliation: also unreachable while eviction is disabled, but harmless to keep
- __wti_evict_walk: skips the tree unless it is evicting dirty pages or pages with updates.
- __evict_try_queue_page: never treats a page as a clean-eviction candidate.
- __evict_review: refuses to evict a clean page.
- __wt_evict: doesn't discard a clean page outright.
- __evict_page_clean_update: doesn't move the page to the victim cache.
- __evict_ckpt_snapshot_required: precise-checkpoint visibility bounds don't apply.
Paths reachable while the table awaits publication: these must keep the broader check
- __wt_conn_dhandle_close: the table skips checkpoint, so its pages are discarded rather than flushed.
- __wt_txn_read: skips the history store lookup.
- _wti_rts_visibility_page_needs_abort and _rts_btree_abort_ondisk_kv: rollback to stable treats the table as having no on-disk state.
- __wti_page_inmem_updates: asserts the table isn't in-memory. Its callers already exclude such tables.
Proposal
- Assert the invariant: assert on entry to reconciliation that the table isn't awaiting publication.
- Narrow the in-memory checks: in the reconciliation group above, check only WT_BTREE_IN_MEMORY. Remove the two direct awaiting-publication branches: the precise-checkpoint selection rule and the keep-dirty case in the in-memory split. The write-path assertion in __rec_write can stay as it is.
- Keep the broader check elsewhere: leave the reachable group, and optionally the eviction group, unchanged. The check is correct and cheap there.
To confirm before landing
- No caller other than publication re-enables eviction on the table while it still awaits publication. Exclusive eviction access is reference-counted, so balanced enable and disable pairs are fine.
- No disaggregated path calls reconciliation directly, bypassing eviction. Running the layered and disaggregated Python tests and the disaggregated format runs with the new assertion will show any such path.