-
Type:
Task
-
Resolution: Unresolved
-
Priority:
Major - P3
-
None
-
Affects Version/s: None
-
Component/s: Reconciliation
-
None
-
Storage Engines, Storage Engines - Transactions
-
133.704
-
SE Transactions - 2026-11-06
-
5
Disaggregated btrees currently build leaf pages at roughly half the size of an identical non-disaggregated table - about 1.7x the page count for the same data, so a configured leaf_page_max=32KB behaves like roughly 16KB. That size is not configured and not chosen. It is where reconciliation happens to land given an interaction between two features described below.
This ticket is to decide what the leaf page size for a disaggregated btree should be, and then to make the code produce that size deliberately.
Current behaviour, measured
Identical workload in all rows - 512MB of 1KB records, 1GB cache, leaf_page_max=32KB, split_pct=90, split target 28KB. Leaf page size measured directly with a debug=(size_stats) tree walk.
| Configuration | Leaf pages | Mean image |
|---|---|---|
| Plain table: | 21,541 | 22.4KB |
| layered: | 34,991 | 13.1KB |
| Plain table: with block_manager=disagg | 37,033 | 13.3KB |
| layered: with the padding term below removed | 17,165 | 26.9KB |
Row three isolates the block manager: no layered table, no ingest, no drain, and the effect is fully present. Row four is the ablation identifying the mechanism - removing one term restores the configured shape, with 16,888 of 17,165 pages landing in the 24KB-28KB histogram bucket.
Independently ruled out: page deltas (13.0KB with leaf and internal deltas disabled), checkpoint cadence (13.3KB at 600s intervals versus 2s), and the disaggregated multiblock-split eviction flag in __sync_check_for_multiblock_rec (13.1KB with it disabled). Configured sizing is identical between the two trees - btree_maxleafpage reads 32KB on the stable constituent with no compression adjustment.
Page size is proportional to leaf_page_max (7.0KB / 13.1KB / 26.7KB at 16KB / 32KB / 64KB) but nearly flat in split_pct (12.0KB -> 13.0KB across 60 -> 90, where the split target moves 47%). So on a disaggregated table split_pct is effectively disconnected, and the size tracks the 50% minimum split boundary rather than the configured target.
How the current size arises
__wti_rec_need_split (src/reconcile/reconcile_inline.h:277) pads each record's length by a tenth of saved-update memory before testing the split boundary:
if (r->page->type == WT_PAGE_ROW_LEAF && page_items > WTI_REC_SPLIT_MIN_ITEMS_USE_MEM)
len += (r->supd_memsize - ((size_t)r->supd_onpage_or_restore * WT_UPDATE_SIZE)) / 10;
Its premise, per its own comment, is that saved updates are content held in memory beyond what the disk image represents, as in update-restore and history-store eviction.
__rec_need_save_upd (src/reconcile/rec_visibility.c:410) saves every selected update on a disaggregated btree not yet flagged WT_UPDATE_DURABLE, for an unrelated reason stated in its own comment: to check whether the reconciliation makes progress, feeding skip-write and delta decisions. A row on its first write is never durable, so every row on the page is saved. A non-disaggregated btree exits earlier at the WT_REC_HS or visible_all checks and saves nothing in the common case.
So one counter is filled and read with two different meanings. Because _wti_rec_split_crossing_bnd (src/reconcile/rec_write.c:1668) gates the minimum boundary on !_wti_rec_need_split(r, 0), and padding applies even at zero length, the minimum boundary stops being a rebalance bookmark and becomes a hard split point.
Two properties separate this from the heuristic's intended use. The padded memory does not survive the write - supd_restore = F_ISSET(r, WT_REC_EVICT) && has_newer_updates (rec_visibility.c:1720) is false for these saves, so writing the image is what releases it, whereas in the update-restore case the chain is restored into the new pages. And the disagg branch is unconditional, while the history-store path is gated on version pile-up, so the term contributes a constant bias rather than a signal.
Related and predating disagg, at rec_visibility.c:725: FIXME-WT-9182: figure out what should be included in the calculation of the size of the saved update chains. The saved-chain size includes the update written to the image, so the estimate is imprecise for every caller. The disagg branch routes every leaf row through it.
Questions to answer
- What leaf page size do we want for a disaggregated btree, and is it the same as for a local btree? Smaller base images may be desirable - cheaper delta application, finer page-log granularity, less rewritten on a full image - but that case has not been made or measured.
- If smaller pages are wanted, what sets the size? Today it is the 50% rebalance boundary, reached incidentally. A deliberate mechanism would be a disagg-specific default or an explicit page-size derivation, and would keep split_pct meaningful.
- If the local size is wanted, how is the memory term corrected? Exempting disaggregated btrees is the direct reading, but skip-write and delta building consume the same saved set, so the choice is between splitting the counter and narrowing the heuristic.
- What does the trade-off actually cost? Page count drives per-page truncate work, WT_REF overhead, internal page fanout, and read amplification, against delta size and page-log write granularity on the other side. No measurement exists for the second half.
- Should WT-9182 be settled first, given both consumers read the same imprecise quantity?
Scope of the measurements
Append-only workload, 1KB values, single table. The save branch keys on WT_UPDATE_DURABLE, so it catches every row on its first reconciliation, making append-heavy tables such as an oplog the worst case. A table that repeatedly rewrites already-durable rows should trip it less; not measured. The exact landing point is also not derived - the padding arithmetic predicts a crossing near 25KB against 13.1KB measured, so the minimum-boundary path explains the direction but not the precise cut.
Related
- is related to
-
WT-18637 Measure fast truncate deletion performance and cache pressure impact
-
- In Code Review
-
-
WT-14695 Merge page deltas into develop
-
- Closed
-
-
WT-16244 Skip writing the leaf page if one-to-one page replacement reconciliation doesn't make any progress in disagg
-
- Closed
-
-
WT-9182 Explore what should be the correct way to calculate upd_memsize in the durable history era
-
- Open
-