-
Type:
Bug
-
Resolution: Fixed
-
Priority:
Major - P3
-
Affects Version/s: None
-
Component/s: Test Python
-
Storage Engines, Storage Engines - Transactions
-
7.935
-
None
-
None
-
8
Failure
FAIL: test_layered_delta15.test_layered_delta15.test_internal_page_delta_random(none.snappy.palite.btree.ts.write_both)
...
File "test_layered_delta15.py", line 159, in test_internal_page_delta_random
self.assertStatGreaterSoon(stat.conn.cache_read_internal_delta, 0)
File "helpers/wttest.py", line 1003, in assertStatGreaterSoon
self.assertGreater(val, threshold, msg)
AssertionError: 0 not greater than 0
Root cause
The test's premise does not hold: writing an internal delta at some point during the test does not guarantee the final on-disk internal page is reachable via a delta chain at read-back time.
The test loop (1-10 iterations) does random inserts + checkpoints, then asserts the cumulative stat rec_page_delta_internal > 0 (true if any checkpoint in the loop wrote an internal delta). It then reopens the connection and verifies, asserting cache_read_internal_delta > 0, which only increments when a page is actually reconstructed from base image + deltas (src/btree/bt_read.c, __page_read_build_full_disk_image). This depends entirely on whatever the last checkpoint actually persisted for the internal page(s) that verify() happens to touch.
Several reconciliation-time rejections in src/reconcile/rec_write.c (around the block_meta->delta_count < max_consecutive_delta check) force a full page rewrite instead of a delta, resetting the chain, and any of them can apply on the last checkpoint even though an earlier checkpoint in the same loop produced a delta:
- Multiblock split result (rec_page_delta_rejected_multiblock) - with internal_page_max=512 and random inserts across 10,000 keys, splits are common.
- Delta size over delta_pct of the full page (rec_page_delta_rejected_size_threshold) - write_both uses delta_pct=80 against a tiny 512-byte page.
- Build failure / zero entries.
- The internal page simply isn't dirtied in the final checkpoint, so its address still points at whatever a prior checkpoint wrote, which may itself have been a full image for the same reasons.
So the test is asserting on a probabilistic outcome ("the specific internal page(s) verify() reads still have a live delta chain") it doesn't control, rather than a deterministic property. This matches the project's guidance on not asserting exact/racy statistics without pinning the condition that produces them.
Fix
After the random loop, perform one additional small, targeted update + checkpoint sized to reliably produce an accepted (non-rejected) internal delta for the page(s) checked by verify(), immediately before the reopen/verify/stat assertions, instead of relying on the accumulated randomness of the loop alone.
- related to
-
WT-18815 failed: unit-test-extra-long on amazon2023-arm64-release-nonstandalone [wiredtiger @ 87272c51]
-
- Needs Scheduling
-