-
Type:
Task
-
Resolution: Unresolved
-
Priority:
Major - P3
-
None
-
Affects Version/s: None
-
Component/s: None
-
None
-
Storage Engines - Transactions
-
34.129
-
SE Transactions - 2026-10-09
-
5
Goal
Milestone 1 of the blind-delete re-land: confirm the project is worth doing before designing anything further. Re-land the original blind-delete implementation (the one that was reverted, not a new design) on a throwaway branch and run it through two patch builds:
- a MongoDB correctness patch build, to see which failures the old implementation still produces on current mainline;
- a sys-perf patch build on the DSC variants, to measure the real performance gain.
This is a measurement/gate ticket, not a productionisation ticket. Nothing from the branch is intended to merge. The gate: if the gain does not hold on current mainline, the project stops here.
Code to restore
The blind remove path was added by WT-17254 (a9678864b5) and reverted wholesale by WT-18390 (ae3bd74a57, re-applied to mongodb-dsc-release-5 as a1bf8114b2). The revert touched three files:
- src/cursor/cur_layered.c — the blind_remove branch in _clayered_remove_from_ingest(), which skipped the stable lookup and used _clayered_lookup_ingest_and_truncate() instead
- dist/api_data.py + generated src/include/wiredtiger.h.in — the overwrite documentation describing the layered-follower remove contract
Reverting ae3bd74a57 restores all three. __clayered_lookup_ingest_and_truncate() itself was left in the tree by the revert and still exists, so only the caller comes back.
Note the abort behaviour comes back with it: a caller-contract violation trips WT_ASSERT_ALWAYS with "overwrite=true should guarantee the key exists for remove()". That is the point — the correctness run is there to find out how often that still fires (AF-19923, AF-19942), and whether WT-18159's lazy stable open has removed the lost-delete class (HELP-97958, SERVER-132631, BF-44855).
Enabling blind deletes in the MongoDB patch build
Restoring the WiredTiger branch is not sufficient. The server side has two gates.
1. The server parameter. wiredTigerBlindWriteRatio defaults to 0.999 and is read by chooseBlindWriteOverwrite() in wiredtiger_cursor_helpers.cpp. It only takes effect when shouldUseBlindWriteWhenSafe() is true, which for the disaggregated provider means the node is not accepting writes — i.e. the follower. No change needed here for a production build.
2. The testing-proctor override — this is the blocker for the correctness run. MONGO_INITIALIZER_WITH_PREREQUISITES(SetBlindWriteRatioUnderTestingProctor) in wiredtiger_cursor_helpers.cpp stores 0.0 into the ratio whenever TestingProctor is enabled. resmoke suites run with testing diagnostics on, so every correctness suite currently runs with blind writes disabled regardless of the parameter default. The correctness patch build must remove or bypass that initializer, otherwise it measures nothing.
For the sys-perf run the proctor is off in a production build, so the 0.999 default applies — but confirm it on the deployed cluster rather than assuming it, by checking wiredTigerBlindWriteRatio via getParameter on the follower in the task logs.
Also confirm blind deletes are actually being taken. The reverted implementation carries no statistic (a stat is a later milestone item), so the check is indirect: compare page-service gets per applied op and follower stable-table search counts between the two legs. The earlier A/B showed page-service gets per applied op dropping from 0.35 to 0.24.
Correctness patch build
Standard patch build on the MongoDB patch, with the disagg variants in scope:
- enterprise-amazon-linux2023-arm64-disagg* variants
- the disagg_* resmoke suites under buildscripts/modules/atlas/suites/ — in particular disagg_replica_sets, disagg_repl_jscore_passthrough, disagg_storage, and the change-stream/fuzzer families
- src/mongo/db/modules/atlas/jstests/disagg_storage/index_remove_phantom_key.js, which already forces wiredTigerBlindWriteRatio: 1.0 and asserts the abort does not fire — it is the closest thing to a direct regression test for the reverted path
Record every distinct failure signature, not just a pass/fail count. The output of this run is the list of failure classes the old implementation still produces, which sizes the remaining design work.
sys-perf patch build
Run both legs (blind on, blind off) so the comparison is against the same mainline commit rather than against a historical baseline.
Variants:
- perf-atlas-dsc.arm.aws (M50-dsc, ARM) — the performance workloads
- mongotune-perf-atlas-dsc.availability.arm.aws — the availability workloads
Tasks:
- tpcc_majority_out_of_cache
- tpcc_majority_agg_out_of_cache — the workload the original +45.5% standby-lag figure came from
- ycsb.out_of_cache.100update.2024-05
- ycsb.out_of_cache.95read5update.2024-05
- ycsb.out_of_cache.100read.2024-05 — as a control; blind deletes should not move a read-only workload
- the availability task list on the availability variant
For each task report throughput, average latency and P99, plus standby lag. P99 is the number that did not reach parity in the earlier A/B (216 ms against ASC's 127 ms) and is the one that decides whether blind deletes alone are enough.
Definition of done
- The failure classes the old implementation produces on current mainline, enumerated, with a verdict on whether the lost-delete class is gone.
- Per-workload throughput/latency/P99/standby-lag deltas for blind-on vs blind-off on the same commit.
- A recommendation on the Milestone 1 gate: does the gain justify continuing.
- Branch and patch-build links recorded on this ticket.
- is related to
-
WT-17254 Require a caller-guaranteed live key for overwrite=true remove on a layered follower cursor
-
- Closed
-
-
WT-18390 Disable blind delete path in __clayered_remove_from_ingest
-
- Closed
-
-
WT-18159 Lazily open the stable table on follower to fix lost tombstone on blind-delete miss (dup-key fasserts)
-
- Closed
-