-
Type:
Task
-
Resolution: Done
-
Priority:
Major - P3
-
None
-
Affects Version/s: None
-
Component/s: DHandles
-
None
-
Storage Engines - Foundations
-
1,730.586
-
SE Foundations - 2026-06-23, SE Foundations - 2026-07-07, SE Foundations - 2026-07-21, SE Foundations - 2026-08-04
-
8
Background
SLS-6071 reports that Disaggregated Storage (DSC) exhibits a regression of over 60% in OperationThroughput compared to classic MongoDB (ASC) on the data_handle_locust workload. Latency is also significantly higher on DSC. This was initially surfaced in PERF-7206.
Observed Numbers
| Test Name | Measurement | ASC (Classic) | DSC | DS vs ASC (%) |
|---|---|---|---|---|
| CreateOne | Latency50thPercentile | 12.782 | 29.2512 | +128.85% |
| CreateOne | Latency95thPercentile | 36.155 | 334.2754 | +824.56% |
| CreateOne | OperationThroughput | 18.19 | 6.72 | -63.03% |
| FindOne | Latency50thPercentile | 0.432 | 0.5132 | +18.80% |
| FindOne | Latency95thPercentile | 0.867 | 1.309 | +50.98% |
| FindOne | OperationThroughput | 13894.99 | 5187.71 | -62.66% |
| UpdateOne | Latency50thPercentile | 12.025 | 29.2082 | +142.90% |
| UpdateOne | Latency95thPercentile | 34.331 | 331.642 | +865.91% |
| UpdateOne | OperationThroughput | 19.47 | 6.70 | -65.50% |
Goal
Investigate the root cause of the DSC performance gap on the high-active-dhandle workload and identify potential improvements. Per the comment in SLS-6071, the regression may be related to DSC's use of layered tables.
Suggested Investigation Areas
- Profile DSC under the data_handle_locust workload to identify hot paths vs ASC
- Examine data handle open/close/sweep costs in DSC (layered table overhead vs standard btree)
- Check whether dhandle cache eviction or sweep behaviour differs significantly between ASC and DSC
- Review lock contention on the dhandle list in the DSC layered-table code path
- Compare checkpoint and reconciliation costs between ASC and DSC under this workload
- Assess whether the 95th-percentile latency spikes (8-9x on write operations) point to periodic stalls (e.g. flush/ingest)
References
- SLS-6071: https://jira.mongodb.org/browse/SLS-6071
- PERF-7206: https://jira.mongodb.org/browse/PERF-7206
- is related to
-
WT-17198 Merge Refactor __checkpoint_db_internal feature branch back to develop
-
- Closed
-
-
WT-17961 Discovery: Table Creation & Collection Scaling for Public Preview (DSC)
-
- Closed
-
- related to
-
WT-17791 Optimise direct ingest creation to mitigate core startup bottlenecks
-
- Closed
-
-
WT-18061 Dhandle scaling Perf: On standby, layered cursors can pick up new stable checkpoints with staggered delays
-
- Open
-
-
WT-18062 Dhandle scaling Perf: Keep a list of modified dhandles for checkpoint prepare
-
- Open
-
-
WT-18063 Dhandle scaling Perf: eliminate point metadata search in checkpoint prepare
-
- Open
-
-
WT-18064 Dhandle scaling Perf: For checkpoint prepare, hoist any checkpoint config parsing out of the dhandle gather loop
-
- Open
-
-
WT-18160 Investigate: Perform a second perf analysis of dhandle locust workloads
-
- Open
-
-
WT-18028 Extend WT dhandle benchmark for disagg and add plateaus
-
- In Code Review
-
-
WT-17924 For disagg, don't checkpoint WiredTiger.wt
-
- Closed
-
-
WT-18060 Dhandle scaling Perf: On standby, do selective pickup of individual btree checkpoints
-
- Closed
-
-
WT-18082 Dhandle scaling Perf: create initial set of ingest tables lazily
-
- Closed
-
-
WT-18159 Lazily open the stable table on follower
-
- Closed
-