-
Type:
Task
-
Resolution: Unresolved
-
Priority:
Minor - P4
-
None
-
Affects Version/s: None
-
Component/s: None
-
None
-
Storage Engines - Persistence
-
85.962
-
StorEng - Defined Pipeline
-
None
This is a follow up ticket for WT-18754, discussed with etienne.petrel@mongodb.com .
The context for this ticket is based on the following pre-conditions:
- After fixing delivery of
WT-17864, we're still facing an issue that the cached disagg database size is inconsistent with the sum of all the checkpoints' sizes. - There's not a reliable way to identify which clusters are suffering from this inconsistency as we don't have a way other than an active wiredTiger repair call to know it.
- For example, if a database size is 2GB, then it's silently drop 500MB for some reason and then be back to normal, we won't be able to distinguish it from from a normal cluster.
- Step up / step down won't reset the cached value, so currently we're relying on all the accumulations are always right without bugs, but in another words, some bugs may be silently hidden inside.
Proposed behaviour change:
- (Auto fix) Each time the last step before step up, we do a database size calibration call.
- If the cached value aligned well with the computed value, we do nothing.
- If the value drifting over a threshold, we do the following steps:
- Fix the cached size.
- Left a verbose error log about what happened
- Increase a metrics stat from 0/1 or simply set it as the fixed gap size, so the mis-alignment can be visible on the grafana dashboard and then we can capture it.
- is related to
-
WT-17864 Investigate continuous constant database size decreasing issue.
-
- Closed
-