-
Type:
Bug
-
Resolution: Fixed
-
Priority:
Major - P3
-
Affects Version/s: None
-
Component/s: Layered Tables
-
Storage Engines - Foundations
-
942.808
-
SE Foundations - 2026-09-29, SE Foundations - 2026-10-13
-
3
Problem:
A checkpoint stamped with schema epoch E claims to cover every schema operation published at or below E. Both leader's and follower's entries leave the shared metadata queue depend on that:
- the leader drains them at checkpoint time,
- a follower prunes them when it picks up a checkpoint.
WT_SESSION::publish checks the requested epoch against the stable epoch and the step-down boundary, but never against last_ckpt_disaggregated_schema_epoch:
- Leader: there's no gap. Its stable epoch can't be set below the last checkpoint's epoch, and publish must be above stable.
- Follower: picking up a checkpoint advances the adopted epoch but not the stable epoch. That leaves a range above stable and at or below the adopted epoch where a publish passes every check.
An entry published in that range ends one of two ways:
- The next pickup prunes it as already covered, although it never reached shared metadata, so the operation is lost.
- It survives until step-up and trips WT_ASSERT(entry->schema_epoch > last_ckpt_epoch).
Solution:
Ensure that the following steps are atomic (under schema lock):
- merge of shared and local metadata
- update of the latest adopted checkpoint epoch
- pruning the schema ops queue based on adopted epoch
This way concurrent publish call sees either all of the above or none of them.