ExportXMLWordPrintableJSON

    • Storage Engines - Server Integration
    • Fully Compatible
    • ALL
    • SESIKhuuBeanz 2026-10-06
    • None
    • None
    • None
    • None
    • None
    • None
    • None

      The problem

      Page serving extents (PSEs) are half-open intervals [start_lsn, end_lsn) that partition the log and map LSNs to page servers. The wire types allow a degenerate, empty extent where start_lsn == end_lsn: PageServingExtentId.start_lsn is "inclusive" and end_lsn is "exclusive" and merely optional when the PSE is open (.../cellmetadata/v1/cell_metadata_service.proto:392). Nothing states that end_lsn must exceed start_lsn.

      An empty PSE is emitted by SLS on a normal path (see below). PALI cannot represent one, and two failures follow — which a node hits, and whether it hits both, depends on the order the two extents are delivered in.

      The mis-routing is the broader case. Given the empty extent, it follows whenever the next PSE starts above it, with no same-start successor needed. The abort additionally needs a successor at the same start, so every abort is preceded by a mis-routing window, and a mis-route into a gap is never followed by one.

      1. The invariant fails and the node aborts

      For extents sharing a start:

      • [10, 10) then [10, 20) — aborts the process (invariant failure) when the successor is processed; the empty extent's own ingest succeeds.
      • [10, 10) then [10, ) — aborts the same way.

      The abort requires the empty extent to be delivered before the successor. A node starting fresh applies the whole PSE list in list order — and so does every stream reconnect, since the snapshot is re-sent on each new subscription. At a shared start that list is ordered by ascending endLsn, so a closed empty [10, 10) comes back ahead of a closed successor [10, 20). While a node is running the same sequence arises when the close that creates the empty extent is published first and the successor is published later.

      These are the only two abort orderings. A successor equal to the empty extent (a re-published [10, 10)) has the same endLsn and skips the guard entirely, and no listing can place a closed non-empty extent ahead of the empty one at the same start — the tie-break is ascending endLsn, so only a closed extent with a larger endLsn can follow it.

      PALI matches PSEs on startLsn alone (_findExistingExtent, page_server_reader.cpp:1313; the read lookup at :1158), so an empty [10, 10) occupies the key 10: an incoming extent at the same start is read as a mutation of the empty extent, never as a distinct PSE. addRange then applies its closed-PSE-mutation guard (:1107-1122):

      if (existing.endLsn != extent.endLsn) {
          invariant(existing.endLsn == kDefaultEndLsn, "Closed PSE mutation: ...");
          existing.endLsn = extent.endLsn;   // "Closing open PSE via split"
      }
      

      The matched entry is the empty extent, already closed at endLsn == 10 while kDefaultEndLsn is uint64_t::max (page_server_reader.h:62), so the invariant fails and mongod aborts via LOGV2_FATAL — on a background PALI thread, where an abort is not attributable to a request. It is the same death path the ModifyClosedPseNotAllowed test already asserts on (page_server_reader_test.cpp:2697), but that test covers a genuine closed-PSE mutation, not a zero-length extent.

      The abort is triggered by the successor's arrival, not by the empty extent itself. [10, 10) is ingested cleanly: _findExistingExtent finds no same-start entry, so addRange inserts it with no comparison and no log line. Only when a second extent at start 10 is processed does the guard fire — and at that point it is the empty extent, not the newcomer, that is bound as existing.

      2. Reads are mis-routed to the wrong page servers

      The corruption requires the open extent to be delivered before the empty one. That is the order any full list arrives in — endLsn orders None < Some, so [10, ) precedes [10, 10) — and it is also what happens while a node is running, since the node already holds [10, ) when SLS publishes the close. The empty extent then reaches PALI as the closed form of an extent PALI already holds open. Matching is on startLsn alone (_findExistingExtent, :1313), so the incoming [10, 10) is read as a mutation of the open entry, not as a distinct PSE. The guard at :1106 sees the existing entry open — existing.endLsn == kDefaultEndLsn — so it passes, and PALI applies the close itself:

      existing.endLsn = 10;                            // :1119 "Closing open PSE via split"
      existing.servers = std::move(extent.servers);    // :1125
      

      The original [10, ) is now the empty [10, 10), and the genuine open upper half is never created — it would have collided on start_lsn. The map is left with no open PSE for the tail, and the node does not crash.

      Reads for that range then resolve to that entry. getPageServersForLsn selects the greatest startLsn <= lsn and never consults endLsn — the extent comparator uses startLsn only (page_server_reader.h:290-292):

      auto it = std::upper_bound(extents.begin(), extents.end(), search);
      if (it == extents.begin() || (--it)->servers.empty()) { /* "No page server serving LSN" */ }
      

      Every LSN at or above 10 that no later extent covers resolves to that entry. Its servers are no longer the segment's: existing.servers = std::move(extent.servers) (:1125) runs on this branch, outside the guard, so the open entry's correct list is replaced by the empty extent's — the outgoing set. The gap is not necessarily adjacent: the next PSE after the empty one may start well above it — [50, ) in the example — so every LSN in [10, 50) resolves to the empty extent. Any re-subscription puts a node in this state — a restart, or the stream reconnecting after an error, since PALI keeps its map across a stream failure rather than clearing it — and it comes back with the empty extent registered in place of the open one, serving reads for that whole range from the wrong servers rather than aborting. The error branch cannot catch it: it fires only when no extent has startLsn <= lsn, which the empty extent makes false for every lsn >= 10, and servers is never empty (_processIncomingExtent drops an extent with no parsable server, :1277-1280). The path is silent — no LOGV2_ERROR, no metric.

      The steady-state sequence is both failure modes, in order. The empty extent corrupts the map when it is applied (this section), and the successor at the same start then aborts the node (section 1). The mis-routed reads happen before the crash, not instead of it — so a single incident can show wrong pages and a later abort that look like two unrelated events.

      Note the collision is per cell: _servers is keyed by cell name and each streamed response is assumed to describe one cell (:1262), so only two extents sharing a start within one cell conflict.

      There is no test for either failure.

      Impact. Silent and self-reproducing: no log line, no metric, and a restart or a stream reconnect rebuilds the same corrupted map rather than clearing it. Reads in the affected range are answered from servers that do not own it, and the only loud symptom — the later abort — does not look related.

      How an empty PSE arises

      Not a corrupt or malformed response — the wire types permit start_lsn == end_lsn, and SLS is the producer. Two paths are known:

      • Same-start PSE add (SLS-3633, upcoming). With an open [10, ) and a new PSE requested at start 10, the same-start branch in add_page_serving_extent — today an inconsistent_config error (page_metadata.rs:222) — is to create the empty predecessor, closing the open extent at its own start.
      • Empty phylog log segment (SLS-10307, "Close PSEs for fully materialized segments", commit 910dd930e). A segment sealed while phylog-empty is reported fully_materialized at its inherited PHYLOG floor, i.e. its own start LSN — the equal-LSN update SLS-10307 began allowing (cellmetadata/src/log_catalog.rs: "an op-log-only seal can be reported at the inherited floor"). Sealing an empty segment is routine (a segment can seal on a size/time boundary with nothing written).

      Either way the requirement is on PALI: an extent with startLsn == endLsn can arrive, in any order relative to a sibling at the same start, and PALI must handle the shape rather than abort or corrupt its map — the two failures above are what it does instead today.

      References

      • Ingest: src/mongo/db/modules/atlas/src/disagg_storage/pali/page_server_reader.cpp:1247
      • Closed-PSE-mutation guard: same file, :1107-1122; server-list copy on the same branch: :1125
      • Lookup: same file, :1146-1178; comparator: page_server_reader.h:290-292; kDefaultEndLsn: page_server_reader.h:62
      • Matcher: same file, _findExistingExtent, :1313; read lookup :1158
      • CMS forwarding: src/mongo/db/modules/atlas/src/disagg_storage/pali/sls_config_manager.cpp:457-459 (stream callback :557, no reordering in the ingest path; a response's extents are applied under one _mutex hold, :1090)
      • Proto: src/mongo/db/modules/atlas/src/disagg_storage/sls-proto/dist/storage/etc/protos/cellmetadata/v1/cell_metadata_service.proto:392-402
      • Existing test: page_server_reader_test.cpp:2697 (closed-PSE mutation)
      • SLS side: SLS-3633 (PM should call CMS AddPageServingExtent without waiting for next phylog; its same-start add produces the empty PSE) — the producing path; SLS-10307 (close PSEs for fully materialized segments; commit 910dd930e)
      • SLS stream: updates are one extent per message, each a single-element list (server.rs publish_page_serving_extent :1927-1957, try_send); AddPageServingExtent publishes only the requested extent (:891), StartMaterialization publishes the covering/created PSE then the closed predecessor (:1039, :1045), ClosePageServingExtents publishes each closed extent (:987); snapshot sent unsorted by pse_stream_manager.rs:53-69
      • SLS extent ordering: storage/components/cellmetadata/src/page_metadata.rs — list order at a shared start_lsn is by ascending end_lsn with open (None) first (get_page_serving_extent_list, :1301-1331); verified by test_get_page_serving_extent_list_orders_open_before_closed_empty_at_same_start and test_get_page_serving_extent_list_orders_empty_before_closed_successor_at_same_start
      • SLS close boundaries: all four close sites use a strictly greater boundary (select_closeable_extents :513-517, record_materialized_lsns_and_maybe_close :660-671, start_materialization :1523 idempotent for a same-start request, add_page_serving_extent :296-316 closes a prev with a lower start); narrowest closed extent producible today is [s, s+1)
      • SLS empty-extent handling: storage/components/cellmetadata/src/page_metadata.rs (find_covering_page_serving_extent, same-start add refused at :222) and storage/components/cellmetadata/src/log_catalog.rs (inherited-floor fully_materialized update)

            Assignee:
            Nic Hollingum
            Reporter:
            Daotang Yang
            Votes:
            0 Vote for this issue
            Watchers:
            3 Start watching this issue

              Created:
              Updated:
              Resolved: