-
Type:
Bug
-
Resolution: Unresolved
-
Priority:
Major - P3
-
Affects Version/s: None
-
Component/s: None
-
None
-
Storage Engines - Foundations
-
ALL
-
200
-
None
-
None
-
None
-
None
-
None
-
None
-
None
drop_replay_after_trim.js fails intermittently (~21% on enterprise-amazon-linux2023-arm64-disagg-no-schema-epochs) at Step 1, before the drop under test happens:
drop_replay_after_trim.js:376 assert.soonNoExcept(..., "the OIS never resolved the source objects for the created collection", TIMEOUT_MS /* 20s */)
The test creates a collection, fsyncs, reads the new pages off the page servers, then polls objectindexd.GetPageAtLSN to map those pages to their S3 source objects. The page servers have the page immediately, but OIS visibility lags behind it by two async stages:
- pagematd buffers phylog entries and only flushes/uploads them when log_upload_threshold_time expires - a recurring timer on pagematd's own clock. pagematd_config in buildscripts/modules/atlas/sls-multicell-docker-compose.yml does not set it, so it uses the sls default of 30s. The byte trigger (log_upload_threshold_bytes) is 1 GiB and is never reached by one small collection.
- objectindexd then folds the uploaded .sidx into an output-index file on sync_period: 5s.
Total latency is therefore ~5s at best and ~35s at worst, depending on where the write lands in pagematd's 30s cycle. The test allows 20s, which sits inside that range - hence the flakiness.
Observed in BF-46365:
04:38:30.965 first GetPageAtLSN -> NotFound "Output index sequence log id 1 domain PHYLOG not found"
04:38:49.400 pagematd-cell1-0 30s tick: writes span 1-0-1-0-7685616802486812700-2.data
(span_start_lsn=0, so it does cover the requested LSN ...682)
04:38:49.470 .sidx uploaded, OIS notified
04:38:49.5+ error changes to "does not have a file covering ... LSN 7685616802486812682"
04:38:51.077 assert.soon expires at 20s <-- test fails here
04:38:53.007 OIS builds output index file ordinal=1 -> query would have succeeded (+22.1s)
Needed 22.1s, had 20s. No product bug - the data was materialized correctly, just not indexed yet.
The test has been flaky since it was introduced in 0a62870c0b6 (SERVER-134515).
- is related to
-
SERVER-134515 Add mongo testing coverage for replaying table drop() on an already discarded table's data
-
- Closed
-