-
Type:
Bug
-
Resolution: Fixed
-
Priority:
Major - P3
-
Affects Version/s: None
-
Component/s: Test Csuite
-
Storage Engines - Foundations
-
35.246
-
None
-
None
Symptom
TSan reports data races in test/csuite/schema_disagg_abort itself (not in WiredTiger) in the two-node lone-follower runs of smoke.sh (-r lf ...). Seen in PR #14814's patch 6ab9dfe7c9593e0007815e7e on amazon2023-arm64-tsan (-r lf -k l8), and locally on a macOS arm64 TSan build in -r lf and -r lf -k l8.
WARNING: ThreadSanitizer: data race
Write of size 1 by thread T24:
#0 thread_reader_run reader.c:79
Previous read of size 1 by main thread:
#0 node_trigger_wait node.c:171
#1 node_run node.c:398
WARNING: ThreadSanitizer: data race
Write of size 4 by thread T22:
#0 pipe_relay_event event_pipe.c:41
#1 apply_event worker.c:293
Previous read of size 4 by thread T21:
#0 pipe_relay_event event_pipe.c:33
#1 apply_event worker.c:293
Cause
TEST_CONFIG.peer_alive and TEST_CONFIG.pipe_write_fd are shared by the node's threads without synchronization.
- peer_alive is cleared by a worker when a relay write fails (event_pipe.c:42) and by the reader on pipe EOF (reader.c:79). The generator (generator.c:51, :257, :354), the workers (worker.c:147) and the main thread (node.c:171) read it with plain loads.
- pipe_write_fd: a leader's workers relay concurrently (worker.c:272, :293, :316). The first to see EPIPE closes the descriptor and stores -1 (event_pipe.c:40-41) while the others load it to write (event_pipe.c:33, :36).
Inference, not observed: beyond the TSan report, two workers can both close the same descriptor number, or write an event into a descriptor number the process has meanwhile reused for another file.
Reproduce
Build with TSan and run bash ./smoke.sh from build/test/csuite/schema_disagg_abort. The connection API race (WT-18743) fires first in the single-node runs unless it is fixed or suppressed; these races show up in the -r lf runs.
Impact
Blocks re-enabling the schema_disagg_abort TSan smoke (FIXME-WT-18743 entries in test/evergreen.yml): with WT-18743 fixed, these races still fail it.
Proposed fix
Load and store peer_alive atomically, and keep the out-pipe open for the process's life instead of closing it from a worker, as the node already does with its in-pipe and self-pipe.