-
Type:
Task
-
Resolution: Unresolved
-
Priority:
Major - P3
-
None
-
Affects Version/s: None
-
Component/s: Replication
-
Replication
-
None
-
None
-
None
-
None
-
None
-
None
-
None
The problem
A step-up trigger (10985438) fires when a standby has had no contact — heartbeat, oplog entry, or checkpoint-end — for the election timeout. Only the heartbeat is periodic, so a trigger means a heartbeat gap of at least one election timeout. But the logs cannot say why, so the pattern "standby steps up while the primary is healthy" (HELP-99821, HELP-99678, HELP-99572) recurs with no attribution.
Questions the logs cannot currently answer
- On a 10985438 trigger, was the primary heartbeating the triggering member at that time — was the member in the primary's heartbeat target set?
- What was the gap: when was the last heartbeat sent to that member (primary side), and when was the last one received (standby side)?
- Did the gap start on the send side (the primary stopped sending) or the receive side (the standby or the network did not deliver it)?
- Is the gap persistent (the member is skipped for as long as the primary stays up — i.e. the skip-self failure, SERVER-134162) or a one-off?
Why none of these are answerable today
- the trigger is a canned message with no gap and no history;
- heartbeats are not logged on either side — not which members the primary targets, not receipt on the standby;
- the only related state, _lastPrimaryContactTimeMillis, is a single value surfaced as the OTEL gauge disaggStorage.stepUp.lastPrimaryContactMs — not heartbeat-specific, no per-target meaning.
Links
- HELP-99821, HELP-99678, HELP-99572 — the "standby steps up while the primary is healthy" pattern.
- SERVER-134162 — the skip-self / wrong-member heartbeat failure.
- SERVER-112825 — evaluate election-timeout behavior for disagg.