-
Type:
Task
-
Resolution: Unresolved
-
Priority:
Major - P3
-
Affects Version/s: None
-
Component/s: None
-
None
-
Replication
-
None
-
None
-
None
-
None
-
None
-
None
-
None
We have a top-level failover SLO of 1.5s TP99. Failovers are rare in the steady state, which means metrics directly on failover are sparse and poses challenges for setting reliable alarms. This change adds a measurement on every node for time since it last heard from a primary. Primaries report 0, and standbys+read replicas update the gauge periodically before hitting the step-up timeout. By taking outliers on this gauge, we can approximate time spent without a primary.