-
Type:
Task
-
Resolution: Unresolved
-
Priority:
Major - P3
-
None
-
Affects Version/s: None
-
Component/s: None
-
None
-
Replication
-
Repl 2026-10-12
-
None
-
None
-
None
-
None
-
None
-
None
-
None
Action item
Investigate and remediate the recurring server-side causes of majority replication lag seen in repeat-offender clusters.
Context
Provides the engineering disposition for customers that remain affected after the Stream 1 prevention rollout — that is, clusters where the released pacing and prevention mechanisms are active and majority replication lag still recurs.
Ownership
- Stream: Stream 4 — Create a repeat-offender process with CSM and TSE support
- Team(s): Replication
- Owner: Unassigned — to be agreed at rollout kickoff
- Rollout gate: Phase 0 — Mobilize (runs throughout)
Background
This is a tracked action item for the Atlas roll-out of majority replication lag prevention, detection, and remediation — Milestone 3 of the Majority Replication Lag SMART goal (GOAL-251), tracked by SPM-4794.
Milestone 1 established fleet-wide observability and RCA foundations. Milestone 2 delivered server enhancements and lag-aware pacing mechanisms. Milestone 3 turns those results into an operational system that detects incidents, explains their causes, gives COEs an approved response path, and reduces recurrence among the worst offenders.
References
- Tracking epic: SPM-4794
- SMART goal: GOAL-251
- Rollout plan: GOAL-251: Atlas Rollout (section 2.5)
- Repl Lag SMART Goal Plan: Milestone plan
- Related: SERVER-129638
- Related: SPM-4631, SPM-4519 (prevention work whose residual gaps this ticket absorbs)
- Related: Engineering Proposal: Capping Majority Replication Lag for 8.0+ Clusters
Filed from a placeholder ticket reference in the rollout plan document. Title, scope, and acceptance criteria are expected to be refined by the owning team.