-
Type:
Task
-
Resolution: Duplicate
-
Priority:
Major - P3
-
None
-
Affects Version/s: None
-
Component/s: Replication
-
None
-
Replication
-
Repl 2026-09-28
-
None
-
None
-
None
-
None
-
None
-
None
-
None
We have had a number of escalations where resharding's cloner was apparently producing load that the replica set primary could handle, but its secondaries could not, leading to increasing replication lag over time:
- HELP-81857
- HELP-84458
- HELP-99223
In these cases, there was no obvious contention on the secondaries to explain why they were unable to keep up (at least, not obvious to someone who was not an expert in replication).
We should do a root cause analysis to determine why the secondaries were unable to keep up, and what about resharding's workload, if anything, exacerbated the situation.
- is duplicated by
-
SERVER-134234 Investigate improvements to OplogWriter for large oplog entries
-
- Investigating
-
- is related to
-
SERVER-134234 Investigate improvements to OplogWriter for large oplog entries
-
- Investigating
-
-
SERVER-135612 Make the steady-state OplogWriter's batch limits runtime-tunable
-
- Backlog
-
- related to
-
SERVER-135037 Ensure sufficient metrics exist to diagnose replication lag
-
- Open
-