Drafted by Claude, working under my direction, from the investigation in SERVER-134234. I reviewed the findings and the numbers.
Problem
The steady-state OplogWriter's batch limits are compile-time constants (src/mongo/db/repl/oplog_writer_batcher.cpp:19):
const OplogWriterBatcher::BatchLimits defaultBatchLimits{ BSONObjMaxInternalSize, // minBytes 16MB + margin BSONObjMaxInternalSize * 2, // maxBytes 32MB + margin 5 * 1000 // maxCount };
replBatchLimitBytes and replBatchLimitOperations look like the corresponding knobs but are not wired to this path – they feed oplog_applier_batcher.cpp and initial sync only. So in the field there was no runtime lever at all on the writer's batch shape, and the mitigation that was tried (replWriterThreadCount: 32) could not help because it acts on a different thread pool.
Why a knob is worth having: the cost is superlinear in batch bytes
Measured on the SERVER-134234 reproducer, secondary oplog writer, at a constant offered byte rate of ~103 MiB/s – so the only variable is how the bytes are packaged:
| entry size | x86 writer cost | vs linear |
|---|---|---|
| 4 MB | 1739 us per MiB | 1.00x |
| 8 MB | 3769 us per MiB | 2.17x |
| 12 MB | 8953 us per MiB | 5.15x |
Per writer batch, from FTDC on the same runs:
| entry size | MiB per batch | ms per batch |
|---|---|---|
| 4 MB | 3.16 | 3.74 |
| 8 MB | 5.54 | 10.12 |
| 12 MB | 7.90 | 27.54 |
Batch bytes grow 2.5x while batch time grows 7.4x. The affected production host was further out still: bytes per batch 7.5x, time per batch 64x (8.6x worse than linear).
Because cost rises faster than batch size, capping the batch smaller should reduce total cost per byte, not just spread it. That is a hypothesis this ticket makes testable – and, in an incident, actionable.
Root cause of the superlinearity (context, not in scope here)
insertDocsToOplog puts the entire writer batch into a single WriteUnitOfWork. perf on the saturated x86 run attributes the writer thread as: insertDocumentsForOplog 93.8% inclusive, of which _wt_evict/_wt_reconcile is 76.6% and the journal write only 5.9%. So the cost is WiredTiger page reconciliation and block writes driven from inside that one transaction, and it scales with how much dirty data the commit carries.
Proposed work
Add server parameters, set_at: [startup, runtime], and use them in place of defaultBatchLimits:
- oplogWriterBatchLimitMinBytes – default BSONObjMaxInternalSize (current behaviour)
- oplogWriterBatchLimitMaxBytes – default BSONObjMaxInternalSize * 2
- oplogWriterBatchLimitOperations – default 5000
Notes for the implementer:
- getNextBatch() already takes a BatchLimits argument; only the default needs to become dynamic. Read the parameters once per getNextBatch() call, not per polled batch, so a
mid-batch change cannot produce an inconsistent limit pair. - Validators must keep minBytes <= maxBytes, and maxBytes must stay at or above BSONObjMaxInternalSize: a single oplog entry can be up to ~16 MB and the batcher must always be able to make progress on one. A lower bound below that risks a stall, so reject it.
- The existing warning at oplog_writer_batcher.cpp for batches over BSONObjMaxInternalSize should keep working against the configured value.
Acceptance criteria
- The three parameters exist, are settable at runtime via setParameter, and take effect on the next writer batch without a restart.
- Defaults reproduce today's behaviour exactly (no change to steady-state replication when unset).
- Validators reject minBytes > maxBytes and maxBytes < BSONObjMaxInternalSize.
- A jstest sets the parameters at runtime on a secondary under load and asserts the observed metrics.repl.write.batchSize / batches.num ratio moves accordingly.
Test plan
- Unit: extend oplog_writer_batcher_test.cpp with non-default limits.
- Integration: jstest as above.
- Performance: use the large-oplog-entry reproducer built for SERVER-134234 at 12 MB entries, sweeping oplogWriterBatchLimitMaxBytes over e.g. 16/24/32 MiB, and compare writer throughput (MiB per second of writer-thread time) and writer duty cycle. Baseline for that configuration is 92.4% writer duty / 112 MiB per writer-second.
Risks
- Smaller batches mean more WriteUnitOfWork commits and more journal flush triggers; the optimum is likely a trough, not a monotonic win. The sweep above is how we find it.
- This is a mitigation lever, not a fix. It does not remove the single-threaded writer (SERVER-134234) and should not be treated as closing that.
- is related to
-
SERVER-134234 Investigate improvements to OplogWriter for large oplog entries
-
- Investigating
-
- related to
-
SERVER-134905 Investigate why secondaries were unable to keep up with primary workload
-
- Closed
-