-
Type:
Task
-
Resolution: Unresolved
-
Priority:
Major - P3
-
None
-
Affects Version/s: None
-
Component/s: None
-
None
-
Storage Execution
-
None
-
None
-
None
-
None
-
None
-
None
-
None
This dashboard should give us insight into the health of the replicated fast count system on customer clusters and the ability to proactively address issues.
Some insights that could be valuable include:
- How lagged the oplog tailer is behind the latest oplog entry (replicatedFastCount.oplogLagSecs)
*Whether the tailer/flusher threads are running - Time since last flush
Some of this work may involve creating/whitelisting new metrics for tracking. Whitelisted metrics can be found here; I don't believe any metrics related to fast count are whitelisted yet.