Create Grafana Dashboard to track fleet-wide Replicated Fast Count Metrics

XMLWordPrintableJSON

    • Type: Task
    • Resolution: Unresolved
    • Priority: Major - P3
    • None
    • Affects Version/s: None
    • Component/s: None
    • None
    • Storage Execution
    • None
    • None
    • None
    • None
    • None
    • None
    • None

      This dashboard should give us insight into the health of the replicated fast count system on customer clusters and the ability to proactively address issues.

      Some insights that could be valuable include:

      • How lagged the oplog tailer is behind the latest oplog entry (replicatedFastCount.oplogLagSecs)
        *Whether the tailer/flusher threads are running
      • Time since last flush

      Some of this work may involve creating/whitelisting new metrics for tracking. Whitelisted metrics can be found here; I don't believe any metrics related to fast count are whitelisted yet.

            Assignee:
            Unassigned
            Reporter:
            Damian Wasilewicz
            Votes:
            0 Vote for this issue
            Watchers:
            2 Start watching this issue

              Created:
              Updated: