Add metrics that can be used to detect when initial replicated fast count flush is stalling

XMLWordPrintableJSON

    • Type: Task
    • Resolution: Duplicate
    • Priority: Major - P3
    • None
    • Affects Version/s: None
    • Component/s: None
    • None
    • Storage Execution
    • None
    • None
    • None
    • None
    • None
    • None
    • None

      The oplog_lag_secs metric is only set to a non-zero value after the first flush succeeds. We should have a way to detect when the initial flush is being held up.

      Some ideas:

      • Log time since last completed flush

      We should also consider logging the persisted valid-as-of timestamp, with this value being unset until the first flush. This would also let us tell at a glance what the last persisted valid-as-of timestamp is set to, without needing to read from containers.

            Assignee:
            Unassigned
            Reporter:
            Damian Wasilewicz
            Votes:
            0 Vote for this issue
            Watchers:
            1 Start watching this issue

              Created:
              Updated:
              Resolved: