ExportXMLWordPrintableJSON

    • Type: Task
    • Resolution: Unresolved
    • Priority: Major - P3
    • None
    • Affects Version/s: None
    • Component/s: Not Applicable
    • Storage Engines - Foundations
    • 33.828
    • None
    • None

      Recently, as part of HELP-99980, we’ve discovered that there are many different timeouts and other interesting rules that can affect how role changes happen. For example, crossing the step-down timeout of 10 seconds can cause a fatal assertion failure, while there is another 60-second step-down fence that prevents us from stepping back up. We also have an election timeout of 2 seconds, which can easily be exceeded by a step-up but, at the moment, doesn’t seem to lead to any serious consequences.

      Given that, I think we should improve our observability around these events. As part of WT-18752, we already introduced metrics to measure different stages of the step-up logic and better understand where the time is being spent. In addition to that, we should introduce metrics to measure the health of role changes across the fleet, with a focus on both step-up and step-down.

      What I currently have in mind is:

      • Step-down: measure the concluding checkpoint time, how long it takes for the current follower to close the oplog lag and become able to pick up the checkpoint, how long it takes to pick up the checkpoint, how much time we spend from setting the step-down timestamp to reconfiguring, and how long the reconfiguration takes.
      • Step-up: measure the overall duration, as well as how much time is spent in the different stages of the process.
      • MongoDB-level timeouts: measure how frequently we cross different timeouts, such as the step-down timeout or election timeout, to understand how often this happens across the fleet. The tricky part is that these timeouts might be changed over a time so we should make sure that our metrics would be able to detect the updated value.

      This should give us a better picture of how healthy role changes are in practice and help us identify cases where we’re getting close to, or exceeding, the limits imposed by the MongoDB layer.

            Assignee:
            [DO NOT USE] Backlog - Storage Engines Team
            Reporter:
            Ivan Kochin
            Votes:
            0 Vote for this issue
            Watchers:
            1 Start watching this issue

              Created:
              Updated: