Fleet-wide and per-cluster dashboards for block cache resize policy

XMLWordPrintableJSON

    • Type: Task
    • Resolution: Unresolved
    • Priority: Critical - P2
    • None
    • Affects Version/s: None
    • Component/s: Observability
    • Storage Engines - Persistence
    • 201.059
    • SE Persistence backlog
    • 5

      Fleet dashboard needs:

      • Signals into improving the policy. Clusters not using the block cache that should be?
      • Errors
      • Links to field guide etc, method of providing feedback, DRI for the policy
      • Dry-run, running, engaged, disabled, ineligible, unknown
      • Number of clusters tuned in selected time frame
      • Breakdown of different clusters running the policy by Mongotune version, MongoDB server version, policy configuration (TODO what is this??)
      • For each alert, how close to firing it is
      • Error count frequencies for poll and tune errors, along with a top-10 list of clusters with highest error frequency, apparently each cluster name needs to be a link to the per-cluster page??
      • Last error codes by cluster count
      • Something about conflicts of control? To see if Mongotune is fighting an Atlas override?

      Per-cluster dashboard needs:

      • Brief description of the policy
      • Links to: Fleet-wide view, Field guide, Troubleshooting wiki/escalation procedure, Method of providing feedback, DRI for the policy
      • Policy state per host (Dry-run indicator, Clear view of when the policy has taken action, Thresholds for decision-making actions)
      • Cluster-health signals and how they should be interpreted (Signals that show how the policy is helping, Signals that show how the policy is hurting)
      • Conflicts with Atlas custom overrides
      • Errors that have occurred: Total error count, Last error code
      • Cluster-specific signals that feed policy alerts – threshold to show when an alert will fire
      • Descriptions for anything non-obvious
      • "Make sure the visualisations provide signal for: dry-run, real engagement, bad configuration, disablement, telemetry gaps, rollback, and recovery scenarios."

            Assignee:
            Will Korteland
            Reporter:
            Will Korteland
            Votes:
            0 Vote for this issue
            Watchers:
            1 Start watching this issue

              Created:
              Updated: