-
Type:
Task
-
Resolution: Unresolved
-
Priority:
Critical - P2
-
None
-
Affects Version/s: None
-
Component/s: Observability
-
Storage Engines - Persistence
-
201.059
-
SE Persistence backlog
-
5
Fleet dashboard needs:
- Signals into improving the policy. Clusters not using the block cache that should be?
- Errors
- Links to field guide etc, method of providing feedback, DRI for the policy
- Dry-run, running, engaged, disabled, ineligible, unknown
- Number of clusters tuned in selected time frame
- Breakdown of different clusters running the policy by Mongotune version, MongoDB server version, policy configuration (TODO what is this??)
- For each alert, how close to firing it is
- Error count frequencies for poll and tune errors, along with a top-10 list of clusters with highest error frequency, apparently each cluster name needs to be a link to the per-cluster page??
- Last error codes by cluster count
- Something about conflicts of control? To see if Mongotune is fighting an Atlas override?
Per-cluster dashboard needs:
- Brief description of the policy
- Links to: Fleet-wide view, Field guide, Troubleshooting wiki/escalation procedure, Method of providing feedback, DRI for the policy
- Policy state per host (Dry-run indicator, Clear view of when the policy has taken action, Thresholds for decision-making actions)
- Cluster-health signals and how they should be interpreted (Signals that show how the policy is helping, Signals that show how the policy is hurting)
- Conflicts with Atlas custom overrides
- Errors that have occurred: Total error count, Last error code
- Cluster-specific signals that feed policy alerts – threshold to show when an alert will fire
- Descriptions for anything non-obvious
- "Make sure the visualisations provide signal for: dry-run, real engagement, bad configuration, disablement, telemetry gaps, rollback, and recovery scenarios."