-
Type:
Task
-
Resolution: Unresolved
-
Priority:
Critical - P2
-
None
-
Affects Version/s: None
-
Component/s: None
-
DB Integration & Observability
-
None
-
None
-
None
-
None
-
None
-
None
-
None
Overview
We hope to add support for retroactive query stats collection — the ability to record query stats for operations that were initially skipped (e.g., due to rate limiting) but whose total runtime exceeded a configurable slow-query threshold. There is a draft PR provided in the comments which does this. The task for this ticket is to make sure that this doesn't add too much overhead to the server's overall performance.
Background
Query stats registration currently happens upfront at request time. If a request is rate-limited, its stats are never recorded — even if the query turns out to be slow. This creates a systematic bias: slow queries that hit the rate limiter are invisible in query stats, despite being precisely the queries most worth analyzing.
Retroactive collection addresses this by deferring key construction to query completion time, then recording stats only for queries whose runtime exceeds a threshold.
Acceptance Criteria
Create a workload or a set of workloads to demonstrate the performance impact. If the impact is high, escalate and consider abandoning this workstream. If the impact is low/manageable, make sure at least one workload should survive as a regression test, and create a report of the performance characteristics. For example:
- If there are only 'slow' queries on the system (most/all queries are >100ms), what is the impact of this change
- If there are some slow queries intermixed with lower-latency traffic, is there any impact?
- Is there a danger of mis-use if someone lowers the slowms threshold? Can we quantify that danger?