-
Type:
Bug
-
Resolution: Unresolved
-
Priority:
Major - P3
-
None
-
Affects Version/s: None
-
Component/s: None
-
None
-
Query Optimization
-
ALL
-
None
-
None
-
None
-
None
-
None
-
None
-
None
Problem
During lowering, the predicates of all edges are attached here. And when multiple edges use the same index, the actual runtime seek key is picked by the last predicate in the stage builder here. If two predicates target the same right-side field (the same index case), then the last predicate in syntactic order wins the seek.
Example
Collections A, B, C with join predicates A.a = B.b, A.a = C.c, and B.x = C.y (the B–C edge is required so {B,C} can be joined first and A probed last), with index {a: 1} on A.
The INLJ candidate probing A is costed with the selectivity of A.a = B.b (first edge), but at runtime the seek value comes from C.c (last predicate). Results are correct, but the cost does not reflect the IO performed: if A.a = B.b is much more selective than A.a = C.c, the INLJ is costed optimistically, can win against HJ on paper, and then fetches a large fraction of A per probe.
Solution
This is even more important in the context of SERVER-131547, given that we want to start picking the best edge for INLJ in a cost based manner. Then we of course expect the actual execution engine to also use that edge. We need to add more info to the INLJ node and modify the stage builders to respect that.
- is related to
-
SERVER-131547 INLJ costing depends on syntactic join-edge order
-
- Open
-
-
SERVER-116534 Consider multiple edges when picking the index for an INLJ
-
- Open
-