ExportXMLWordPrintableJSON

    • Type: Bug
    • Resolution: Unresolved
    • Priority: Major - P3
    • None
    • Affects Version/s: None
    • Component/s: None
    • None
    • Cluster Scalability
    • ALL
    • ClusterScalability 14Sep-28Sep, ClusterScalability 28Sep-12Oct
    • None
    • None
    • None
    • None
    • None
    • None
    • None

      Symptom

      analyzeShardKey fails outright on a valid candidate shard key when the collection carries two indexes with the same key pattern – legal when they differ only by collation, for example a case-insensitive unique index alongside a simple one. The command hints its cardinality aggregation by index key pattern rather than by index name, so the hint matches both indexes and the query planner refuses it:

      (IndexNotFound) error enumerating plans for query: ns=sample_supplies.sales1 planner returned
      error :: caused by :: Hint matched multiple indexes, must hint by index name.
      Matched: kp: { customer.email: 1, _id: 1 } unique name: 'customer.email_1__id_1_unique_ci'
          and kp: { customer.email: 1, _id: 1 } name: 'customer.email_1__id_1'
      

      Reproduction

      // Two indexes with the same key pattern, differing only by collation.
      db.sales1.createIndex({"customer.email": 1, _id: 1});
      db.sales1.createIndex(
          {"customer.email": 1, _id: 1},
          {name: "customer.email_1__id_1_unique_ci",
           unique: true,
           collation: {locale: "en", strength: 2}});
      
      db.adminCommand({
          analyzeShardKey: "sample_supplies.sales1",
          key: {"customer.email": 1, _id: 1},
          keyCharacteristics: true,
          readWriteDistribution: true
      });
      

      The same failure occurs for key: {"customer.email": "hashed", _id: 1}, since the supporting index is selected by field name either way.

      Analysis

      Index selection is not the problem – it already resolves to exactly one index. findCompatiblePrefixedIndex skips any index with a non-simple collation:

      if (!indexDesc->collation().isEmpty()) {
          continue;
      }
      

      so customer.email_1_id_1_unique_ci is skipped and the simple customer.email_1_id_1 is chosen. That is also the correct choice on the merits: a strength-2 case-insensitive index collapses distinct values, so cardinality read off it would be understated for a case-sensitive shard key.

      The selected index is then described to the planner by key pattern – aggRequest.setHint(hintIndexKey) – and that pattern matches both indexes. So the command discards the one piece of information that identifies its own choice unambiguously: IndexSpec carries {keyPattern, isUnique} and never the index name, although indexDesc->indexName() is available in the selection loop.

      Reaching the planner at all confirms selection succeeded: the hint is built only after an index is chosen, and a collection with no compatible index takes the documented IllegalOperation "no supporting index" path instead.

      Code references (master, 986761aa688b7e2a6c493eb450256ec5a1bbfd55, src/mongo/db/s/analyze_shard_key_cmd_util.cpp):

      • Collation skip during selection: L454
      • Prefix match by field name: L458
      • struct IndexSpec – holds the key pattern and uniqueness, not the name: L401
      • The hint: aggRequest.setHint(hintIndexKey) at L292

      Suggested fix

      Keep index selection exactly as it is and identify the selected index to the planner by name instead of by key pattern: carry indexDesc->indexName() out on IndexSpec and hint with BSON("$hint" << indexName), which is the representation parseHint already produces for a string hint. No IDL change is needed – the aggregate command's hint field is typed indexHint, "either a string index name or an object index specification". The key pattern still has to be passed for building the pipeline (hashed field detection, the $group), so only the hint changes.

      This also covers duplicate key patterns arising from options other than collation – partialFilterExpression or sparse – which selection skips but a pattern hint still matches.

      For the clustered-collection branch, ClusteredIndexSpec does not always carry a name; falling back to the pattern hint there is safe, since a clustered collection has exactly one clustered index and the hint cannot be ambiguous.

      Two alternatives seem worse:

      • Skipping indexes whose key pattern is shared would discard a usable simple-collation index and report "no supporting index" where full metrics are computable.
      • Catching IndexNotFound and retrying without a hint would lose the index-only scan the aggregation relies on (see SERVER-119912).

      Relationship to SERVER-135375

      Same theme – an index is selected and then cannot be used – but a distinct defect, and neither fix resolves the other. SERVER-135375 throws BadValue from validateIndexKey at L193, before the aggregation is built; this one throws IndexNotFound from the planner after the hint is set at L292. Hinting by name does not help SERVER-135375, because validation runs first; skipping descending indexes during selection does not help here, because the ambiguity is between two ascending indexes.

      Impact

      This affects MongoDB Atlas's Sharding Advisor, which calls analyzeShardKey to evaluate candidate shard keys. Duplicate key patterns differing by collation are an ordinary layout – case-insensitive lookups are the usual reason – so an otherwise valid candidate key becomes unanalyzable because of an unrelated index on the collection. Atlas cannot work around it: the candidate key is valid, and the only client-side remedy would be asking the customer to drop a working index.

            Assignee:
            Jordan Glassley
            Reporter:
            Alex Dambrouski
            Votes:
            0 Vote for this issue
            Watchers:
            3 Start watching this issue

              Created:
              Updated: