-
Type:
Task
-
Resolution: Unresolved
-
Priority:
Major - P3
-
None
-
Affects Version/s: None
-
Component/s: None
-
None
-
Cluster Scalability
-
🟦 Shard Catalog
-
None
-
None
-
None
-
None
-
None
-
None
SERVER-134771 brought up a generic problem in the migration mechanics.
In that issue they do moveChunk back and forth and since auth shards we have a new check failing in this scenario (more info on it here). The intention was to return a ConflictingOperationInProgress (from the destination manager side), however the user gets an OperationFailed.
The root issue has always been there, none of any error codes escaped from the destination manager (callee) as the source manager (caller) folds every errors coming from the destination into an OperationFailed. To be clear: it's not even a "returning a status" rather writing a doc and the source polls that doc for result. None the less the destination never returned the error code, just the error message. In some cases the destination manager code just sets the error message without even using an error code internally.
The case here is that there is a new error path which is way easier to hit now with auth shards that just returns OperationFailed instead of the original intention of ConflictingOperationInProgress. The reasoning behind ConflictingOperationInProgress is that eventually - in 5 min - it will be resolved automagically and you can retry.
Instead of folding the error into OperationFailed on the source side we always should use appropriate error code if the destination goes into a Fail state.
This ticket is about
- Find a way to propagate error codes from destination to the source
- Find the appropriate error codes for each fail state on the destination
- is related to
-
SERVER-134771 v9 reports ConflictingOperationInProgress as wrong error code
-
- Closed
-