-
Type:
Bug
-
Resolution: Unresolved
-
Priority:
Major - P3
-
None
-
Affects Version/s: None
-
Component/s: None
-
DB Integration & Observability
-
ALL
-
None
-
None
-
None
-
None
-
None
-
None
-
None
Background
The search executor shutdown backstop (SERVER-129849) bounds mongod shutdown when a search cursor is pinned by an in-flight getMore and cannot be disposed. In gRPC egress mode (useGrpcForSearch=true, used by Search in Community / self-managed deployments), this exposes a transport-layer problem: mongod cannot shut down cleanly while a pinned search cursor keeps the mongot task executor alive.
Repro
jstests/with_mongot/search_mocked/search_executor_shutdown_with_inflight_cursor.js under the search_community suite (currently excluded there under a TODO referencing this ticket). It opens a $search cursor, pins it with a getMore blocked at the getMoreHangAfterPinCursor failpoint, then shuts mongod down with searchTaskExecutorShutdownTimeoutMS=8000 and asserts EXIT_CLEAN plus elapsedMs >= 5000.
Root cause (two independent problems):
- execPtr->join() hangs. shutdownTaskExecutor() calls execPtr->shutdown() then execPtr->join(). With the mongot executor on a gRPC network interface, join() blocks forever: ThreadPoolTaskExecutor::join() -> NetworkInterface shutdown -> GRPCAsyncClientFactory::shutdown() blocks on _cv.wait(_shutdownComplete), which requires _numActiveHandles == 0. The pinned cursor's leased gRPC stream keeps _numActiveHandles > 0, and the handle is only released when the cursor is disposed; this happens only after killAllOperations(), i.e. after search executor shutdown. This results in mongod hanging.
- invariant(_clients.empty()) aborts at transport-layer shutdown even if the join is bounded. GRPCTransportLayerImpl::shutdown() asserts _clients.empty() at tlm->shutdown(). GRPCAsyncClientFactory::startup() registers one GRPCClient per NetworkInterface in _clients; it is erased only when the factory's _client ref is released via the destructor. The pinned cursor holds a shared_ptr to the executor (ClientCursor -> pipeline -> TaskExecutorCursor -> PCTE -> executor), so the executor and its gRPC client outlive the timeout, and the invariant aborts.
Fix direction
Either
- (a) make GRPCTransportLayerImpl::shutdown() tolerate live clients (gracefully shut them down instead of asserting empty), while ensuring no use-after-free if a client outlives the transport layer; or
- (b) release the factory's client ref during factory shutdown so _clients empties even while the executor lives.
A bounded-join change in search_task_executors.cpp (running shutdown+join on a helper thread capped by searchTaskExecutorShutdownTimeoutMS, detaching on timeout) could also be part of the complete fix, but might be unfeasible/undesirable.
Note: This only affects gRPC egress (useGrpcForSearch=true). The MongoRPC path does not register a transport-layer client and the inflight test passes there.
- depends on
-
SERVER-129849 Fix mongod shutdown hang waiting for search task executor
-
- Closed
-
- is blocked by
-
SERVER-129849 Fix mongod shutdown hang waiting for search task executor
-
- Closed
-