• Type: Bug
    • Resolution: Unresolved
    • Priority: Major - P3
    • None
    • Affects Version/s: None
    • Component/s: None
    • DB Integration & Observability
    • ALL
    • None
    • None
    • None
    • None
    • None
    • None
    • None

      Background

      The search executor shutdown backstop (SERVER-129849) bounds mongod shutdown when a search cursor is pinned by an in-flight getMore and cannot be disposed. In gRPC egress mode (useGrpcForSearch=true, used by Search in Community / self-managed deployments), this exposes a transport-layer problem: mongod cannot shut down cleanly while a pinned search cursor keeps the mongot task executor alive.

      Repro

      jstests/with_mongot/search_mocked/search_executor_shutdown_with_inflight_cursor.js under the search_community suite (currently excluded there under a TODO referencing this ticket). It opens a $search cursor, pins it with a getMore blocked at the getMoreHangAfterPinCursor failpoint, then shuts mongod down with searchTaskExecutorShutdownTimeoutMS=8000 and asserts EXIT_CLEAN plus elapsedMs >= 5000.

      Root cause (two independent problems):

      1. execPtr->join() hangs. shutdownTaskExecutor() calls execPtr->shutdown() then execPtr->join(). With the mongot executor on a gRPC network interface, join() blocks forever: ThreadPoolTaskExecutor::join() -> NetworkInterface shutdown -> GRPCAsyncClientFactory::shutdown() blocks on _cv.wait(_shutdownComplete), which requires _numActiveHandles == 0. The pinned cursor's leased gRPC stream keeps _numActiveHandles > 0, and the handle is only released when the cursor is disposed; this happens only after killAllOperations(), i.e. after search executor shutdown. This results in mongod hanging.
      2. invariant(_clients.empty()) aborts at transport-layer shutdown even if the join is bounded. GRPCTransportLayerImpl::shutdown() asserts _clients.empty() at tlm->shutdown(). GRPCAsyncClientFactory::startup() registers one GRPCClient per NetworkInterface in _clients; it is erased only when the factory's _client ref is released via the destructor. The pinned cursor holds a shared_ptr to the executor (ClientCursor -> pipeline -> TaskExecutorCursor -> PCTE -> executor), so the executor and its gRPC client outlive the timeout, and the invariant aborts.

      Fix direction

      Either

      • (a) make GRPCTransportLayerImpl::shutdown() tolerate live clients (gracefully shut them down instead of asserting empty), while ensuring no use-after-free if a client outlives the transport layer; or
      • (b) release the factory's client ref during factory shutdown so _clients empties even while the executor lives.

      A bounded-join change in search_task_executors.cpp (running shutdown+join on a helper thread capped by searchTaskExecutorShutdownTimeoutMS, detaching on timeout) could also be part of the complete fix, but might be unfeasible/undesirable.

      Note: This only affects gRPC egress (useGrpcForSearch=true). The MongoRPC path does not register a transport-layer client and the inflight test passes there.

            Assignee:
            Unassigned
            Reporter:
            Mariano Shaar
            Votes:
            0 Vote for this issue
            Watchers:
            2 Start watching this issue

              Created:
              Updated: