ExportXMLWordPrintableJSON

    • Type: Bug
    • Resolution: Unresolved
    • Priority: Major - P3
    • None
    • Affects Version/s: None
    • Component/s: Logging
    • None
    • Storage Engines - Persistence
    • 11.994
    • StorEng - Defined Pipeline
    • None

      Problem
      When a slot has WTI_SLOT_SYNC_DIRTY set (and not WTI_SLOT_SYNC), __wti_log_release starts the asynchronous flush on log->log_fh, the current log file handle:

      if (F_ISSET_ATOMIC_16(slot, WTI_SLOT_SYNC_DIRTY) && !F_ISSET_ATOMIC_16(slot, WTI_SLOT_SYNC)) {
          WT_FH *log_fh = __wt_atomic_load_ptr_acquire(&log->log_fh);
          if ((ret = __wt_fsync(session, log_fh, false)) != 0) {
      

      The slot's data was written to slot->slot_fh, which _wti_log_slot_activate captures when the slot is activated. _log_newfile can switch log->log_fh to the next log file while an earlier slot is still being released: the switch happens under the slot lock when a new slot is set up, and release runs without that lock. When that happens, the non-blocking flush runs on the new log file rather than the one holding the slot's dirty data, so the write-back that log.os_cache_dirty_pct asked for never starts on the file that needs it.

      __log_slot_dirty_max_check only schedules the dirty sync when the release LSN and the last sync LSN are in the same file. It checks this when the slot is closed, so a file switch between closing and releasing the slot is not covered.

      Impact
      The dirty flush is best-effort. Durable syncs go through __log_fsync_file, which opens its own handle for the target LSN's file, so this should not lose data. What goes wrong is that the flush hits the wrong file, so dirty pages build up in the OS cache on the old file until it is closed and fsynced.

      Suggested fix
      Flush slot->slot_fh instead of log->log_fh in the WTI_SLOT_SYNC_DIRTY branch of _wti_log_release. Check that the slot's file handle cannot be closed by the log file close server before the release finishes. The close server waits for write_lsn to pass log_close_lsn, and _wti_log_release advances write_lsn just before this flush, so this ordering needs checking.

            Assignee:
            [DO NOT USE] Backlog - Storage Engines Team
            Reporter:
            Etienne Petrel
            Votes:
            0 Vote for this issue
            Watchers:
            1 Start watching this issue

              Created:
              Updated: