-
Type:
Bug
-
Resolution: Unresolved
-
Priority:
Major - P3
-
None
-
Affects Version/s: None
-
Component/s: Layered Tables
-
Storage Engines - Foundations
-
146.188
-
None
-
None
Problem
conn->next_file_id is uint32_t (src/include/connection.h:1022), incremented once per file created - local tables, indexes, and both constituents of every layered table (src/schema/schema_create.c:193, 1274). It's a single shared counter, cluster-lifetime not per-process: IDs are never reused on drop, and a follower inherits a leader's high-water mark via checkpoint pick-up (__raise_next_file_id, src/conn/conn_layered_checkpoint_pick_up.c:1365-1372) rather than starting fresh, so restarts don't reduce exposure.
The counter doesn't need to hit UINT32_MAX to break - WT_BTREE_ID_NAMESPACED(base_id, ns) = base_id << 3 | ns (src/include/btree.h:106-107) also returns uint32_t, so the effective ceiling is 2^29 (~537M). Past that, the top 3 bits of the raw id are silently discarded by the shift before the namespace tag is OR'd in. _validate_file_id (src/schema/schema_create.c:130-156) only asserts non-zero and non-collision with a handful of reserved special IDs - no check against 2^29, no wraparound detection anywhere in _wt_generate_file_id.
Impact
A wrapped/truncated id can collide with a low numeric id already assigned to a live, unrelated file. File ids are load-bearing identity (routing log records to the correct btree, cache/eviction bookkeeping). Depending on where the collision lands, outcome is either an immediate __wt_panic (layered table manager hard-asserts on seeing the same id twice while live, src/conn/conn_layered_table_manager.c:90-92) or, worse, a live btree silently receiving state/log records meant for a different file.
Reachability
Not just theoretical - from churn rates observed on production clusters (sustained ~12 file-id increments/sec on one cluster), 2^29 is reachable in under two years of continuous table/collection churn (e.g. tenant provisioning, time-series bucket rotation), well within normal cluster lifetime.
Fix direction
Widen the counter (e.g. uint64_t or size_t) and/or add an explicit bounds check in __validate_file_id that fails loudly (return EINVAL/WT_PANIC, not a silent wrap) before crossing the shift boundary. Needs a decision on wire/on-disk format impact if the namespaced id width changes, since it's persisted in metadata and log records.