Raise the compile-time minimum NCCL version to 2.27 and remove the
preprocessor and runtime version gates that guarded features introduced in
NCCL 2.27 or earlier. The introduction version of every gated feature was
cross-checked against the NCCL release notes before its gate was removed:
AVG (2.10), PreMulSum (2.11), ncclGetLastError/ncclRemoteError (2.13),
non-blocking comms + ncclConfig_t (2.14), ncclCommSplit (2.18),
ncclCommRegister/ncclMemAlloc (2.19), group abort (2.22),
ncclCommInitRankScalable (2.23), FP8 e5m2/e4m3 (2.24), QoS/trafficClass
(2.26), and ncclCommShrink / ncclCommWindowRegister / symmetric memory /
collnetEnable / CTAPolicy / nvlsCTAs (2.27).
Suggested review order:
- NCCLUtils.hpp: the static_assert is bumped from 2.7 to 2.27 and the
now-unconditional feature macros are deleted. This is the source of
truth that makes every gate below dead code. - NCCLUtils.cpp, ProcessGroupNCCL.{hpp,cpp}, init.cpp: the dead
#ifdef/#ifndef/#elsebranches for those macros are removed and the
live branch inlined; the runtimegetNcclVersionNumber() >= 2.22
(group abort) andversion < 2.19(mem allocator) checks are dropped. - cuda/nccl.{h,cpp}: the local 2.13/2.14 feature macros and the ancient
NCCL 1.x / 2.0 / 2.7 / 2.12.10 guards are removed; only NCCL-presence
gates (USE_NCCL, defined(NCCL_MAJOR)) and >= 2.28 gates remain. - symm_mem/nccl_dev_cap.hpp and the C++ test: NCCL_HAS_SYMMEM_SUPPORT is
made unconditional within USE_NCCL.
Gates for versions newer than 2.27 are kept as-is (>= 2.28 ncclAlltoAll,
NCCL_HAS_SYMMEM_DEVICE_SUPPORT, NCCL_HAS_COMM_OFFLOAD 2.29.7,
NCCL_HAS_MAX_P2P_PEERS 2.30, etc.), as is the defensive runtime
version >= 2.27 check inside ProcessGroupNCCL::shrink().
Two macros were not deleted outright because they encode NCCL availability,
not just a version, and are evaluated in non-NCCL build configurations:
- HAS_NCCL_BF16_DATATYPE (nccl.h) keeps its CUDA-bf16/ROCm structure; only
theNCCL_MINOR >= 10subexpression is dropped (nowNCCL_MAJOR >= 2),
preserving the "NCCL absent => disabled" behavior its consumers rely on. - NCCL_HAS_SYMMEM_SUPPORT is kept as a macro but defined unconditionally
inside#if USE_NCCL, since the symmetric-memory source files use it to
gate on NCCL presence as well as version.
Test Plan:
Build against NCCL 2.29.7 (satisfies the new static_assert):
uv pip install -e . --no-build-isolation -v
Lint the changed files:
lintrunner -a torch/csrc/distributed/c10d/NCCLUtils.hpp \
torch/csrc/distributed/c10d/NCCLUtils.cpp \
torch/csrc/distributed/c10d/ProcessGroupNCCL.cpp \
torch/csrc/distributed/c10d/ProcessGroupNCCL.hpp \
torch/csrc/distributed/c10d/init.cpp \
torch/csrc/cuda/nccl.cpp torch/csrc/cuda/nccl.h \
torch/csrc/distributed/c10d/symm_mem/nccl_dev_cap.hpp \
test/cpp/c10d/ProcessGroupNCCLTest.cpp
This PR was authored with the assistance of Claude, an AI coding assistant.