DDP training on RTX 4090 (ADA, cu118)

export NCCL_P2P_DISABLE=1 sorta works for models like GitHub - The-AI-Summer/pytorch-ddp: code for the ddp tutorial. (DDP works, slowly; DP gives NaN loss).
Yet for the life of me I cannot get it working on my models (~= CLIP transformer). If NCCL is enabled, it hangs with 100% volatile GPU utilization, but the processes can be killed with ^C or kill -9. If NCCL is disabled, it hard freezes the system.

This was working perfectly well a few days ago on two 2080Ti with otherwise identical hardware. Model trains fine on either one of the single 4090s. IOMMU is disabled in BIOS. memtest good, gpu_burn reports no errors either; hardware seems fine.

These GPUs need nvidia-driver >= 520, (using 525.78.01) which comes with cuda 12.0. (Related issue: torch_compile also doesn’t work b/c they need sm_89 etc). I might just train on one GPU until the new hardware bugs get ironed out …