计算机科学
同步(交流)
排队
排队论
网络拓扑
计算机网络
星团(航天器)
缩放比例
控制(管理)
网络拥塞
还原(数学)
服务器
软件部署
基站
标度律
通信系统
交通拥挤
计算机集群
拓扑(电路)
比例(比率)
实时计算
集体行为
线性比例尺
分布式计算
电信网络
钥匙(锁)
吞吐量
作者
Taoran Qi,Shuo Li,Xingqi Zou,Liangce Deng,Guodong Wei
标识
DOI:10.1109/infocom55648.2025.11044740
摘要
The evolution of AI has driven the trend towards utilizing larger models with increasing parameter sizes, necessitating deployment in datacenters comprising tens of thousands of GPUs. However, as cluster sizes expand, the overall computational capacity of the system fails to scale linearly with the number of compute nodes. This non-linear scaling law is primarily due to decreased GPU utilization caused by network congestion, which significantly impacts model training speed. In this paper, we analyze the network topology and traffic characteristics of AI clusters and propose PC4, a precision collective communication congestion control algorithm tailored for AI clusters. PC4 optimizes congestion control based on collective communication operations (e.g., all-reduce, all-to-all) inherent in collective communication libraries (e.g., MPI, NCCL) used in distributed training. By leveraging information from these collective communication operations, we compute a base rate that enables rapid system convergence. Furthermore, PC4 employs a datacenter time synchronization mechanism, utilizing one-way delay as congestion signal. This approach allows for more precise control of link queue lengths and queuing time, thereby efficiently managing congestion. In our all-to-all evaluation, PC4 achieves a 72% reduction in tail FCT compared to DCQCN.
科研通智能强力驱动
Strongly Powered by AbleSci AI