互连
计算机科学
多处理
可靠性(半导体)
公制(单位)
表征(材料科学)
图形
容错
分布式计算
拓扑(电路)
并行计算
钥匙(锁)
理论(学习稳定性)
断层(地质)
网络拓扑
局域网
算法
比例(比率)
理论计算机科学
性能指标
图论
连接部件
本地系统
有向图
组分(热力学)
作者
W. C. Zheng,Shuming Zhou,Eddie Cheng,Lulu Yang
标识
DOI:10.1093/comjnl/bxag009
摘要
Abstract With the growing scale and complexity of high-performance computing systems, ensuring reliability through robust fault diagnosis becomes increasingly critical. System-level diagnosis plays a key role in identifying faulty processors and maintaining system stability of multiprocessor systems. However, traditional diagnosability, as a global reliability metric for multiprocessor systems, overlooks local diagnostic capability, topological criticality, and fault distribution. In order to better capture the local characteristics of a system around a given node, this work proposes a novel fault diagnosis strategy, called cyclic local diagnosability, where the cyclic fault pattern requires that at least two components contain cycles. We propose some characterizations of cyclic local diagnosability of interconnection networks under PMC and MM* models. As applications, we determine the cyclic local diagnosabilities of data center network DCell ($D_{k,n}$), $(n,k)$-star graph ($S_{n,k}$) and $(n,k)$-bubble-sort graph ($B_{n,k}$) under PMC and MM* models. Finally, we show the superiority of the cyclic local diagnosability through comparison with other conditional diagnosabilities.
科研通智能强力驱动
Strongly Powered by AbleSci AI