<p>As the demand for computational power continues to rise, large-scale hierarchical networks play an increasingly pivotal role in high-performance computing (HPC) systems. Fault diagnosis is essential for ensuring the stability and reliability of these complex networks. However, traditional fault diagnosis methods face significant challenges in scalability and real-time performance as system sizes and complexities grow. This paper introduces a parallel adaptive system-level fault diagnosis algorithm (PAD-HCN) that leverages the Hamiltonian structure to efficiently and accurately diagnose faults in large-scale hierarchical cubic networks (HCNs). Specially, we first prove that HCNs are Hamiltonian, a property that forms the theoretical foundation for designing an efficient fault diagnosis mechanism. Building upon this property, we propose a parallel and adaptive fault diagnosis algorithm that significantly enhances diagnostic performance. Simulation results demonstrate that the proposed algorithm achieves exceptional diagnostic accuracy, maintaining nearly <InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11227_2025_7897_Article_IEq1.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="39" /> </InlineMediaObject> <EquationSource Format="TEX">\(100\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>100</mn> <mo>%</mo> </mrow> </math></EquationSource> </InlineEquation> accuracy when the number of faulty vertices remains within the fault-tolerance threshold. Furthermore, the parallelization of the algorithm leads to a substantial reduction in diagnosis time, with the parallel approach achieving up to a <InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11227_2025_7897_Article_IEq2.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="51" /> </InlineMediaObject> <EquationSource Format="TEX">\(99.97\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>99.97</mn> <mo>%</mo> </mrow> </math></EquationSource> </InlineEquation> reduction in time compared to traditional non-parallel methods. These results underscore the scalability of PAD-HCN, demonstrating its ability to efficiently handle the increasing size and complexity of modern HPC systems while preserving high fault diagnosis accuracy and real-time responsiveness.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Adaptive system-level fault diagnosis of hierarchical cubic networks

  • Mengjie Lv,
  • Sixiao Di,
  • Weibei Fan

摘要

As the demand for computational power continues to rise, large-scale hierarchical networks play an increasingly pivotal role in high-performance computing (HPC) systems. Fault diagnosis is essential for ensuring the stability and reliability of these complex networks. However, traditional fault diagnosis methods face significant challenges in scalability and real-time performance as system sizes and complexities grow. This paper introduces a parallel adaptive system-level fault diagnosis algorithm (PAD-HCN) that leverages the Hamiltonian structure to efficiently and accurately diagnose faults in large-scale hierarchical cubic networks (HCNs). Specially, we first prove that HCNs are Hamiltonian, a property that forms the theoretical foundation for designing an efficient fault diagnosis mechanism. Building upon this property, we propose a parallel and adaptive fault diagnosis algorithm that significantly enhances diagnostic performance. Simulation results demonstrate that the proposed algorithm achieves exceptional diagnostic accuracy, maintaining nearly \(100\%\) 100 % accuracy when the number of faulty vertices remains within the fault-tolerance threshold. Furthermore, the parallelization of the algorithm leads to a substantial reduction in diagnosis time, with the parallel approach achieving up to a \(99.97\%\) 99.97 % reduction in time compared to traditional non-parallel methods. These results underscore the scalability of PAD-HCN, demonstrating its ability to efficiently handle the increasing size and complexity of modern HPC systems while preserving high fault diagnosis accuracy and real-time responsiveness.