错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Cascading-Failure-Aware Distributed Computing System with Performance Sharing: Reliability and Robustness Analysis

  • Ankit Gupta,
  • Dharmendra Prasad Mahato

摘要

In distributed computing systems, enhancing reliability has been a significant focus of research through various approaches such as task allocation optimization, software and hardware redundancy, and performance sharing mechanisms. However, the issue of cascading failures poses a challenge to the reliability of such systems, as it involves a feedback loop leading to a complete system failure. This paper presents a comprehensive analysis of a distributed computing system that incorporates performance sharing, aiming to address the problem of cascading failures and improve system reliability. A model for evaluating the reliability and robustness of the distributed system with cascading failures is proposed. By utilizing graph networks and simulation techniques, the dependability of the system is assessed, considering various attack and defense strategies. The research findings indicate that defending against cascading failures yields higher network robustness compared to inducing failures. Moreover, a strong correlation between robustness and reliability is observed, implying that enhancing the robustness of the network leads to increased reliability of the distributed computing system. The proposed methodology and findings contribute to the development of higher-quality distributed systems, particularly in the face of cascading failures. The research outcomes have significant implications for various applications reliant on distributed systems, enhancing their performance and resilience.