Resilience of Networked Systems: A Taxonomy of Challenges, Faults, Disciplines, and Attributes
摘要
Failures in networked systems are inevitable. They may occur due to various challenges, including forces of nature (such as hurricanes or earthquakes), human errors (e.g., cable cuts), or malicious attacks, to mention a few. Despite the visible diversity of their characteristics, they share a common feature: There is no way to eliminate them entirely. This chapter discusses the taxonomy of challenges and faults, errors, and failures, and describes the disciplines of resilience referring to network design approaches to provide service continuity (such as survivability, fault tolerance, traffic tolerance, and disruption tolerance mechanisms), as well as measurable characteristics, including the attributes of network dependability (such as reliability and availability), security, or performability. The latter part of this chapter explains the techniques for evaluating and improving the total availability and reliability of serial, parallel, and mixed architectures of systems.