Distributed Systems: Maximizing Resilience
摘要
WeDistributed system claim that system redundancyRedundancy (natural or artificial, deliberately introduced) should be applied for the purposes of performance, reliabilityReliability and energy efficiency or “PRE-smartness”. A process of an algorithm of application of available redundancyRedundancy for PRE-smartness is proposed showing that it is further extension generalized algorithm of fault toleranceGeneralized algorithm of fault tolerance when property of fault detection is replaced as property of PRE-smartness. We present a system-level implementation step of PRE-smartness. Using IT-ACS Ltd. forward and backward tracing algorithms applied for distributed computer system we demonstrate that efficiency (reliabilityReliability, especially availability; performance of detection and recovery) grows up to the level when real-time applications of distributed systemDistributed system grow substantially. An estimation of reliabilityReliability gain is modeled and demonstrated suggesting a policy of implementation of PRE-smartness as a permanent ongoing process required for the system successful functioning.