错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Generalized Algorithm of Fault Tolerance (GAFT)

  • Igor Schagaev,
  • Jürg Gutknecht

摘要

Fault toleranceFault tolerance so far wasGeneralized algorithm of fault tolerance considered as a property of a system. In fact and instead we introduce A Generalized Algorithm of Fault ToleranceFault tolerance (GAFT) that considers property of fault toleranceFault tolerance as a system process. GAFT implementation analysis—if we want to make it rigorous—should be using classification of redundancyRedundancy types. Various redundancyRedundancy types have different “power” of use at various steps of GAFT. Properties of GAFT implementation impact on overall performance of the system, coverage of faults and ability of reconfigurationReconfiguration. Clear that separation of malfunctions from permanent fault simply must be implemented and reliabilityReliability gain is analyzed. A ratio of malfunctions to permanent faults is achieving 105–7 and simple exclusion from working configuration a malfunctioned element is no longer feasible. Further we must consider GAFT extension in terms of generalization and application for support of system safety of complex systems. Our algorithms of searching correct state, “guilty” element and analysis of potential damages become powerful extension of GAFT for challenging applications like avionic systems, aircraft. In Chap.  3 , we showed that fault toleranceFault tolerance should be treated as a process. In this chapter, we elaborate further this process into a clearly defined algorithm and develop a framework to the design of fault tolerant systems, the generalized algorithm of fault toleranceGeneralized algorithm of fault tolerance—GAFT. We also introduce a theoretical model to quantify the impact of the additional redundancyRedundancy to the reliabilityReliability of the whole system and derive an answer to the question of how much added redundancyRedundancy leads to the system with highest reliabilityReliability. A question that GAFT cannot answer is how the real source of a detected fault can be identified, as the fault manifestation might have occurred in another hardware element and spread in the system due to nonexistent fault containment. We will show an algorithm that based on the dependencies of the elements of a system can identify the possible fault sources and also predict which elements an identified fault might have affected. We now start in a first step by further elaborating the process of fault toleranceFault tolerance.