Recovery: Searching and Monitoring of Correct Software States
摘要
The last of the three GAFT processes is called recovery and recovery monitoring. After the detection of an error and possible reconfigurationReconfiguration, the last step is recovering the software, which means that the effect of the error on the software must be eliminated. In line with the previous chapters and (Sogomonian and Schagaev in Hardware and software fault tolerance of computer systems. Avtomatika i Telemekhanika, pp. 3–39, 1988 [1]; Schagaev in Avtomat Telemekh 4, 1989 [2]; Schagaev in Autom Remote Control 51(3), 1990 [3]; Schagaev in Reliability of malfunction tolerance. International Multi-conference on Computer Science and Information Technology (IMCSIT 2008), pp. 733–737, 2008 [4]; Schagaev et al. in ERA: Evolving reconfigurable architecture. 11th ACIS International Conference, pp. 215–220, 2010 [5]; Castano and Schagaev in Resilient computer system design. Springer, Berlin, 2015 [6]), recovery consists of restoring the last recovery pointRecovery points and continue processing. But is this really sufficient? What if latent faults exist in the system and manifest themselves in the system but trigger some detection schemes an arbitrary time later? Assuming this reasonable and unpleasant sequence of events it becomes clear that just restoring data and program from the last stored recovery pointRecovery points is not enough. We have to admit that we do not have any guarantee that fault is now eliminated: even when hardware is restored or even reconfigured—we have erroneous states of software recorded in recovery pointsRecovery points. Thus we have to consider the recovery processRecovery process itself, analyze which classic algorithms are applicable and fit the purpose of efficient recovery. We introduce and analyze three recovery algorithms that are able to ensure successful recovery by iteratively go through all stored recovery pointsRecovery points.