Recovery of Real-Time Clusters with the Division of Computing Resources into the Execution of Functional Queries and the Restoration of Data Generated Since the Last Backup
摘要
The possibilities of increasing the readiness of fault-tolerant cluster systems for the timely execution of functional requests are investigated. Duplicated systems containing two computers and two two-input memory nodes are considered as cluster nodes, which ensures direct accessibility of any memory node for two computers. For accelerated recovery of duplicated cluster nodes, the results of the last backup are preliminarily entered into the memory node intended to replace the failed one. As a result of this decision, after replacing a failed memory node, its information recovery requires only the entry of up-to-date information generated after the last backup. Up-to-date data can be replicated from a healthy storage node of the same cluster node that contains data generated since the last backup. The purpose of this article is to investigate the possibilities of increasing the readiness of a cluster system for the timely execution of functional requests based on the rationale for the distribution of computing resources of duplicate cluster nodes for restoring actual data and for executing functional requests. Variants of division of computing resources into information recovery of actual data and execution of functional queries are considered. Markov models of duplicated cluster nodes are proposed, on the basis of which the dependences of the system readiness for timely execution of requests on resource allocation options that have retained the operability of computing nodes to perform functional tasks and restore information in memory, including the stages of entering the results of the last backup and replicating relevant information from memory of healthy nodes. The risks of using information based on the results of the last backup in the process of servicing functional requests are assessed. The effectiveness of the proposed solutions for restoring the system after failures is evaluated by the intensity of profit from servicing information requests, taking into account the risks of using outdated information.