The Basic Design Space of Out-of-Order Execution in CPU Cores
摘要
In the quest for higher performance, the microarchitecture of client CPU cores substantially evolved in the past, resulting in many advanced microarchitectures published with various levels of detail. Although there are several publications discussing the design spaces of individual tasks of the microarchitecture, such as branch prediction (Zhou et al. in Design space exploration of TAGE branch predictor with ultra-small RAM. GLSVLSI '17: Proceedings of the Great Lakes Symposium on VLSI 2017, IEEE Micro, pp 281–286, 2017; Rodrigues et al. in Performance analysis of Big.LITTLE system with various branch prediction part of the transactions on computer systems and networks book series (TCSN). Data Science, pp 59–72, 2021) register renaming (Sima in The design space of register renaming techniques. IEEE Micro, pp 70–83, 2000; Mishev in Exploring the register renaming design space using the supersim simulator. Institute of Informatics, Faculty of Natural Sciences and Mathematics, Ss. Cyril and Methodius University in Skopje, Macedonia, CIIT, 2001) or scheduling (shelving) (Sima in The desing space of shelving J Syst Architect 863–885, 1999; Wong et al. in ACM Trans Reconfigurable Technol Syst 11(1):1–22, 2018), no published effort has been made yet to explore the design space of out-of-order execution of CPU cores. Our paper aims at closing this shortfall. In superscalar processors, out-of-order execution has four major tasks, such as register renaming, operand fetching, scheduling, and reordering, used to provide sequential consistency of instruction execution. Consequently, in spanning the design space of out-of-order execution, we have to consider the already explored design spaces of register renaming and scheduling and further feasible alternatives for implementing operand fetching and instruction reordering. Our paper considers only the FX- and FP-execution of the cores; this is why we use the designation basic design space. As shown, there are 32 alternatives of out-of-order instruction execution in CPU cores under the assumptions made. We identified which choice may be considered the most advantageous and showed which design options and development paths have been preferred in major CPU cores in the long way of their evolution.