In this chapter, we experiment on the effectiveness of the workload-dependent multi-timescale mitigation approach for performance variability, with a cycle-accurate simulation flow we have developed. The targeted applications for the approach are highly dynamic workloads with strict real-time requirements. The state-of-the-art approaches fail to provide practically usable full guarantees for these workloads, so we consider them the most suited benchmarks. We select two representative workloads from widely different application domains. One is the lower sub-band quantization block of the ADPCM encoding application from TACLeBench Falk et al (TACLeBench: a benchmark collection to support worst-case execution time research. In: Schoeberl M (ed) 16th International Workshop on Worst-Case Execution Time Analysis (WCET 2016). OpenAccess Series in Informatics (OASIcs), vol 55, pp 2:1–2:10. Schloss Dagstuhl–Leibniz-Zentrum für Informatik, Dagstuhl, 2016). The other one is the JPEG decoding application from MiBench Guthaus et al (MiBench: a free, commercially representative embedded benchmark suite. In: Proceedings of the Fourth Annual IEEE International Workshop on Workload Characterization, pp 3–14, 2001. https://doi.org/10.1109/WWC.2001.990739). Both workloads are typical real-time signal and data processing applications, where the input of workload is a sequence of frames. Deadlines for processing consecutive frames are periodic, where the interval between two consecutive deadlines is called frame period. In conventional real-time scheduling, the finest unit is the job of processing a single frame, but in our approach, more fine-grained mitigation can be achieved by cutting the job of processing a frame into multiple TNs, each of which has an execution time of tens to hundreds of microseconds. The execution time of these workloads is highly dynamic with different input data. Therefore, they are representative of the highly dynamic workloads that our approach targets. This chapter is structured as follows: Sect. 6.1 elaborates on the experiment setup, in which the cycle-accurate simulation flow is introduced. Then, Sects. 6.2 and 6.3 discuss the case studies with the workloads of lower-band quantization of ADPCM encoding and JPEG decoding. The comparison of the results of the case studies is discussed in Sect. 6.4.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Workload-Dependent Multi-Timescale Mitigation Approach for Performance Variability: Experiments

  • Ji-Yung Lin,
  • Michalis Noltsis,
  • Dimitrios Soudris,
  • Francky Catthoor

摘要

In this chapter, we experiment on the effectiveness of the workload-dependent multi-timescale mitigation approach for performance variability, with a cycle-accurate simulation flow we have developed. The targeted applications for the approach are highly dynamic workloads with strict real-time requirements. The state-of-the-art approaches fail to provide practically usable full guarantees for these workloads, so we consider them the most suited benchmarks. We select two representative workloads from widely different application domains. One is the lower sub-band quantization block of the ADPCM encoding application from TACLeBench Falk et al (TACLeBench: a benchmark collection to support worst-case execution time research. In: Schoeberl M (ed) 16th International Workshop on Worst-Case Execution Time Analysis (WCET 2016). OpenAccess Series in Informatics (OASIcs), vol 55, pp 2:1–2:10. Schloss Dagstuhl–Leibniz-Zentrum für Informatik, Dagstuhl, 2016). The other one is the JPEG decoding application from MiBench Guthaus et al (MiBench: a free, commercially representative embedded benchmark suite. In: Proceedings of the Fourth Annual IEEE International Workshop on Workload Characterization, pp 3–14, 2001. https://doi.org/10.1109/WWC.2001.990739). Both workloads are typical real-time signal and data processing applications, where the input of workload is a sequence of frames. Deadlines for processing consecutive frames are periodic, where the interval between two consecutive deadlines is called frame period. In conventional real-time scheduling, the finest unit is the job of processing a single frame, but in our approach, more fine-grained mitigation can be achieved by cutting the job of processing a frame into multiple TNs, each of which has an execution time of tens to hundreds of microseconds. The execution time of these workloads is highly dynamic with different input data. Therefore, they are representative of the highly dynamic workloads that our approach targets. This chapter is structured as follows: Sect. 6.1 elaborates on the experiment setup, in which the cycle-accurate simulation flow is introduced. Then, Sects. 6.2 and 6.3 discuss the case studies with the workloads of lower-band quantization of ADPCM encoding and JPEG decoding. The comparison of the results of the case studies is discussed in Sect. 6.4.