<p>Fast yet accurate performance and timing prediction of complex parallel data flow applications on multiprocessor systems remains a challenging discipline. The main reason is that the applications contain numerous degrees of parallelism (task, pipeline, data) but they necessitate sharing resources, such as communication buses and memories, within execution platforms. Executing such applications on resource-limited platforms leads to timing interferences that are difficult to accurately express in pure analytical approaches or to excessive simulation duration for cycle-accurate models. In this work, we propose a message-level communication model for fast yet accurate performance prediction for data flow applications executed on MPSoCs with shared memories and buses. This approach combines a high level executable model of the communication infrastructure with a formal description of the synchronization instants related to low-level communication mechanisms. This combination significantly reduces the number of simulation events while still accurately predicting the effects of contention at shared resources. We evaluated our work against measurements from a real prototype and cycle-accurate performance prediction models on two case-studies from the computer vision domain. We illustrated how the computational complexity of our approach can be adapted to deliver high simulation efficiency. In our experiment, we achieved an average accuracy of <InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(\varvec{99\%}\)</EquationSource> </InlineEquation> in latency prediction compared to real implementation for the different use-cases and mappings we considered. In the experiments where the approach is used with limited computational complexity, simulation speed-ups of an order of magnitude of <InlineEquation ID="IEq2"> <EquationSource Format="TEX">\(\varvec{10}^{\varvec{4}}\)</EquationSource> </InlineEquation> compared to cycle-accurate models are achieved, while maintaining a fully acceptable accuracy (more than <InlineEquation ID="IEq3"> <EquationSource Format="TEX">\(\varvec{98\%}\)</EquationSource> </InlineEquation>). Good suitability for fast and accurate exploration of the design space is demonstrated by the proposed models.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Hybrid Message-level Modeling Approach for Fast Yet Accurate Simulation of Multiprocessor Shared Bus Effects on Data Flow Applications Execution

  • Sébastien Le Nours,
  • Hai-Dang Vu,
  • Ralf Stemmer,
  • Sébastien Pillement

摘要

Fast yet accurate performance and timing prediction of complex parallel data flow applications on multiprocessor systems remains a challenging discipline. The main reason is that the applications contain numerous degrees of parallelism (task, pipeline, data) but they necessitate sharing resources, such as communication buses and memories, within execution platforms. Executing such applications on resource-limited platforms leads to timing interferences that are difficult to accurately express in pure analytical approaches or to excessive simulation duration for cycle-accurate models. In this work, we propose a message-level communication model for fast yet accurate performance prediction for data flow applications executed on MPSoCs with shared memories and buses. This approach combines a high level executable model of the communication infrastructure with a formal description of the synchronization instants related to low-level communication mechanisms. This combination significantly reduces the number of simulation events while still accurately predicting the effects of contention at shared resources. We evaluated our work against measurements from a real prototype and cycle-accurate performance prediction models on two case-studies from the computer vision domain. We illustrated how the computational complexity of our approach can be adapted to deliver high simulation efficiency. In our experiment, we achieved an average accuracy of \(\varvec{99\%}\) in latency prediction compared to real implementation for the different use-cases and mappings we considered. In the experiments where the approach is used with limited computational complexity, simulation speed-ups of an order of magnitude of \(\varvec{10}^{\varvec{4}}\) compared to cycle-accurate models are achieved, while maintaining a fully acceptable accuracy (more than \(\varvec{98\%}\) ). Good suitability for fast and accurate exploration of the design space is demonstrated by the proposed models.