Grid workflows across applications show significant differences in resource usage patterns. Sharing resources concurrently comes with the challenge of fluctuating allocations and varying loads on the execution machines. To achieve fair behaviour with other workflows and prevent job terminations due to resource overconsumption, controlled and predictable CPU usage is crucial. This paper evaluates the impact of incorporating CPU constraining mechanisms within the Grid middleware itself, specifically targeting a heterogeneous and dynamic environment like the LHC ALICE experiment Grid. Leveraging existing tools in the operating systems of the execution machines, CPU limitations can be configured at various levels, ranging from pinning the execution on specific CPU cores to setting a dedicated CPU bandwidth. We present performance results demonstrating that CPU-constrained execution environments lead to highly predictable CPU efficiency for jobs. Additionally, the effects of NUMA-aware scheduling for processes spawned by these jobs are explored, evaluating how it impacts performance on different execution environments.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluation of CPU Constraining Mechanisms in the LHC ALICE Experiment Grid

  • Marta Bertran Ferrer,
  • Costin Grigoras,
  • Maksim Storetvedt,
  • Rosa M Badia,
  • Latchezar Betev

摘要

Grid workflows across applications show significant differences in resource usage patterns. Sharing resources concurrently comes with the challenge of fluctuating allocations and varying loads on the execution machines. To achieve fair behaviour with other workflows and prevent job terminations due to resource overconsumption, controlled and predictable CPU usage is crucial. This paper evaluates the impact of incorporating CPU constraining mechanisms within the Grid middleware itself, specifically targeting a heterogeneous and dynamic environment like the LHC ALICE experiment Grid. Leveraging existing tools in the operating systems of the execution machines, CPU limitations can be configured at various levels, ranging from pinning the execution on specific CPU cores to setting a dedicated CPU bandwidth. We present performance results demonstrating that CPU-constrained execution environments lead to highly predictable CPU efficiency for jobs. Additionally, the effects of NUMA-aware scheduling for processes spawned by these jobs are explored, evaluating how it impacts performance on different execution environments.