<p>Accelerators in mission-critical applications consume substantial power even during low workloads. Due to strict availability requirements, they cannot be stopped or restarted. This leads to unnecessary power consumption during low-load periods. We propose an accelerator pool architecture that aggregates accelerators beyond traditional server constraints. This approach enables dynamic resource allocation through middleware-level management while keeping applications running. It allows online resource scaling rather than relying on application start/stop scheduling. Our contributions include: (1) a novel pooling architecture for resource-intensive applications (virtual Radio Access Network (vRAN), generative Artificial Intelligence (AI), automated driving) that reduces power consumption while maintaining performance, (2) a comprehensive evaluation of the architecture’s practicality for vRAN use cases with identified implementation challenges and mitigation strategies, (3) a mathematical model formulating the accelerator allocation problem as a bin-packing optimization, and (4) empirical validation through comparative simulations. Preliminary simulations using vRAN demonstrate up to 64.3% reduction in operating accelerators compared to conventional architectures. They also show a 29.0% reduction compared to existing server-internal sharing approaches. While this architectural concept offers limited benefits for applications with consistently high accelerator utilization, it shows significant potential in scenarios with geographical load imbalances or temporal variations.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Accelerator pool: accelerator sharing architecture for energy efficiency

  • Shogo Saito,
  • Ko Natori,
  • Ikuo Otani,
  • Kei Fujimoto

摘要

Accelerators in mission-critical applications consume substantial power even during low workloads. Due to strict availability requirements, they cannot be stopped or restarted. This leads to unnecessary power consumption during low-load periods. We propose an accelerator pool architecture that aggregates accelerators beyond traditional server constraints. This approach enables dynamic resource allocation through middleware-level management while keeping applications running. It allows online resource scaling rather than relying on application start/stop scheduling. Our contributions include: (1) a novel pooling architecture for resource-intensive applications (virtual Radio Access Network (vRAN), generative Artificial Intelligence (AI), automated driving) that reduces power consumption while maintaining performance, (2) a comprehensive evaluation of the architecture’s practicality for vRAN use cases with identified implementation challenges and mitigation strategies, (3) a mathematical model formulating the accelerator allocation problem as a bin-packing optimization, and (4) empirical validation through comparative simulations. Preliminary simulations using vRAN demonstrate up to 64.3% reduction in operating accelerators compared to conventional architectures. They also show a 29.0% reduction compared to existing server-internal sharing approaches. While this architectural concept offers limited benefits for applications with consistently high accelerator utilization, it shows significant potential in scenarios with geographical load imbalances or temporal variations.