<p>Efficient hyperparameter tuning and model training in distributed heterogeneous environments remain challenging due to resource underutilization, training staleness, and checkpointing bottlenecks. This study addresses these issues by enhancing the Ray Tune framework with three key strategies. First, we propose an Automatic Heterogeneous Resource Allocation Strategy, which dynamically adapts resource distribution based on Node configurations, optimizing utilization while reducing manual setup complexity. Second, our Heterogeneous Resource Fusion Scheduling Algorithm classifies Workers by computational speed and employs generation-aware scheduling to minimize staleness, ensuring synchronized progress among Trials and enhancing Population-Based Training evolution. Third, an Optimized Checkpoint Management Mechanism integrates RAM and disk storage, reducing access delays while maintaining robust persistence. Experimental results validate the effectiveness of our approach. Optimizing checkpoint storage location reduces training time by 1.22×, while our heterogeneous resource fusion strategy improves overall performance by 2.36×, demonstrating enhanced workload balancing and resource efficiency. To foster further research, we open-source our implementation within Ray Tune, providing a scalable and adaptive solution for distributed hyperparameter optimization.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Efficient configuration of heterogeneous resources and task scheduling strategies in deep learning auto-tuning systems

  • Pao-Yi Ken,
  • Chao-Chin Wu

摘要

Efficient hyperparameter tuning and model training in distributed heterogeneous environments remain challenging due to resource underutilization, training staleness, and checkpointing bottlenecks. This study addresses these issues by enhancing the Ray Tune framework with three key strategies. First, we propose an Automatic Heterogeneous Resource Allocation Strategy, which dynamically adapts resource distribution based on Node configurations, optimizing utilization while reducing manual setup complexity. Second, our Heterogeneous Resource Fusion Scheduling Algorithm classifies Workers by computational speed and employs generation-aware scheduling to minimize staleness, ensuring synchronized progress among Trials and enhancing Population-Based Training evolution. Third, an Optimized Checkpoint Management Mechanism integrates RAM and disk storage, reducing access delays while maintaining robust persistence. Experimental results validate the effectiveness of our approach. Optimizing checkpoint storage location reduces training time by 1.22×, while our heterogeneous resource fusion strategy improves overall performance by 2.36×, demonstrating enhanced workload balancing and resource efficiency. To foster further research, we open-source our implementation within Ray Tune, providing a scalable and adaptive solution for distributed hyperparameter optimization.