<p>Data is generated at high volume and speed, creating both challenges and opportunities. Cloud computing has further amplified some of these challenges, particularly in terms of storage options, availability of storage locations, and data access latency. To simplify data management and avoid the complexities of migrating data between cloud providers, most cloud users prefer subscribing to a single provider for hosting their applications and datasets. However, multicloud environments mitigate vendor lock-in risks while offering cost benefits and access to specialized services. In this paper, we propose an <b>a</b>daptive <b>da</b>ta <b>p</b>lacemen<b>t</b> framework (ADAPT) framework designed to optimize data storage costs and enhance data availability in multicloud environments. We approach the selection of optimal storage locations and data availability as classification problems. To evaluate our method, we applied four popular machine learning models and assessed their performance. Our results indicate that the XGBoost model effectively improves cost efficiency from 6.58% to 24.26% and data file availability. We then integrated XGBoost into the ADAPT framework and presented prototype results demonstrating its effectiveness.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

ADAPT: An effective data-aware multicloud data placement framework

  • John Agyekum,
  • Somnath Mazumdar,
  • Christoph Scheich

摘要

Data is generated at high volume and speed, creating both challenges and opportunities. Cloud computing has further amplified some of these challenges, particularly in terms of storage options, availability of storage locations, and data access latency. To simplify data management and avoid the complexities of migrating data between cloud providers, most cloud users prefer subscribing to a single provider for hosting their applications and datasets. However, multicloud environments mitigate vendor lock-in risks while offering cost benefits and access to specialized services. In this paper, we propose an adaptive data placement framework (ADAPT) framework designed to optimize data storage costs and enhance data availability in multicloud environments. We approach the selection of optimal storage locations and data availability as classification problems. To evaluate our method, we applied four popular machine learning models and assessed their performance. Our results indicate that the XGBoost model effectively improves cost efficiency from 6.58% to 24.26% and data file availability. We then integrated XGBoost into the ADAPT framework and presented prototype results demonstrating its effectiveness.