<p>As the volume of data continues to grow and becomes increasingly distributed across geographically dispersed regions, ensuring timely access and managing these massive datasets presents significant challenges. Data replication has emerged as a key solution, enhancing data availability and reducing access latency by storing multiple copies of data across different sites, enabling users to retrieve the nearest replica efficiently. Among existing approaches, the Dynamic Hierarchical Replication (DHR) algorithm adopts a hierarchical strategy, but its reliance on a single parameter (access frequency) and its sequential search logic often lead to suboptimal placements and additional processing overhead. To overcome these limitations, we propose CORF (Cuckoo-based Optimized Replication Framework), which integrates DHR with the Cuckoo Optimization Algorithm (COA) to achieve multi-criteria, profit-driven replica placement. CORF evaluates candidate sites in parallel using parameters such as access frequency, memory availability, access cost, and locality, thereby avoiding sequential bottlenecks and enabling scalability for large-scale, HPC-class grid environments. By embedding COA’s evolutionary search logic, CORF adaptively balances exploration and exploitation, reinforcing high-performing sites while penalizing suboptimal ones, which improves both efficiency and resilience. The framework is implemented and tested in the OptorSim simulation environment and benchmarked against baseline algorithms (DHR, DHRT, and RDT). Experimental results demonstrate that CORF achieves a 16–18% reduction in job execution time and a 17–19% improvement in effective network utilization compared to baselines. These consistent gains underscore CORF’s suitability for real-time, data-intensive supercomputing workloads, where parallelism, scalability, and adaptive optimization are essential for managing massive distributed datasets.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

CORF: a cuckoo optimized replication framework for data placement in grid computing

  • Faramarz Safi-Esfahani,
  • Habib Larian,
  • Roza Majidi,
  • Ghassan Beydoun

摘要

As the volume of data continues to grow and becomes increasingly distributed across geographically dispersed regions, ensuring timely access and managing these massive datasets presents significant challenges. Data replication has emerged as a key solution, enhancing data availability and reducing access latency by storing multiple copies of data across different sites, enabling users to retrieve the nearest replica efficiently. Among existing approaches, the Dynamic Hierarchical Replication (DHR) algorithm adopts a hierarchical strategy, but its reliance on a single parameter (access frequency) and its sequential search logic often lead to suboptimal placements and additional processing overhead. To overcome these limitations, we propose CORF (Cuckoo-based Optimized Replication Framework), which integrates DHR with the Cuckoo Optimization Algorithm (COA) to achieve multi-criteria, profit-driven replica placement. CORF evaluates candidate sites in parallel using parameters such as access frequency, memory availability, access cost, and locality, thereby avoiding sequential bottlenecks and enabling scalability for large-scale, HPC-class grid environments. By embedding COA’s evolutionary search logic, CORF adaptively balances exploration and exploitation, reinforcing high-performing sites while penalizing suboptimal ones, which improves both efficiency and resilience. The framework is implemented and tested in the OptorSim simulation environment and benchmarked against baseline algorithms (DHR, DHRT, and RDT). Experimental results demonstrate that CORF achieves a 16–18% reduction in job execution time and a 17–19% improvement in effective network utilization compared to baselines. These consistent gains underscore CORF’s suitability for real-time, data-intensive supercomputing workloads, where parallelism, scalability, and adaptive optimization are essential for managing massive distributed datasets.