<p>With high-speed technological advancements and Next-Generation Sequencing (NGS), the discovered genome dataset (i.e., DNA/RNA/Protein) is increasing exponentially, magnifying every month. Timely analysis of the genome dataset is very important for understanding biological activities and drug development. However, due to the vast amounts of sequences and the complex structure of genome datasets, the storage and timely analysis of genome datasets is becoming challenging for traditional analysis techniques.&#xa0;Distributed and cluster computing platforms are becoming significant for big data analytics and are now required in computational biology. This paper presents a computational model named Sprak-Pi-DNN, based on a parallel deep neural network for the timely classification of large RNA sequences into piRNAs and non-piRNAs. The proposed model takes advantage of parallel and distributed computing platforms. The performance of the proposed Sprak-Pi-DNN was extensively evaluated in two parts. In the first part, we compare and analyze the performance of two widely used cluster-based big data analytics platforms (Apache Hadoop and Apache Spark) on the same benchmark dataset. The computational-based metrics include computation times, speedup, and scalability. The second part assessed the proposed model’s effectiveness using performance metrics such as accuracy, specificity, sensitivity, and Matthews’s correlation coefficient. We employed the PseKNC algorithm for feature extraction with varied K-sizes. The experimental results revealed that Apache Spark’s execution time performed better and faster than Apache Hadoop. Moreover, the evaluation results in both cases showed that the proposed model improved computation speedup without affecting accuracy.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Optimizing performance of parallel computing platforms for large-scale genome data analysis

  • Sumaiya Noor,
  • Hamid Hussain Awan,
  • Amber Sarwar Hashmi,
  • Aamir Saeed,
  • Salman Khan,
  • Salman A. AlQahtani

摘要

With high-speed technological advancements and Next-Generation Sequencing (NGS), the discovered genome dataset (i.e., DNA/RNA/Protein) is increasing exponentially, magnifying every month. Timely analysis of the genome dataset is very important for understanding biological activities and drug development. However, due to the vast amounts of sequences and the complex structure of genome datasets, the storage and timely analysis of genome datasets is becoming challenging for traditional analysis techniques. Distributed and cluster computing platforms are becoming significant for big data analytics and are now required in computational biology. This paper presents a computational model named Sprak-Pi-DNN, based on a parallel deep neural network for the timely classification of large RNA sequences into piRNAs and non-piRNAs. The proposed model takes advantage of parallel and distributed computing platforms. The performance of the proposed Sprak-Pi-DNN was extensively evaluated in two parts. In the first part, we compare and analyze the performance of two widely used cluster-based big data analytics platforms (Apache Hadoop and Apache Spark) on the same benchmark dataset. The computational-based metrics include computation times, speedup, and scalability. The second part assessed the proposed model’s effectiveness using performance metrics such as accuracy, specificity, sensitivity, and Matthews’s correlation coefficient. We employed the PseKNC algorithm for feature extraction with varied K-sizes. The experimental results revealed that Apache Spark’s execution time performed better and faster than Apache Hadoop. Moreover, the evaluation results in both cases showed that the proposed model improved computation speedup without affecting accuracy.