Algorithms for Short-Read Viral Haplotype Reconstruction: Challenges, Solutions, and Perspectives
摘要
RNA viruses, such as HIV, HCV, and SARS-CoV-2, show high levels of intrahost genetic diversity. Many different haplotypes can be present in a single infection, which can be studied using next-generation sequencing. However, full-length haplotype reconstruction from short reads is computationally challenging due to the presence of low-frequency mutants, as well as sequencing errors. Moreover, reads may not be long enough to span regions between neighboring mutations. Finally, the sequencing depths needed to discover such low-frequency mutants result in large datasets, which require highly efficient algorithms. In this review, we provide an overview of current strategies to address these challenges and identify potential directions for increasing the accuracy and efficiency of viral haplotype reconstruction. Such developments will be key to advancing our understanding of viral evolution, improving treatment strategies, and informing public health interventions.