Viruses are of great diversity and variability and impact human society deeply and broadly. Advances in sequencing technologies enable easier detection and analysis of environmental samples, aiding clinical and virology research. Efficient viral classification remains a key challenge, with ongoing methodological improvements pursuing more reliable outputs. Time-consuming alignment-based methods are becoming obsolete with the explosion of data, while alignment-free especially machine learning methods take the lead. Along this line, we introduce VirB, a hierarchical classification method based on the latest BERT model ModernBERT, to classify the viral contigs or genomes at the order and family level. Integration of BPE and Transformer architecture with optimized embedding and attention strategies enables our model to process ultra-long sequences effectively, demonstrating superior performance over all compared methods in test cases. Experiments conducted with diverse real-world datasets confirmed the generalization power of the proposed model by presenting remarkable performance in predicting unseen sequences and sequences with noises.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

VirB: A Virus Hierarchical Classification Method Based on ModernBERT

  • Haizhen Huang,
  • Haodi Feng,
  • Daming Zhu

摘要

Viruses are of great diversity and variability and impact human society deeply and broadly. Advances in sequencing technologies enable easier detection and analysis of environmental samples, aiding clinical and virology research. Efficient viral classification remains a key challenge, with ongoing methodological improvements pursuing more reliable outputs. Time-consuming alignment-based methods are becoming obsolete with the explosion of data, while alignment-free especially machine learning methods take the lead. Along this line, we introduce VirB, a hierarchical classification method based on the latest BERT model ModernBERT, to classify the viral contigs or genomes at the order and family level. Integration of BPE and Transformer architecture with optimized embedding and attention strategies enables our model to process ultra-long sequences effectively, demonstrating superior performance over all compared methods in test cases. Experiments conducted with diverse real-world datasets confirmed the generalization power of the proposed model by presenting remarkable performance in predicting unseen sequences and sequences with noises.