Background <p>Infectious diseases continue to pose unprecedented challenges to public health and the global economy. Virulence factors (VFs) enable pathogens to adhere, reproduce, and cause damage to host cells, while antibiotic resistance genes (ARGs) enable pathogens to withstand treatments that would otherwise be effective. The concurrent identification of VFs and ARGs is crucial for efficient pathogen surveillance. However, existing tools for predicting VFs or ARGs typically suffer from high false negative rates and limitations in identifying only high-identity genes against known reference VF or ARG databases.</p> Results <p>To address these challenges, we developed SEVA, an advanced model that integrates protein language models (pLMs) with structural and evolutionary protein features to predict VFs and ARGs from genome sequencing data. Integrating multiple homologous sequences can identify latent virulence or drug resistance caused by site mutations, reducing false negative rates. Meanwhile, the protein structure remains conserved despite the low sequence identity in some functional domains of VFs or ARGs. The aggregate of protein structure information further improves the identification abilities of VF and ARG. In addition, pLMs enable the model to capture high-dimensional feature representations more effectively. SEVA rigorously collected three datasets with over 20,000 genes and five reference databases. It outperforms state-of-the-art methods, including Diamond, VRprofile, FoldSeek, PreVFs-RG, PLM-ARG, ARG-BERT, and HyperVR, achieving an accuracy of 97.13% and confirming the efficacy of its key components, such as refined feature selection and multiple sequence alignment subsampling.</p> Conclusion <p>SEVA takes protein sequences as input and derives evolutionary, structural, and statistical representations for prediction, making our model a reliable tool for VF and ARG prediction. This capability is particularly valuable in epidemic prevention and control, where accurate identification of VFs and ARGs is crucial. By providing concurrent and reliable predictions of VFs and ARGs, SEVA enhances our ability to respond to microbial threats effectively. This finding supports robust efforts to mitigate the spread of infectious diseases and safeguard public health, addressing a critical gap in contemporary epidemic response strategies. The SEVA model and data are available at <a href="https://github.com/kaiqili2/SEVA">https://github.com/kaiqili2/SEVA</a>.</p> <p><MediaObject ID="MOESM2"><VideoObject FileRef="MediaObjects/40168_2026_2467_MOESM2_ESM.mp4" VideoID="Eu-nM7NJcwNYi196_J2PBu"><Caption Language="En" xml:lang="en"><CaptionContent><p>Video Abstract</p></CaptionContent></Caption></VideoObject></MediaObject></p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

SEVA: structural and evolutionary feature integration for predicting virulence factors and antibiotic resistance genes

  • Kaiqi Li,
  • Xin Peng,
  • Xiuwei Qian,
  • Shuaicheng Li,
  • Xianglilan Zhang

摘要

Background

Infectious diseases continue to pose unprecedented challenges to public health and the global economy. Virulence factors (VFs) enable pathogens to adhere, reproduce, and cause damage to host cells, while antibiotic resistance genes (ARGs) enable pathogens to withstand treatments that would otherwise be effective. The concurrent identification of VFs and ARGs is crucial for efficient pathogen surveillance. However, existing tools for predicting VFs or ARGs typically suffer from high false negative rates and limitations in identifying only high-identity genes against known reference VF or ARG databases.

Results

To address these challenges, we developed SEVA, an advanced model that integrates protein language models (pLMs) with structural and evolutionary protein features to predict VFs and ARGs from genome sequencing data. Integrating multiple homologous sequences can identify latent virulence or drug resistance caused by site mutations, reducing false negative rates. Meanwhile, the protein structure remains conserved despite the low sequence identity in some functional domains of VFs or ARGs. The aggregate of protein structure information further improves the identification abilities of VF and ARG. In addition, pLMs enable the model to capture high-dimensional feature representations more effectively. SEVA rigorously collected three datasets with over 20,000 genes and five reference databases. It outperforms state-of-the-art methods, including Diamond, VRprofile, FoldSeek, PreVFs-RG, PLM-ARG, ARG-BERT, and HyperVR, achieving an accuracy of 97.13% and confirming the efficacy of its key components, such as refined feature selection and multiple sequence alignment subsampling.

Conclusion

SEVA takes protein sequences as input and derives evolutionary, structural, and statistical representations for prediction, making our model a reliable tool for VF and ARG prediction. This capability is particularly valuable in epidemic prevention and control, where accurate identification of VFs and ARGs is crucial. By providing concurrent and reliable predictions of VFs and ARGs, SEVA enhances our ability to respond to microbial threats effectively. This finding supports robust efforts to mitigate the spread of infectious diseases and safeguard public health, addressing a critical gap in contemporary epidemic response strategies. The SEVA model and data are available at https://github.com/kaiqili2/SEVA.

Video Abstract