A trimodal protein language model enables advanced protein searches
摘要
ProTrek unifies protein sequence, structure and natural language function in a trimodal language model through contrastive learning, enabling comprehensive searches between any two modalities, including within modality. ProTrek surpasses current alignment tools (for example, Foldseek and MMseqs2) in speed and accuracy for identifying functionally related proteins. Computational and wet-lab experimental validations show that the ProTrek server (www.search-protrek.com), with precomputed embeddings for over 5 billion proteins, efficiently processes and analyzes large-scale protein repositories.