Contrastive Fitness Learning: Reprogramming Protein Language Models for Low-N Learning of Protein Fitness Landscape
摘要
Machine learning (ML) is revolutionizing our ability to model the fitness landscape of protein sequences. Recently, the protein language model (pLM) has become the foundation of state-of-the-art ML solutions for many problems in protein biology. However, significant challenges remain in leveraging pLMs for protein fitness prediction, in part due to the disparity between the scarce number of sequences functionally characterized by high-throughput assays and the massive data samples required for training large pLMs. To bridge this gap, we introduce Contrastive Fitness Learning (ConFit), a pLM-based ML method for learning the protein fitness landscape with limited (low-N) experimental fitness measurements as training data (Fig. 1). We propose a novel contrastive learning strategy to fine-tune the pre-trained pLM, tailoring it to achieve protein-specific fitness prediction while avoiding overfitting.