<p>Predicting protein-ligand binding affinity from three-dimensional (3D) structural data is a central task in structure-based drug discovery, yet it remains challenging due to limited data availability, structural complexity, and the sparse nature of 3D molecular representations. In this study, we investigate the application of vision transformers (ViTs) to the problem of affinity prediction. Unlike other neural networks used in this problem, the ViT framework can capture global, long-range interactions across the entire protein-ligand complex via self-attention, without relying on local receptive fields or predefined interaction cutoffs. We evaluate this advantage in representation of spatial information across two benchmark datasets, demonstrating competitive performance and, in some cases, surpassing state-of-the-art models. We study the model’s behavior using explainable AI (XAI) techniques, revealing that spatially proximal patches with similar attention scores cluster around biologically relevant regions, confirming the model’s ability to capture key interaction features. Furthermore, we show that data augmentation strategies can yield performance improvements, highlighting the potential for further enhancement. Despite challenges related to data sparsity and conformational variability, ViTs show strong performance and high robustness in structure-based affinity prediction tasks. Our findings underscore their effectiveness in learning spatial patterns and suggest broader applicability to related tasks, such as protein-protein or protein-nucleic acid interaction modeling.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Application of vision transformers to protein-ligand affinity prediction

  • Jakub Poziemski,
  • Pawel Siedlecki

摘要

Predicting protein-ligand binding affinity from three-dimensional (3D) structural data is a central task in structure-based drug discovery, yet it remains challenging due to limited data availability, structural complexity, and the sparse nature of 3D molecular representations. In this study, we investigate the application of vision transformers (ViTs) to the problem of affinity prediction. Unlike other neural networks used in this problem, the ViT framework can capture global, long-range interactions across the entire protein-ligand complex via self-attention, without relying on local receptive fields or predefined interaction cutoffs. We evaluate this advantage in representation of spatial information across two benchmark datasets, demonstrating competitive performance and, in some cases, surpassing state-of-the-art models. We study the model’s behavior using explainable AI (XAI) techniques, revealing that spatially proximal patches with similar attention scores cluster around biologically relevant regions, confirming the model’s ability to capture key interaction features. Furthermore, we show that data augmentation strategies can yield performance improvements, highlighting the potential for further enhancement. Despite challenges related to data sparsity and conformational variability, ViTs show strong performance and high robustness in structure-based affinity prediction tasks. Our findings underscore their effectiveness in learning spatial patterns and suggest broader applicability to related tasks, such as protein-protein or protein-nucleic acid interaction modeling.