Machine learning-based molecular property prediction has promising potential in drug discovery and materials design. The lack of labeled data, however, poses a great challenge. By leveraging self-supervised pre-training on unlabeled data, molecular representation learning can alleviate this problem. Most existing pre-training frameworks utilize either a single-mode molecular representation or unaligned multiple representations, limiting the expressive power of molecular features. In this paper, we propose a pre-training framework, called CRAFT (Consistent RepresentAtional Fusion of Three molecular modalities), which can efficiently fuse multiple molecular representations, including SELFIES, 2D molecular graphs, and IUPAC names. The three molecular representations are aligned through contrastive learning with a momentum model and fused through cross-attention mechanisms. By utilizing this consistent fusion approach, CRAFT is able to learn domain knowledge about chemistry with unsupervised pre-training on a small dataset, and it outperforms the state-of-the-art models in five downstream tasks of molecular property prediction.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

CRAFT: Consistent Representational Fusion of Three Molecular Modalities

  • Kun Hu,
  • Mingjun Xiao,
  • Qian Cao,
  • Binlan Wu

摘要

Machine learning-based molecular property prediction has promising potential in drug discovery and materials design. The lack of labeled data, however, poses a great challenge. By leveraging self-supervised pre-training on unlabeled data, molecular representation learning can alleviate this problem. Most existing pre-training frameworks utilize either a single-mode molecular representation or unaligned multiple representations, limiting the expressive power of molecular features. In this paper, we propose a pre-training framework, called CRAFT (Consistent RepresentAtional Fusion of Three molecular modalities), which can efficiently fuse multiple molecular representations, including SELFIES, 2D molecular graphs, and IUPAC names. The three molecular representations are aligned through contrastive learning with a momentum model and fused through cross-attention mechanisms. By utilizing this consistent fusion approach, CRAFT is able to learn domain knowledge about chemistry with unsupervised pre-training on a small dataset, and it outperforms the state-of-the-art models in five downstream tasks of molecular property prediction.