Pre-trained molecular language models with random functional group masking
摘要
Recent advancements in computational chemistry utilize transformer-based models pre-trained on Simplified Molecular Input Line Entry System (SMILES) sequences to predict molecular properties. To improve upon existing methods, we propose MLM-FG, a molecular language model with a novel pre-training strategy that randomly masks subsequences corresponding to chemically significant functional groups. This technique compels the model to better infer molecular structures and properties by learning the context of these key units. Extensive evaluations across 11 benchmark tasks demonstrate the superiority of MLM-FG, outperforming existing SMILES- and graph-based models in 9 of the 11 tasks. Remarkably, MLM-FG surpasses even some 3D-graph-based models, highlighting its exceptional capacity for representation learning without explicit 3D structural information. These results indicate that MLM-FG effectively learns to interpret molecular properties from SMILES, offering a powerful new tool for computational chemistry and related disciplines.