<p>Molecular property prediction is central to early-stage drug discovery, where graph neural networks, pretrained string Transformers and classical descriptors each offer complementary inductive biases, yet it remains unclear when combining them improves accuracy enough to justify the added complexity, especially on the small, analog-rich datasets typical of lead optimization. Here we show that MolDualNet, a no-pretraining multi-modal architecture, is competitive on such small-data molecular-property prediction. The graph and character-level SMILES branches interact through gated bidirectional cross-attention, while expert chemical descriptors and lightweight bond-level 3D geometry are encoded by their own branches and provide auxiliary chemical priors through joint warm-up training and ablation-controlled inclusion (rather than being concatenated to the cross-attended representation). Across four standard regression endpoints (ESOL, FreeSolv, Lipophilicity, BACE) under matched per-task random 80/10/10 splitting against nine representative baselines spanning fingerprint, graph neural network, 3D-aware and pretrained SMILES Transformer families, MolDualNet attains the lowest mean RMSE on Lipophilicity (0.584), is within 1<i>σ</i> of the best method on ESOL (0.633 vs DimeNet++ 0.624) and FreeSolv (0.965 vs DimeNet++ 0.902), and remains competitive on BACE (0.734) while trailing MolFormer-XL and fingerprint baselines. On Lipophilicity, MolDualNet’s three-seed RMSE standard deviation of 0.004 is the tightest of all methods, an order of magnitude below the pretrained Transformer baselines. Under per-task Bemis–Murcko scaffold splitting, performance becomes task-dependent: MolDualNet remains competitive on small-data physicochemical endpoints but trails the 1.1 B-molecule-pretrained MolFormer-XL on data-rich drug-like tasks (Lipophilicity, BACE). Multi-seed ablations indicate that cross-attention, expert descriptors and bond-level geometry act in a task- and dataset-size dependent manner: they contribute most when limited dataset size prevents single-modality models from learning the underlying signal end-to-end, and saturate on data-rich endpoints. These results position MolDualNet as a competitive no-pretraining multi-modal architecture for analog-rich small-data molecular-property prediction.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MolDualNet as a multi-modal architecture for small-data analog-space molecular property prediction

  • Zihan Zhang,
  • Xuezhou Zhao,
  • Dan Wu

摘要

Molecular property prediction is central to early-stage drug discovery, where graph neural networks, pretrained string Transformers and classical descriptors each offer complementary inductive biases, yet it remains unclear when combining them improves accuracy enough to justify the added complexity, especially on the small, analog-rich datasets typical of lead optimization. Here we show that MolDualNet, a no-pretraining multi-modal architecture, is competitive on such small-data molecular-property prediction. The graph and character-level SMILES branches interact through gated bidirectional cross-attention, while expert chemical descriptors and lightweight bond-level 3D geometry are encoded by their own branches and provide auxiliary chemical priors through joint warm-up training and ablation-controlled inclusion (rather than being concatenated to the cross-attended representation). Across four standard regression endpoints (ESOL, FreeSolv, Lipophilicity, BACE) under matched per-task random 80/10/10 splitting against nine representative baselines spanning fingerprint, graph neural network, 3D-aware and pretrained SMILES Transformer families, MolDualNet attains the lowest mean RMSE on Lipophilicity (0.584), is within 1σ of the best method on ESOL (0.633 vs DimeNet++ 0.624) and FreeSolv (0.965 vs DimeNet++ 0.902), and remains competitive on BACE (0.734) while trailing MolFormer-XL and fingerprint baselines. On Lipophilicity, MolDualNet’s three-seed RMSE standard deviation of 0.004 is the tightest of all methods, an order of magnitude below the pretrained Transformer baselines. Under per-task Bemis–Murcko scaffold splitting, performance becomes task-dependent: MolDualNet remains competitive on small-data physicochemical endpoints but trails the 1.1 B-molecule-pretrained MolFormer-XL on data-rich drug-like tasks (Lipophilicity, BACE). Multi-seed ablations indicate that cross-attention, expert descriptors and bond-level geometry act in a task- and dataset-size dependent manner: they contribute most when limited dataset size prevents single-modality models from learning the underlying signal end-to-end, and saturate on data-rich endpoints. These results position MolDualNet as a competitive no-pretraining multi-modal architecture for analog-rich small-data molecular-property prediction.