<p>Accurate and generalizable molecular representation remains challenging due to data scarcity, uneven distribution, and substantial heterogeneity across chemical domains, which often leads to negative transfer and limits model generalization in transfer learning. To address these issues, this study proposes a hierarchical Domain-to-Task Adaptation (D2TA) transfer learning framework for molecular learning. The framework is designed to extract universal chemical features from large-scale chemical data, generate domain-specific features through fragment-driven molecular generation strategy and transfer learning, and ultimately adapt these representations into task-specific features tailored to downstream objectives. Using the prediction of enthalpy of formation (EOF) in energetic molecules as a case study, the D2TA framework achieves a prediction accuracy of <i>R</i>² = 0.958, with a mean absolute error (MAE) of 81.54 kJ·mol⁻¹ and a root mean square error (RMSE) of 115.69 kJ·mol⁻¹, representing an improvement of approximately 5 kJ mol⁻¹ in MAE compared with the first-stage transfer learning. In the β-site APP Cleaving Enzyme (BACE) molecular activity classification task, the ROC-AUC (%) reaches 92%, improving by 13% relative to the first-stage transfer learning. Extensive evaluations across multiple downstream tasks and benchmark datasets further demonstrate the effectiveness of D2TA in improving predictive accuracy. These results demonstrate that D2TA effectively enhances the adaptability and generalization of molecular representation across heterogeneous domains, which provides a new approach for improving molecular learning with small sample sizes in various chemical domains.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Hierarchical domain-to-task adaptation for generalizable molecular transfer learning

  • Ruihui Wang,
  • Linhu Pan,
  • Mingren Fan,
  • Yi Wang,
  • Xiujuan Qi,
  • Siwei Song,
  • Qinghua Zhang

摘要

Accurate and generalizable molecular representation remains challenging due to data scarcity, uneven distribution, and substantial heterogeneity across chemical domains, which often leads to negative transfer and limits model generalization in transfer learning. To address these issues, this study proposes a hierarchical Domain-to-Task Adaptation (D2TA) transfer learning framework for molecular learning. The framework is designed to extract universal chemical features from large-scale chemical data, generate domain-specific features through fragment-driven molecular generation strategy and transfer learning, and ultimately adapt these representations into task-specific features tailored to downstream objectives. Using the prediction of enthalpy of formation (EOF) in energetic molecules as a case study, the D2TA framework achieves a prediction accuracy of R² = 0.958, with a mean absolute error (MAE) of 81.54 kJ·mol⁻¹ and a root mean square error (RMSE) of 115.69 kJ·mol⁻¹, representing an improvement of approximately 5 kJ mol⁻¹ in MAE compared with the first-stage transfer learning. In the β-site APP Cleaving Enzyme (BACE) molecular activity classification task, the ROC-AUC (%) reaches 92%, improving by 13% relative to the first-stage transfer learning. Extensive evaluations across multiple downstream tasks and benchmark datasets further demonstrate the effectiveness of D2TA in improving predictive accuracy. These results demonstrate that D2TA effectively enhances the adaptability and generalization of molecular representation across heterogeneous domains, which provides a new approach for improving molecular learning with small sample sizes in various chemical domains.