<p>Accurate prediction of toluene/water partition coefficients of neutral species is crucial in drug discovery and separation processes; however, data-driven modeling of these coefficients remains challenging due to limited available experimental data. To address the limitation of available data, we apply multi-fidelity learning approaches leveraging a quantum chemical dataset (low fidelity) of approximately 9000 entries generated by COSMO-RS and an experimental dataset (high fidelity) of about 250 entries collected from the literature. We explore the <i>transfer learning</i>, <i>feature-augmented learning</i>, and <i>multi-target learning</i> approaches in combination with graph neural networks, validating them on two external datasets: one with molecules similar to training data (EXT-Zamora) and one with more challenging molecules (EXT-SAMPL9). Our results show that <i>multi-target learning</i> significantly improves predictive accuracy, achieving a root-mean-square error of 0.44 <InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="13321_2025_1057_Article_IEq1.gif" Format="GIF" Height="17" Rendition="HTML" Resolution="72" Type="Linedraw" Width="39" /> </InlineMediaObject> <EquationSource Format="TEX">\(\log {P}\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mo>log</mo> <mi>P</mi> </mrow> </math></EquationSource> </InlineEquation> units for the EXT-Zamora, compared to a root-mean-square error of 0.63 <InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="13321_2025_1057_Article_IEq1.gif" Format="GIF" Height="17" Rendition="HTML" Resolution="72" Type="Linedraw" Width="39" /> </InlineMediaObject> <EquationSource Format="TEX">\(\log {P}\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mo>log</mo> <mi>P</mi> </mrow> </math></EquationSource> </InlineEquation> units for single-task models. For the EXT-SAMPL9 dataset, <i>multi-target learning</i> achieves a root-mean-square error of 1.02 <InlineEquation ID="IEq3"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="13321_2025_1057_Article_IEq1.gif" Format="GIF" Height="17" Rendition="HTML" Resolution="72" Type="Linedraw" Width="39" /> </InlineMediaObject> <EquationSource Format="TEX">\(\log {P}\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mo>log</mo> <mi>P</mi> </mrow> </math></EquationSource> </InlineEquation> units, indicating reasonable performance even for more complex molecular structures. These findings highlight the potential of multi-fidelity learning approaches that leverage quantum chemical data to improve toluene/water partition coefficient predictions and address challenges posed by limited experimental data. We expect the applicability of the methods used beyond just toluene/water partition coefficients.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-fidelity graph neural networks for predicting toluene/water partition coefficients

  • Thomas Nevolianis,
  • Jan G. Rittig,
  • Alexander Mitsos,
  • Kai Leonhard

摘要

Accurate prediction of toluene/water partition coefficients of neutral species is crucial in drug discovery and separation processes; however, data-driven modeling of these coefficients remains challenging due to limited available experimental data. To address the limitation of available data, we apply multi-fidelity learning approaches leveraging a quantum chemical dataset (low fidelity) of approximately 9000 entries generated by COSMO-RS and an experimental dataset (high fidelity) of about 250 entries collected from the literature. We explore the transfer learning, feature-augmented learning, and multi-target learning approaches in combination with graph neural networks, validating them on two external datasets: one with molecules similar to training data (EXT-Zamora) and one with more challenging molecules (EXT-SAMPL9). Our results show that multi-target learning significantly improves predictive accuracy, achieving a root-mean-square error of 0.44 \(\log {P}\) log P units for the EXT-Zamora, compared to a root-mean-square error of 0.63 \(\log {P}\) log P units for single-task models. For the EXT-SAMPL9 dataset, multi-target learning achieves a root-mean-square error of 1.02 \(\log {P}\) log P units, indicating reasonable performance even for more complex molecular structures. These findings highlight the potential of multi-fidelity learning approaches that leverage quantum chemical data to improve toluene/water partition coefficient predictions and address challenges posed by limited experimental data. We expect the applicability of the methods used beyond just toluene/water partition coefficients.