Objective <p>Amorphous solid dispersion (ASD) is widely utilized to enhance the solubility and bioavailability of water-insoluble drugs. However, conventional experimental approaches for ASD development are often resource-intensive and time-consuming. Machine learning (ML) algorithms have great potential to predict ASD formulations but face the challenge of extensive data to construct reliable models. Current study aims to predict the formation of both binary and ternary ASD by combined high-throughput screening (HTS) and ML approaches.</p> Methods <p>Micro-quantity HTS was conducted to generate 1272 binary and ternary solid dispersions using solvent evaporation method. The Powder X-Ray Diffraction (PXRD) was used to characterize the amorphous state of formulations. The results indicated that 188 formulations successfully formed amorphous solid dispersions (ASDs), while 1084 resulted in crystalline formations. Models development employed nested cross-validation with four algorithms: Light Gradient Boosting Machine (LGBM), Random Forest (RF), Support Vector Machine (SVM), and Multi-Layer Perceptron (MLP).</p> Results <p>The RF model for ASD formation achieved 96.7% accuracy on the in-house HTS dataset, with a precision of approximately 87.9% and an F1 score of 83.6%. Furthermore, the RF model trained with milligram-scale HTS experimental data could effectively predict the large-scale ASD formulations from the literature, highlighting its promise as a powerful tool for advancing ASD prediction.</p> Conclusion <p>In summary, the combination of HTS experiments and ML techniques provides a valuable reference framework for ASD development, greatly minimizing both time and material usage in the selection of formulations during the early stages of drug discovery with a limited quantity of API.</p> Graphical Abstract <p>The workflow of this study involves a micro-quantity HTS experiment that generated over a thousand homogeneous data points for ML model development. The RF model demonstrated strong performance on a large-scale external dataset for ASD prediction, significantly saving time and material usage in formulation development, particularly during early drug discovery.</p> <p></p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Combining High-Throughput Screening and Machine Learning to Predict the Formation of Both Binary and Ternary Amorphous Solid Dispersion Formulations for Early Drug Discovery and Development

  • Tianshu Lu,
  • Yiyang Wu,
  • Ping Xiong,
  • Hao Zhong,
  • Yang Ding,
  • Haifeng Li,
  • Defang Ouyang

摘要

Objective

Amorphous solid dispersion (ASD) is widely utilized to enhance the solubility and bioavailability of water-insoluble drugs. However, conventional experimental approaches for ASD development are often resource-intensive and time-consuming. Machine learning (ML) algorithms have great potential to predict ASD formulations but face the challenge of extensive data to construct reliable models. Current study aims to predict the formation of both binary and ternary ASD by combined high-throughput screening (HTS) and ML approaches.

Methods

Micro-quantity HTS was conducted to generate 1272 binary and ternary solid dispersions using solvent evaporation method. The Powder X-Ray Diffraction (PXRD) was used to characterize the amorphous state of formulations. The results indicated that 188 formulations successfully formed amorphous solid dispersions (ASDs), while 1084 resulted in crystalline formations. Models development employed nested cross-validation with four algorithms: Light Gradient Boosting Machine (LGBM), Random Forest (RF), Support Vector Machine (SVM), and Multi-Layer Perceptron (MLP).

Results

The RF model for ASD formation achieved 96.7% accuracy on the in-house HTS dataset, with a precision of approximately 87.9% and an F1 score of 83.6%. Furthermore, the RF model trained with milligram-scale HTS experimental data could effectively predict the large-scale ASD formulations from the literature, highlighting its promise as a powerful tool for advancing ASD prediction.

Conclusion

In summary, the combination of HTS experiments and ML techniques provides a valuable reference framework for ASD development, greatly minimizing both time and material usage in the selection of formulations during the early stages of drug discovery with a limited quantity of API.

Graphical Abstract

The workflow of this study involves a micro-quantity HTS experiment that generated over a thousand homogeneous data points for ML model development. The RF model demonstrated strong performance on a large-scale external dataset for ASD prediction, significantly saving time and material usage in formulation development, particularly during early drug discovery.