Enhancing Predictive Accuracy in Molecular Solubility: A Comprehensive Study on the Impact of Datasets in Linear Regression Models
摘要
This research delves into the realm of molecular solubility prediction by leveraging linear regression models. Focusing on the application of these models to LogS (logarithmic solubility) prediction, our study systematically investigates the critical role played by diverse datasets in shaping model performance. Utilizing four distinct datasets, we explore the nuances of their impact on the predictive accuracy of linear regression models. The datasets vary in size, composition, and source, providing a comprehensive understanding of the model's adaptability to different molecular contexts. We meticulously calculate molecular descriptors (Mu features) to enrich our dataset, enabling a detailed exploration of LogS prediction. The linear regression models, trained and tested on each dataset, consistently exhibit remarkable accuracy, as evidenced by low mean squared errors and high coefficients of determination. The evaluation is complemented by visual analyses, including scatter plots, showcasing the alignment between predicted and experimental LogS values. Notably, our study unravels the model's robustness across diverse datasets while shedding light on dataset-specific influences. These insights contribute to the broader landscape of predictive modeling in biochemistry and drug discovery, emphasizing the significance of dataset characteristics in enhancing the reliability and generalizability of linear regression models for molecular solubility prediction.