QSAR: A Tool of Predictive Cheminformatics
摘要
Quantitative structure-activity relationship (QSAR) and related statistical modeling techniques (such as Quantitative structure-property/toxicity relationship or QSPR/QSTR) are routinely used in data gap filling for chemical compounds for which the experimental response (activity/property/toxicity) values are unavailable. Such predictions have shown useful applications in drug design, materials science, predictive toxicology, nanoscience, food science, and agricultural science, among others. However, the quality of predictions from such models is dependent on various factors, including the quality of the modeling data set (e.g., experimental error, data set size, chemical diversity, distribution of response values, presence of influential observations, etc.) and the modeling workflow (e.g., curation and splitting of the data set, variable selection, internal and external validation, consideration of applicability domain, etc.) utilized to develop and validate the model. To develop acceptable QSAR models, the best practices should be followed in light of the guidelines recommended by the Organisation for Economic Co-operation and Development (OECD). It is also important to identify the modelability of a QSAR data set, as the quality of the final models greatly depends on several dataset-specific factors.