Quantitative structure–activity relationship (QSAR) models have been widely and successfully used in many research areas providing insights into the factors that influence molecular properties and helping in the design of new compounds with improved characteristics. Such models are typically derived using machine learning (ML) algorithms. In this chapter we therefore provide an overview of the main steps involved in the derivation of QSAR models including data collection, training/test splitting, data preparation, outlier removal, model derivation, model validation, model interpretation, and model application. We then provide an in-depth description of many of the ML algorithms used for this purpose such as linear regression, logistic regression, support vector machines, neural networks, and decision trees. Following the description of each algorithm, we provide relevant examples taken from the literature. We complete this chapter by discussing the various metrics used for the evaluation of both classification and regression models. We hope that this chapter will serve as an entry point for newcomers into the field of ML as practiced in the derivation of QSAR models as well as a valuable reference for more experienced practitioners of the field.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Statistical Methods in QSAR

  • Paul F. A. Clarke,
  • Hanoch Senderowitz

摘要

Quantitative structure–activity relationship (QSAR) models have been widely and successfully used in many research areas providing insights into the factors that influence molecular properties and helping in the design of new compounds with improved characteristics. Such models are typically derived using machine learning (ML) algorithms. In this chapter we therefore provide an overview of the main steps involved in the derivation of QSAR models including data collection, training/test splitting, data preparation, outlier removal, model derivation, model validation, model interpretation, and model application. We then provide an in-depth description of many of the ML algorithms used for this purpose such as linear regression, logistic regression, support vector machines, neural networks, and decision trees. Following the description of each algorithm, we provide relevant examples taken from the literature. We complete this chapter by discussing the various metrics used for the evaluation of both classification and regression models. We hope that this chapter will serve as an entry point for newcomers into the field of ML as practiced in the derivation of QSAR models as well as a valuable reference for more experienced practitioners of the field.