错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Structure Based Machine Learning Prediction of Retention Times for LC Method Development of Pharmaceuticals

  • Jonathan Fine,
  • Amanda K. Peterson Mann,
  • Pankaj Aggarwal

摘要

Purpose

Significant resources are spent on developing robust liquid chromatography (LC) methods with optimum conditions for all project in the pipeline. Although, data-driven computer assisted modelling has been implemented to shorten the method development timelines, these modelling approaches require project-specific screening data to model retention time (RT) as function of method parameters. Sometimes method re-development is required, leading to additional investments and redundant laboratory work. Cheminformatics techniques have been successfully used to predict the RT of metabolites & other component mixtures for similar use cases. Here we will show that these techniques can be used to model structurally diverse molecules and predictions of these models trained on multiple LC conditions can be used for downstream data-driven modelling.

Methods

The Molecular Operating Environment (MOE) was used to calculate over 800 descriptors using the strucutres of the analytes. These descriptors were used to model the RT of the analytes under four chromatographic conditions. These models were then used to create data-driven models using LC-SIM.

Results

A structural-based Random Forest (RF) model outperformed other techniques in cross-validation studies and predicted the RTs of a randomized test set with a median percentage error less than 4% for all LC conditions. RTs predicted by this structure-based model were used to fit a data-driven model that identifies optimum LC conditions without any additional experimental work.

Conclusions

These results show that small training sets yield pharmaceutically relevant models when used in a combination of structure-based and data-driven model.