<p>Estimation of enzymatic activities still heavily relies on experimental assays, which can be cost and time-intensive. We present CatPred, a deep learning framework for predicting in vitro enzyme kinetic parameters, including turnover numbers (<i>k</i><sub><i>cat</i></sub>), Michaelis constants (<i>K</i><sub><i>m</i></sub>), and inhibition constants (<i>K</i><sub><i>i</i></sub>). CatPred addresses key challenges such as the lack of standardized datasets, performance evaluation on enzyme sequences that are dissimilar to those used during training, and model uncertainty quantification. We explore diverse learning architectures and feature representations, including pretrained protein language models and three-dimensional structural features, to enable robust predictions. CatPred provides accurate predictions with query-specific uncertainty estimates, with lower predicted variances correlating with higher accuracy. Pretrained protein language model features particularly enhance performance on out-of-distribution samples. <i>CatPred</i> also introduces benchmark datasets with extensive coverage (~23 k, 41 k, and 12 k data points for <i>k</i><sub><i>cat</i></sub>, <i>K</i><sub><i>m</i></sub>, and <i>K</i><sub><i>i</i></sub> respectively). Our framework performs competitively with existing methods while offering reliable uncertainty quantification.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

CatPred: a comprehensive framework for deep learning in vitro enzyme kinetic parameters

  • Veda Sheersh Boorla,
  • Costas D. Maranas

摘要

Estimation of enzymatic activities still heavily relies on experimental assays, which can be cost and time-intensive. We present CatPred, a deep learning framework for predicting in vitro enzyme kinetic parameters, including turnover numbers (kcat), Michaelis constants (Km), and inhibition constants (Ki). CatPred addresses key challenges such as the lack of standardized datasets, performance evaluation on enzyme sequences that are dissimilar to those used during training, and model uncertainty quantification. We explore diverse learning architectures and feature representations, including pretrained protein language models and three-dimensional structural features, to enable robust predictions. CatPred provides accurate predictions with query-specific uncertainty estimates, with lower predicted variances correlating with higher accuracy. Pretrained protein language model features particularly enhance performance on out-of-distribution samples. CatPred also introduces benchmark datasets with extensive coverage (~23 k, 41 k, and 12 k data points for kcat, Km, and Ki respectively). Our framework performs competitively with existing methods while offering reliable uncertainty quantification.