<p>Enzymes drive cellular metabolism, yet predicting catalytic properties from amino acid sequences remains challenging. Existing protein language models (PLMs) provide powerful general-purpose representations but are often inefficient for high-throughput screening and insufficiently adapted to enzyme-specific tasks. Here, we propose EnzGFM, an enzyme-specific PLM based on a Mamba-Transformer hybrid architecture with hierarchical pre-training to capture enzyme-specific patterns. Across enzyme property prediction benchmarks, EnzGFM consistently outperforms Transformer-based PLMs with 2–5-fold acceleration, achieving relative improvements of 16.67% in kinetic parameter prediction, 15.69% in enzyme-reaction mapping, 13.19% in EC number classification, and 20.04% in mutation effect assessment. Building on EnzGFM, we develop EnzGFM-Agent, an enzyme-focused agentic pipeline. Experimental validation further suggests that EnzGFM-Agent can enrich beneficial variants within small candidate pools. Together, these results demonstrate that EnzGFM captures enzyme-specific sequence-function patterns, while EnzGFM-Agent translates these predictions into experimentally actionable candidates and can help reduce wet-lab screening burden for practical enzyme engineering.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An enzyme-specific protein language model for catalytic property prediction

  • Chong Wang,
  • Mengyao Li,
  • Shaolei Geng,
  • Weidong Li,
  • Xuezhi Zhou,
  • Yu Guang Wang,
  • Yi Yu,
  • Tianyun Wang,
  • Yiqing Shen

摘要

Enzymes drive cellular metabolism, yet predicting catalytic properties from amino acid sequences remains challenging. Existing protein language models (PLMs) provide powerful general-purpose representations but are often inefficient for high-throughput screening and insufficiently adapted to enzyme-specific tasks. Here, we propose EnzGFM, an enzyme-specific PLM based on a Mamba-Transformer hybrid architecture with hierarchical pre-training to capture enzyme-specific patterns. Across enzyme property prediction benchmarks, EnzGFM consistently outperforms Transformer-based PLMs with 2–5-fold acceleration, achieving relative improvements of 16.67% in kinetic parameter prediction, 15.69% in enzyme-reaction mapping, 13.19% in EC number classification, and 20.04% in mutation effect assessment. Building on EnzGFM, we develop EnzGFM-Agent, an enzyme-focused agentic pipeline. Experimental validation further suggests that EnzGFM-Agent can enrich beneficial variants within small candidate pools. Together, these results demonstrate that EnzGFM captures enzyme-specific sequence-function patterns, while EnzGFM-Agent translates these predictions into experimentally actionable candidates and can help reduce wet-lab screening burden for practical enzyme engineering.