<p>The rapid expansion of chemical data, coupled with limitations of traditional computational methods, has accelerated the integration of artificial intelligence (AI) into cheminformatics. Among AI techniques, transformer-based large language models (LLMs) have emerged as powerful tools for representing and reasoning about complex molecular structures, enabling advancements in tasks ranging from molecular property prediction to de novo drug design. This review provides a comprehensive synthesis of how transformers are reshaping cheminformatics, tracing their evolution from early rule-based systems to modern generative and multimodal architectures. Through a systematic literature analysis using the PRISMA framework, we examine over 200 relevant studies to assess the current capabilities, challenges, and trends in applying LLMs to chemical problems. Key strengths include improved scalability, data efficiency, and the ability to learn global molecular patterns from sequence-based formats like SMILES and SELFIES. However, persistent challenges—such as limited interpretability, tokenization errors, and high computational costs—highlight the need for hybrid architectures, self-supervised learning, and integration with domain knowledge. Emerging trends such as 3D-aware models, quantum-enhanced systems, and agentic AI workflows signal a future where LLMs not only assist but autonomously drive discovery in chemical research. This review provides a critical exposition of the interplay between AI architectures and domain-specific challenges in molecular modeling. By bridging deep learning with domain-specific chemistry, this review offers a roadmap for researchers aiming to leverage transformers and LLMs in next-generation cheminformatics.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Large Language Models Meet Molecules: A Systematic Review of Advances and Challenges in AI-Driven Cheminformatics

  • Muhammad Saad Umer,
  • Muhammad Nabeel,
  • Usama Athar,
  • Iseult Lynch,
  • Antreas Afantitis,
  • Sami Ullah,
  • Muhammad Moazam Fraz

摘要

The rapid expansion of chemical data, coupled with limitations of traditional computational methods, has accelerated the integration of artificial intelligence (AI) into cheminformatics. Among AI techniques, transformer-based large language models (LLMs) have emerged as powerful tools for representing and reasoning about complex molecular structures, enabling advancements in tasks ranging from molecular property prediction to de novo drug design. This review provides a comprehensive synthesis of how transformers are reshaping cheminformatics, tracing their evolution from early rule-based systems to modern generative and multimodal architectures. Through a systematic literature analysis using the PRISMA framework, we examine over 200 relevant studies to assess the current capabilities, challenges, and trends in applying LLMs to chemical problems. Key strengths include improved scalability, data efficiency, and the ability to learn global molecular patterns from sequence-based formats like SMILES and SELFIES. However, persistent challenges—such as limited interpretability, tokenization errors, and high computational costs—highlight the need for hybrid architectures, self-supervised learning, and integration with domain knowledge. Emerging trends such as 3D-aware models, quantum-enhanced systems, and agentic AI workflows signal a future where LLMs not only assist but autonomously drive discovery in chemical research. This review provides a critical exposition of the interplay between AI architectures and domain-specific challenges in molecular modeling. By bridging deep learning with domain-specific chemistry, this review offers a roadmap for researchers aiming to leverage transformers and LLMs in next-generation cheminformatics.