This study explores the potential of large language models (LLMs) as independent research generators, leveraging a dataset of over 1.2 million DBLP papers (2019–2023) across diverse domains. Utilizing cutting-edge LLMs, including Llama-3, Mistral, Mixtral, and Gemma, we subjected them to supervised fine-tuning and direct preference optimization (DPO) using an automated preference dataset. Our experiments reveal that DPO-optimized models surpass solely supervised fine-tuned models like GPT-3.5 Turbo, Davinci-002, and Gemini-1.0 by 27% in the novel creativity index, which evaluates originality, feasibility, impact, and reliability. Additionally, these models achieved a 42% improvement in automated user satisfaction scores, with 89% of the generated research ideas being validated as highly relevant and promising by domain experts. This research demonstrates the significant potential of LLMs as autonomous researchers, setting a new standard for efficiency and creativity in ideation.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Empowering AI as Autonomous Researchers: Evaluating LLMs in Generating Novel Research Ideas Through Automated Metrics

  • Debajyoti Dasgupta,
  • Arijit Mondal,
  • Partha P. Chakrabarti

摘要

This study explores the potential of large language models (LLMs) as independent research generators, leveraging a dataset of over 1.2 million DBLP papers (2019–2023) across diverse domains. Utilizing cutting-edge LLMs, including Llama-3, Mistral, Mixtral, and Gemma, we subjected them to supervised fine-tuning and direct preference optimization (DPO) using an automated preference dataset. Our experiments reveal that DPO-optimized models surpass solely supervised fine-tuned models like GPT-3.5 Turbo, Davinci-002, and Gemini-1.0 by 27% in the novel creativity index, which evaluates originality, feasibility, impact, and reliability. Additionally, these models achieved a 42% improvement in automated user satisfaction scores, with 89% of the generated research ideas being validated as highly relevant and promising by domain experts. This research demonstrates the significant potential of LLMs as autonomous researchers, setting a new standard for efficiency and creativity in ideation.