<p>Publications related to experimentation with Large Language Models (LLMs) in healthcare are rapidly increasing. While human evaluation remains the gold standard for evaluating LLMs, there is still a lack of standardization in its implementation. In this review article, we systematically examine studies involving LLMs in healthcare that have conducted human evaluations. We analyze the metrics used, assess their variability across studies. We also propose a standardized framework along with an interactive open web application <i>HumanELY</i>, to facilitate human evaluation. We believe that use of <i>HumanELY</i> will provide an opportunity for consistent, comprehensive, reliable, reproducible, and measurable human evaluations of LLM in healthcare. HumanELY is publicly available at <a href="https://www.brainxai.com/humanely">https://www.brainxai.com/humanely</a>.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Human evaluation of large language models in healthcare: gaps, challenges, and the need for standardization

  • Raghav Awasthi,
  • Atharva Bhattad,
  • Sai Prasad Ramachandran,
  • Shreya Mishra,
  • Ashish K. Khanna,
  • Jacek B. Cywinski,
  • Kamal Maheshwari,
  • Dwarikanath Mahapatra,
  • Izabella DiRosa,
  • Anabelle Cohen,
  • Hajra Arshad,
  • Aarit Atreja,
  • Asma Alshukaili,
  • Aryan Vohra,
  • Nishant Singh,
  • Francis A. Papay,
  • Ashish Atreja,
  • Rahul Kashyap,
  • Piyush Mathur

摘要

Publications related to experimentation with Large Language Models (LLMs) in healthcare are rapidly increasing. While human evaluation remains the gold standard for evaluating LLMs, there is still a lack of standardization in its implementation. In this review article, we systematically examine studies involving LLMs in healthcare that have conducted human evaluations. We analyze the metrics used, assess their variability across studies. We also propose a standardized framework along with an interactive open web application HumanELY, to facilitate human evaluation. We believe that use of HumanELY will provide an opportunity for consistent, comprehensive, reliable, reproducible, and measurable human evaluations of LLM in healthcare. HumanELY is publicly available at https://www.brainxai.com/humanely.