<p>The integration of retracted scientific literature into large language models (LLMs) poses significant risks to research integrity, particularly in fast-moving fields like stem cell research. To develop and apply a multi-method framework for detecting the presence of retracted stem cell articles in ChatGPT-5.0’s training data and analyze citation patterns. We identified 117 retracted stem cell articles from PubMed (1998–2025) using title keywords and collected 30 non-retracted control articles matched by publication year and journal impact factor. We employed two detection methods: (1) Question-answering tests based on article conclusions; (2) Text autocompletion tasks measuring similarity between ChatGPT-generated completions and original abstracts. Statistical analyses included chi-square tests, correlation analysis, and multivariate regression. The text autocompletion method showed high similarity (ROUGE-L score &gt; 0.7) for 67.5% of retracted articles, strongly indicating their presence in training data. In QA testing, ChatGPT-5.0 utilized 61.5% (72/117) of retracted articles, with only 41.7% (30/72) of these responses including retraction notices. Retraction time interval significantly influenced citation rates (χ<sup>2</sup> = 16.28, <i>p</i> &lt; 0.001), with highest rates for articles retracted within 24&#xa0;months (86.7%). Controlled analysis revealed that Asian-authored retracted articles were cited 28.1% more frequently than controls (<i>p</i> = 0.028), while European/American articles showed no significant difference. Multiple detection methods confirm that retracted stem cell literature is incorporated into ChatGPT-5.0’s training data. The model demonstrates inconsistent retraction recognition and exhibits geographical disparities in citation patterns. These findings highlight the urgent need for improved filtering mechanisms in LLM training pipelines and dynamic retraction verification systems.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Retracted articles on stem cells are continuously cited in publications and used by ChatGPT

  • Xinhe Zhang,
  • Tianshu Gu,
  • Sidharth Loganathan,
  • Yanjun Xie,
  • Jinghong Chen

摘要

The integration of retracted scientific literature into large language models (LLMs) poses significant risks to research integrity, particularly in fast-moving fields like stem cell research. To develop and apply a multi-method framework for detecting the presence of retracted stem cell articles in ChatGPT-5.0’s training data and analyze citation patterns. We identified 117 retracted stem cell articles from PubMed (1998–2025) using title keywords and collected 30 non-retracted control articles matched by publication year and journal impact factor. We employed two detection methods: (1) Question-answering tests based on article conclusions; (2) Text autocompletion tasks measuring similarity between ChatGPT-generated completions and original abstracts. Statistical analyses included chi-square tests, correlation analysis, and multivariate regression. The text autocompletion method showed high similarity (ROUGE-L score > 0.7) for 67.5% of retracted articles, strongly indicating their presence in training data. In QA testing, ChatGPT-5.0 utilized 61.5% (72/117) of retracted articles, with only 41.7% (30/72) of these responses including retraction notices. Retraction time interval significantly influenced citation rates (χ2 = 16.28, p < 0.001), with highest rates for articles retracted within 24 months (86.7%). Controlled analysis revealed that Asian-authored retracted articles were cited 28.1% more frequently than controls (p = 0.028), while European/American articles showed no significant difference. Multiple detection methods confirm that retracted stem cell literature is incorporated into ChatGPT-5.0’s training data. The model demonstrates inconsistent retraction recognition and exhibits geographical disparities in citation patterns. These findings highlight the urgent need for improved filtering mechanisms in LLM training pipelines and dynamic retraction verification systems.