Background <p>There is huge interest in the use of artificial intelligence (AI) in the production and assessment of academic material; however, the role of AI remains unclear.</p> Aim <p>The purpose of this study was to perform a reviewer-blinded assessment of the quality of scientific discussion generated by an advanced AI language model (ChatGPT-4, Open AI) and determine whether this could be recommended for high-impact journal publication.</p> Methods <p>The introduction, methods and results sections of a recently published article from a high-impact journal were input into a current AI model. The AI application then produced a discussion and conclusion based on the provided text using a standardized prompt. Six experienced blinded reviewers scored all five sections of the hybrid article. A one-way analysis of variance (ANOVA) was used to assess significant differences between scores of each section. Reviewers recommended a decision regarding the suitability of the article for publication.</p> Results <p>AI composed a scientific discussion and conclusion. The median score was 80 (IQR 70–90) for introduction, 77.5 (IQR 70–90) for methods, 82.5 (IQR 50–90) for results, 60 (IQR 40–75) for discussion and 60 (IQR 40–80) for the conclusion. The median scores for the AI-generated sections were non-significantly lower than other sections (<i>p</i> = 0.37). The majority of reviewers (5/6, 83%) recommended “acceptance for publication after major revision”. One reviewer recommended “resubmission with no guarantee of acceptance”. There were no recommendations for rejection.</p> Conclusion <p>Current AI large language models are now capable of generating content that passes experienced peer review and is acceptable for publication in a high-impact orthopaedic journal, after revision. There are still many concerns regarding the integration of AI into the process of scientific writing, mainly the tendency of AI to rely on advanced pattern recognition and fabricated or inadequate references.</p> <p><b>Level of evidence:</b> Level IV</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Can artificial intelligence generate scientific discussion that passes peer review for publication in a high-impact orthopaedic journal?

  • Gerard A. Sheridan,
  • Lisa C. Howard,
  • Michael E. Neufeld,
  • Tom R. Doyle,
  • Andrew J. Hughes,
  • Peter K. Sculco,
  • David E. Beverland,
  • Donald S. Garbuz,
  • Bassam A. Masri

摘要

Background

There is huge interest in the use of artificial intelligence (AI) in the production and assessment of academic material; however, the role of AI remains unclear.

Aim

The purpose of this study was to perform a reviewer-blinded assessment of the quality of scientific discussion generated by an advanced AI language model (ChatGPT-4, Open AI) and determine whether this could be recommended for high-impact journal publication.

Methods

The introduction, methods and results sections of a recently published article from a high-impact journal were input into a current AI model. The AI application then produced a discussion and conclusion based on the provided text using a standardized prompt. Six experienced blinded reviewers scored all five sections of the hybrid article. A one-way analysis of variance (ANOVA) was used to assess significant differences between scores of each section. Reviewers recommended a decision regarding the suitability of the article for publication.

Results

AI composed a scientific discussion and conclusion. The median score was 80 (IQR 70–90) for introduction, 77.5 (IQR 70–90) for methods, 82.5 (IQR 50–90) for results, 60 (IQR 40–75) for discussion and 60 (IQR 40–80) for the conclusion. The median scores for the AI-generated sections were non-significantly lower than other sections (p = 0.37). The majority of reviewers (5/6, 83%) recommended “acceptance for publication after major revision”. One reviewer recommended “resubmission with no guarantee of acceptance”. There were no recommendations for rejection.

Conclusion

Current AI large language models are now capable of generating content that passes experienced peer review and is acceptable for publication in a high-impact orthopaedic journal, after revision. There are still many concerns regarding the integration of AI into the process of scientific writing, mainly the tendency of AI to rely on advanced pattern recognition and fabricated or inadequate references.

Level of evidence: Level IV