Objective <p>To evaluate the performance of four artificial intelligence (AI) systems (ChatGPT 4o, Claude 3.7, Gemini 2.0, and Grok 2) in analysing nasal deformities.</p> Methods <p>The artificial intelligence chatbots were compared to experts in terms of their capacity to analyse nasal deformities. A quantitative analysis compared AI-generated MIRA scores with expert MIRA scores using error measures, Bland–Altman analysis, concordance metrics, and intraclass correlation coefficients to evaluate agreement and systematic bias. A qualitative evaluation was conducted using a 5-point Likert scale to characterise the major nasal type (tension nose, saddle nose, deviated nose, etc.).</p> Results <p>Fifty adult patients seeking rhinoplasty were evaluated by the chatbots and two experts based on standardised photographs. The evaluations by the two surgeons demonstrated very strong concordance (ICC = 0.997) for nasal analysis using the MIRA scale. Only Claude 3.7 and the experts had comparable total MIRA score evaluations (p &gt; 0.05). Detailed analysis of MIRA sub-scores showed a significant difference between chatbots and experts across all models (p &lt; 0.05), including Claude. Grok 2 (p &lt; 0.001) demonstrated the poorest performance. The qualitative description of the nose by ChatGPT 4o achieved the best results, with an accuracy rate reaching 70%.</p> Conclusions <p>No model achieved significant performance on MIRA sub-scores in the quantitative analysis of nasal deformity. The qualitative assessment shows that ChatGPT4o could assist, under supervision, with rhinoplasty assessments to analyse major nose types. However, it was effective in only two-thirds of cases. To date, AI tools are not reliable for analysing nasal deformities.</p> Level of Evidence V <p>This journal requires that authors assign a level of evidence to each article. For a full description of these Evidence-Based Medicine ratings, please refer to the Table of Contents or the online Instructions to Authors <a href="http://www.springer.com/00266">www.springer.com/00266</a>.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Qualitative and Quantitative Assessment of Four Artificial Intelligence Systems for Nasal Deformity Analysis in Rhinoplasty

  • Abdullah Alqahtani,
  • Dario Ebode,
  • Martin Penicaud,
  • Stephane Gargula,
  • Justin Michel,
  • Thomas Radulesco

摘要

Objective

To evaluate the performance of four artificial intelligence (AI) systems (ChatGPT 4o, Claude 3.7, Gemini 2.0, and Grok 2) in analysing nasal deformities.

Methods

The artificial intelligence chatbots were compared to experts in terms of their capacity to analyse nasal deformities. A quantitative analysis compared AI-generated MIRA scores with expert MIRA scores using error measures, Bland–Altman analysis, concordance metrics, and intraclass correlation coefficients to evaluate agreement and systematic bias. A qualitative evaluation was conducted using a 5-point Likert scale to characterise the major nasal type (tension nose, saddle nose, deviated nose, etc.).

Results

Fifty adult patients seeking rhinoplasty were evaluated by the chatbots and two experts based on standardised photographs. The evaluations by the two surgeons demonstrated very strong concordance (ICC = 0.997) for nasal analysis using the MIRA scale. Only Claude 3.7 and the experts had comparable total MIRA score evaluations (p > 0.05). Detailed analysis of MIRA sub-scores showed a significant difference between chatbots and experts across all models (p < 0.05), including Claude. Grok 2 (p < 0.001) demonstrated the poorest performance. The qualitative description of the nose by ChatGPT 4o achieved the best results, with an accuracy rate reaching 70%.

Conclusions

No model achieved significant performance on MIRA sub-scores in the quantitative analysis of nasal deformity. The qualitative assessment shows that ChatGPT4o could assist, under supervision, with rhinoplasty assessments to analyse major nose types. However, it was effective in only two-thirds of cases. To date, AI tools are not reliable for analysing nasal deformities.

Level of Evidence V

This journal requires that authors assign a level of evidence to each article. For a full description of these Evidence-Based Medicine ratings, please refer to the Table of Contents or the online Instructions to Authors www.springer.com/00266.