<p>This study investigates the readability, clinical reliability, and temporal consistency of artificial intelligence (AI) chatbots regarding pneumothorax information. A question bank comprising 40 patient-centered queries was deployed across three large language models (ChatGPT, Gemini, Copilot), stratified by two access tiers and two prompting strategies (zero-shot versus the optimized PROMPORT strategy). Queries were replicated longitudinally on Days 1, 3, and 7 under strict session-control protocols. Text accessibility was quantified using five automated readability indices, while two independent, blinded thoracic surgeons evaluated clinical quality using modified DISCERN (mDISCERN), JAMA benchmarks, and PEMAT-P indices. Readability metrics demonstrated absolute structural stability across the tracking intervals (<i>p</i> &gt; 0.05). Unprompted configurations consistently generated complex, high-school-level outputs, whereas the PROMPORT strategy successfully compressed linguistic variances and neutralized chronological algorithmic drift (<i>p</i> &gt; 0.05). Conversely, unprompted architectures exhibited significant temporal volatility in mDISCERN and JAMA profiles (<i>p</i> &lt; 0.05), which was successfully stabilized by optimized prompt constraints. Inter-rater reliability was high across all structural evaluations. In conclusion, while unprompted models exhibit marked baseline linguistic and quality variations, the strategic integration of robust prompt engineering successfully enforces the temporal stability and clarity required for reliable digital public health communication.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluation of the performance and temporal variability of large language models in patient education regarding pneumothorax: a seven-day analysis

  • Ömer Önal,
  • Suzan Temiz Bekce

摘要

This study investigates the readability, clinical reliability, and temporal consistency of artificial intelligence (AI) chatbots regarding pneumothorax information. A question bank comprising 40 patient-centered queries was deployed across three large language models (ChatGPT, Gemini, Copilot), stratified by two access tiers and two prompting strategies (zero-shot versus the optimized PROMPORT strategy). Queries were replicated longitudinally on Days 1, 3, and 7 under strict session-control protocols. Text accessibility was quantified using five automated readability indices, while two independent, blinded thoracic surgeons evaluated clinical quality using modified DISCERN (mDISCERN), JAMA benchmarks, and PEMAT-P indices. Readability metrics demonstrated absolute structural stability across the tracking intervals (p > 0.05). Unprompted configurations consistently generated complex, high-school-level outputs, whereas the PROMPORT strategy successfully compressed linguistic variances and neutralized chronological algorithmic drift (p > 0.05). Conversely, unprompted architectures exhibited significant temporal volatility in mDISCERN and JAMA profiles (p < 0.05), which was successfully stabilized by optimized prompt constraints. Inter-rater reliability was high across all structural evaluations. In conclusion, while unprompted models exhibit marked baseline linguistic and quality variations, the strategic integration of robust prompt engineering successfully enforces the temporal stability and clarity required for reliable digital public health communication.