Performance of GPT-5.5 thinking for Manchester Triage System–based emergency department triage: agreement with an expert reference standard and prediction of short-term outcomes
摘要
To evaluate the agreement between Manchester Triage System (MTS) category assignments generated by GPT-5.5 Thinking and an expert-derived reference standard in a selected cohort of emergency department (ED) patients, and to assess the model’s ability to predict hospital admission and 72-hour mortality.
MethodsThis single-center observational study prospectively sampled consecutive ED presentations during a prespecified continuous 72-hour window, with triage-time variables and secondary outcomes obtained from hospital records. Patients classified as MTS Category 5 (Blue; nonurgent) by the expert reference standard were excluded from the analytic cohort. GPT-5.5 Thinking was prompted in a standardized single-shot format using structured clinical vignettes derived from triage data. The primary endpoint was agreement between model-assigned MTS categories and an expert reference standard established by two independent blinded emergency medicine specialist rater teams, with disagreements resolved through senior adjudication. Exact agreement, within-1-level agreement, overtriage, undertriage, and dangerous undertriage were calculated. Secondary analyses evaluated prediction of hospital admission and 72-hour mortality using diagnostic performance metrics.
ResultsA total of 1,841 patients were included. Overall agreement between GPT-5.5 Thinking and the expert reference standard was substantial (κ = 0.721; 95% CI 0.691–0.740). Exact agreement rate was 62.0%, and within-1-level agreement rate was 99.0%. Overtriage occurred in 35.3% of cases, undertriage in 2.7%, and dangerous undertriage in 0.2%. Exact agreement rate was significantly lower in patients aged ≥ 65 years than in younger adults (57.1% vs. 64.2%, p = 0.013). For hospital admission, accuracy rate was 79.9%, sensitivity 83.8%, positive predictive value (PPV) 45.2%, and negative predictive value (NPV) 95.9%. For 72-hour mortality, accuracy rate was 86.5%, sensitivity 72.3%, PPV 12.6%, and NPV 99.1%.
ConclusionWithin this selected ED cohort, GPT-5.5 Thinking showed substantial agreement with expert MTS-based ED triage and generally preserved the ordinal hierarchy of clinical urgency. Its performance was characterized by frequent adjacent-level agreement, predominant overtriage, and very rare but nonzero dangerous undertriage. The model also showed stronger rule-out than rule-in characteristics for hospital admission and 72-hour mortality.