<p>Early identification of clinical deterioration remains crucial for ICU patient outcomes, yet conventional prediction models rely exclusively on numerical vital signs, ignoring contextual information in clinical narratives. To develop a computationally efficient multimodal system that integrates lightweight large language model (LLM)-derived structured clinical summaries with temporal vital-sign data for simultaneous prediction of in-hospital mortality, vasopressor-dependent shock, and mechanical ventilation requirements. We analyzed 50,318 ICU stays from the MIMIC-IV v2.2 databases (raw: 73,181 stays; analytic cohort: 50,318 stays from approximately 36,800 unique patients) after applying exclusion criteria. Stays were partitioned at the patient level into training (≈ 68%), validation (≈ 14%), and test (≈ 18%) sets. Clinical notes from the first 24&#xa0;h were transformed into structured embeddings using template-guided Flan-T5-Large to minimize hallucination. These embeddings were fused with hourly vital signs and laboratory time series in a multi-task LSTM architecture. Class imbalance was addressed with focal loss for the sequential model and synthetic oversampling for tabular baselines. Primary outcomes were in-hospital mortality (11.3%), vasopressor-dependent shock (18.7%), and mechanical ventilation initiation (22.4%) within 48&#xa0;h following the 24-hour prediction window. The multimodal model achieved AUROCs of 0.876 (95% CI: 0.868–0.883) for mortality, 0.842 (95% CI: 0.835–0.849) for shock, and 0.831 (95% CI: 0.824–0.838) for ventilation improvements of 0.047, 0.038, and 0.032 over LSTM-only baselines, respectively (all <i>p</i> &lt; 0.001). Larger gains were observed in AUPRC (0.063, 0.054, and 0.048). The model showed good calibration (ECE 0.043–0.051) and identified 71.3% of patients with low SOFA scores who died, compared with near-zero sensitivity for traditional scores. Template-guided generation was associated with a lower observed hallucination rate in the reviewed sample (3.2% versus 14.7% for unconstrained generation). The system supports inference in 12–18 ms on a single 24 GB GPU. Constrained lightweight LLM summarization, when combined with temporal modeling, meaningfully enhances early prediction of mortality, shock, and ventilation requirements while remaining feasible for real-world clinical deployment.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Lightweight LLM summarization with LSTM for early ICU deterioration prediction

  • Mohammadreza Momenzadeh,
  • Atiyeh Oshaghi

摘要

Early identification of clinical deterioration remains crucial for ICU patient outcomes, yet conventional prediction models rely exclusively on numerical vital signs, ignoring contextual information in clinical narratives. To develop a computationally efficient multimodal system that integrates lightweight large language model (LLM)-derived structured clinical summaries with temporal vital-sign data for simultaneous prediction of in-hospital mortality, vasopressor-dependent shock, and mechanical ventilation requirements. We analyzed 50,318 ICU stays from the MIMIC-IV v2.2 databases (raw: 73,181 stays; analytic cohort: 50,318 stays from approximately 36,800 unique patients) after applying exclusion criteria. Stays were partitioned at the patient level into training (≈ 68%), validation (≈ 14%), and test (≈ 18%) sets. Clinical notes from the first 24 h were transformed into structured embeddings using template-guided Flan-T5-Large to minimize hallucination. These embeddings were fused with hourly vital signs and laboratory time series in a multi-task LSTM architecture. Class imbalance was addressed with focal loss for the sequential model and synthetic oversampling for tabular baselines. Primary outcomes were in-hospital mortality (11.3%), vasopressor-dependent shock (18.7%), and mechanical ventilation initiation (22.4%) within 48 h following the 24-hour prediction window. The multimodal model achieved AUROCs of 0.876 (95% CI: 0.868–0.883) for mortality, 0.842 (95% CI: 0.835–0.849) for shock, and 0.831 (95% CI: 0.824–0.838) for ventilation improvements of 0.047, 0.038, and 0.032 over LSTM-only baselines, respectively (all p < 0.001). Larger gains were observed in AUPRC (0.063, 0.054, and 0.048). The model showed good calibration (ECE 0.043–0.051) and identified 71.3% of patients with low SOFA scores who died, compared with near-zero sensitivity for traditional scores. Template-guided generation was associated with a lower observed hallucination rate in the reviewed sample (3.2% versus 14.7% for unconstrained generation). The system supports inference in 12–18 ms on a single 24 GB GPU. Constrained lightweight LLM summarization, when combined with temporal modeling, meaningfully enhances early prediction of mortality, shock, and ventilation requirements while remaining feasible for real-world clinical deployment.