A multi-agent large language model framework for structured clinical interviewing and psychiatric screening: a proof-of-concept
摘要
Structured psychiatric interviews improve diagnostic reliability but are resource-intensive to administer. We developed a multi-agent large language model (LLM) framework to support protocolized structured interviewing and module-level psychiatric screening. The system orchestrates four agents, Questioner, Evaluator, Navigator, and Diagnoser, to administer interview modules, interpret free-text responses, enforce bounded clarification loops, identify safety-related red flags and trigger escalation, and apply encoded decision rules to produce screening outcomes. Prior to any model processing, responses undergo automated redaction of personally identifiable and protected health information. We conducted a simulation-based technical evaluation using 1350 interviews across 15 neuropsychiatric screening modules (e.g., anxiety, depression, and bipolar disorders) with systematic variation in symptom direction, disclosure depth, and emotional tone. To test robustness to language variability, we additionally evaluated reconstructed conversation bundles derived from the deidentified MentalChat16K dataset. In the primary evaluation, module-level decisions showed 87.8% concordance with rule-based expected outcomes (95% CI: 85.9–89.6), with sensitivity 88.9% (95% CI: 86.5–91.2) and specificity 86.7% (95% CI: 84.0–89.2) relative to predefined item-level endorsements. Post hoc comparator analyses showed higher F2 for the multi-agent framework (88.5%) than single-pass LLM (49.5%) or deterministic rule-based (61.9%) transcript-scoring baselines. On a 2000-sentence red flag benchmark stratified by explicit/implicit and active/passive ideation, the classifier achieved 93.9% accuracy, 94.2% sensitivity, and 93.6% specificity. In pilot feedback from mental health clinicians, descriptive ratings suggested favorable perceptions of clarity, workflow fit, and perceived safety for clinician-supervised use. These results support proof-of-concept technical feasibility for a protocol-adherent, auditable agentic LLM framework for structured psychiatric screening.