A reproducible symptom-based computational phenotyping framework for longitudinal fibromyalgia analysis using real-world registry data: a cross-cohort validation study
摘要
Fibromyalgia is a heterogeneous condition where clinical response to therapies like low-intensity repetitive Transcranial Magnetic Stimulation (Li-rTMS) varies significantly. Existing phenotyping is often limited by cross-sectional designs or lack of external validation. This study aims to develop and validate a reproducible informatics framework to identify stable symptom-based computational phenotypes and evaluate their longitudinal trajectories using multi-centre real-world data (RWD).
MethodsWe developed a dual-stage informatics pipeline for data curation and consistency filtering, which was applied to a registry of over 1,600 patients across three independent cohorts (Discovery, Structural Validation, and Longitudinal Validation). Unsupervised learning was performed using a K-means + + architecture on baseline clinical variables (FIQ, WPI, and SSS). Model generalizability was tested using a “frozen model” strategy, where centroids and scaling parameters from the discovery cohort were applied to the validation cohorts without re-fitting. Longitudinal stability was assessed by analysing transitions between these clinical symptom phenotypes across the treatment period.
ResultsThe framework identified three distinct and stable symptom-based computational phenotypes: CP1 (Low Symptom Burden), CP2 (Pain-Dominant), and CP3 (High Multidimensional Burden). The curation pipeline ensured high data quality, maintaining structural consistency across all cohorts. CP1 demonstrated the highest longitudinal stability, acting as a clinical “sink,” while CP2 showed the most significant clinical improvement following Li-rTMS intervention. The results confirm that the identified symptom phenotypes are not artifacts of a specific dataset but represent reproducible clinical symptom states.
ConclusionsThe proposed informatics framework provides a reliable and transferable methodology for extracting stable clinical symptom phenotypes from heterogeneous real-world registries. By demonstrating cross-cohort validation and longitudinal consistency, this work offers a scalable template for personalized clinical decision-making and stratified management in chronic pain populations, aligning with FAIR data principles.