Identifying Hidden Patterns from Health Administrative Claims by Means of “HAC2Vec” Embedding
摘要
The field of artificial intelligence (AI) has recently seen a significant role for Generative AI, particularly large language models (LLMs), and Natural Language Processing (NLP) techniques in healthcare applications. This paper explores the utility of language technologies, in deepening the understanding of Health Administrative Claims (HAC) data, a critical healthcare data source containing codes related to healthcare services. HAC data often lack essential clinical details, making it challenging to analyze disease phases, forms and subtypes. However, distinctive patterns of codes within HAC data can potentially signify specific disease phenotypes, making language technologies valuable tools for analysis. To address this, we introduce the “HAC2vec-mean” method, which utilizes skip-gram neural networks to convert HAC sequences into numerical vectors. We employ random forest models for binary and multiclass classification tasks, achieving an Area under the Receiver Operating Characteristic Curve of 0.86 for International Classification of Diseases v10. The paper presents data visualizations indicating the effectiveness of the approach in reducing data dimensionality and identifying patterns in patient profiles. Furthermore, it highlights the potential of this approach for cohort selection and index date specification. In conclusion, our study demonstrates the potential of NLP embeddings in enhancing the analysis of HAC data. This flexible framework offers improved insights into patient journeys and healthcare conditions, mitigating the limitations associated with traditional methods. Future work includes exploring the clinical relevance of identified patterns and enhancing explainability. Overall, this research opens doors to uncovering hidden structures with prognostic and therapeutic potential within HAC data.