A Method for Estimating the Number of Diseases in Computed Tomography Reports of the Japanese Medical Image Database (J-MID): Variations Among Facilities
摘要
A medical image database is vital for research related to diagnosis, training, quality control, disease prevalence, and treatment outcomes. Accurate identification of disease names and their quantities within these databases is essential for efficient data utilization and gaining insights into disease types and distributions. It also plays a significant role as a resource for machine learning-based research. However, accurately determining the number of diseases in image databases has been a challenge in many healthcare facilities, hindering data management efficiency and subject selection for machine learning. This study aimed to analyze Japanese-language image diagnostic reports in the Japan Medical Image Database (J-MID) computed tomography image database and estimate disease counts. Disease name extraction was aided by a lexicon of affirmative and negative predicates. Since J-MID involves multiple medical institutions, we aimed to understand predicate usage and disease distribution variations among institutions. By using the lexicon, we identified differences in disease distribution among facilities and effectively narrowed down machine learning-suitable diseases within the dataset during the limited data acquisition period.