Artificial intelligence for automating the establishment of an Arabic benchmark dataset for enhancing health information quality assessment
摘要
Recent studies emphasize online health information quality. However, little research focuses on Arabic content for emergency conditions like heart disease, hypertension, and stroke despite the high demand for this information. This study aims to address this gap by using artificial intelligence to create a benchmark dataset for Arabic health information and evaluate its quality.
MethodsWe assessed the quality of health information across three criteria: source quality, treatment quality, and content trustworthiness. The Kruskal-Wallis test was used to analyze quality differences across content types (General Health Information, Medical Advice, Treatment Description) and website categories (Government, Journalistic, Portal, Professional). Data augmentation techniques such as paraphrasing, back translation, and RandAugment were also employed to enhance quality assessment using the Arabic BERT model. The study also proposes a novel architecture termed the Mixture of Classification. In this approach, each health document is processed in parallel by three instances of an Arabic BERT model: the first identifies the type of health information, the second determines the category of the provider, and the third estimates a continuous quality score via regression.
ResultsSignificant quality differences were observed among website categories and content types. Portal sites achieved the highest mean score (11.64), while Journalistic sites scored the lowest (3.33). Treatment Descriptions had the highest score (18.74), while Medical Advice scored the lowest (4.73). These differences were statistically significant with large effect sizes (Cohen’s
Content type, provider category, and quality score are key factors in enhancing the ranking of Arabic health information. Paraphrased data augmentation contributes to improved model reliability in distinguishing between quality classes. Future research should extend this approach to other languages and health-related topics. However, a major challenge remains: achieving a balanced dataset, particularly for binary classification between high- and low-quality content, as well as for the new classification based on provider category, content type, and quality score. The goal is to ensure an equal distribution across all these categories.