BI-RADS Classified Mammography Dataset for AI-Based Diagnostic Models
摘要
The use of Artificial Intelligence (AI)-based techniques to support the diagnostic interpretation of breast pathologies has become a valuable tool. These techniques show promising potential for the early detection of breast cancer through mammography. A key component in the development of these systems is the training process, particularly in the context of supervised learning. In medical image processing ap-plications, the availability of labeled datasets—where images are paired with confirmed diagnoses or clinical findings—is essential for training, validating, and evaluating supervised learning models. However, access to such databases remains limited, largely due to the need for large, high-quality datasets that are accurately labeled and statistically representative of the target population. Given the constraints and limitations of existing publicly available mammographic databases, this study presents a research protocol aimed at constructing a dedicated dataset through the systematic collection of mammographic images and their corresponding clinical diagnoses, including BI-RADS classifications. The work is structured in two main phases: the initial phase involves the design and preparation of the research protocol and all necessary documentation required for ethical and institutional approval. The second phase outlines the procedural steps taken to build the database. The resulting dataset is substantial, comprising approximately 21,000 individually classified images from 5100 distinct studies. This volume of data exceeds that of most existing public databases, making it a valuable resource for the training and evaluation of AI models in breast cancer diagnosis.