Making Data AI-Ready: A Comprehensive Framework for Automated Data Pipelines
摘要
Accurate and high-quality datasets are needed to deliver the desired AI/ML model outputs. However, many projects face challenges in collecting and structuring this data. The lack of data can directly impact the success of model performances. While researchers aim to simplify ML processes with less effort through automated models in the age of AI, they also have significant challenges in obtaining AI-ready data for AI/ML projects. Therefore, researchers need to optimize their data acquisition, collection, storage, and quality check processes; develop sound data strategies based on them; and make this data ready for AI/ML models. In our study, we aim to develop a framework, which can serve as a preliminary step in automating processing workflows for AI/ML models and contribute to improving outcomes. For this purpose, we first conduct a review of related studies to establish a basis for our framework by highlighting recent developments related to AI-ready data. The studies in the relevant literature are discussed considering the phases of data preparation such as data understanding, infrastructure, storage, collection, preparation, and quality. Based on the literature review, key findings and assessments are combined within the proposed framework.