Connection and Curation of Corpus (Labeled and Unlabeled)
摘要
The availability of large amount of unlabeled and unstructured real-time textual data poses a challenge for training models in various tasks. Labeled data is essential for both training and testing the machine learning models, but obtaining labeled data can be difficult and time-consuming. This limitation prevents us from harnessing the valuable information hidden within this medical textual data. However, recent advancements in biomedical databases have prompted the development of techniques to extract and annotate this data, making it structured and usable. Data curation plays a vital role in transforming unstructured data into annotated and structured data, that can be utilized for training and evaluating models effectively. Curation involves processes such as data preprocessing, feature extraction, annotation, and quality control. These steps aim to transform raw data into a well-organized and labeled dataset that can serve as a valuable resource for various biomedical applications. This chapter elucidates the need of labeled data and importance of integrating different types of data accompanied by the explanation of data curation process and case study.