Entity Labeling and Data Analysis Framework
摘要
Training and test data are required in the field of machine learning-based text mining research in order to train machine learning models and assess prediction outcomes. Application domain specialists manually label the train and test data. In text data, labelled data refers to people, locations, products, dates/times, and concepts. Finding an entity or notion in the text and categorizing it are the two steps in labelling. There are no software tools available to facilitate this task, hence the process requires a long time. The paper aims to provide a framework that facilitates entity/concept labelling and classification in text, generating training and assessment data for text data mining research. To assist research teams in analyzing and suggesting avenues for improvement, the tool also supports the function of evaluating prediction results on the difference between machine learning outcomes and evaluation data.