Big Data Analytics Approach with Multiple Text Types: The Case of the Computer Gaming
摘要
This study explores the possibilities of processing multi-textual data (objects represented by a set of texts and metadata about them). During the research, a dataset of such data was collected, containing information about 13,117 video games. Each game is represented by 2 different types of texts, the quantity of which is not known in advance. When building models, the authors were guided by the following assumptions: each individual text is meaningful and complete (therefore separate processing of texts contributes to improving predictions), texts of different types are significantly different (therefore different types of texts should be processed separately), and inclusion of metadata increases the amount of information about the object (therefore contributes to improving predictions). Within the study, 2 models using classical NLP methods and 3 models using the above assumptions were created. As a result of 2 series of experiments, the proposed models showed higher results compared to classical methods.