Emergent Topic-Based Classification of Literature Books Using Social Data Analysis
摘要
The aim of this research is to propose a method for emergent topical classification of books, based on data gathered from social media. The proposed method collects publisher provided book descriptions and data available in relevant social media, extracts features from the data and compares books considering these features. Features include keywords extracted from publisher-provided book descriptions and genres appreciated by the community of readers. The method was experimentally validated using data scraped from Goodreads – a popular subsidiary website to Amazon that provides a “social cataloging system”. The experimental dataset covers information about books tagged by users as published during the year 2022. In this contribution, we present the developed system and the experimental results showing that the proposed method allows good discrimination between book descriptions, thus supporting interested readers in choosing a book to read. Experimental results show that adding social data is beneficial, by refining the discrimination of book descriptions. The developed prototype introduces also an intuitive visualization of connections between books, based on discovered similarities.