WADCAT: Web Application for Data Collection and Analysis of Topics
摘要
Web Application development has become important for Data Collection and Analysis, showing the requirement for efficient solutions to manage large-scale data processing. In this paper, we present a Web Application for Data Collection and Analysis of Topics (WADCAT), specifically designed for handling textual data. This highly scalable web application addresses the challenges of data retrieval and analysis from diverse sources. By using APIs from prominent platforms such as Google, Reddit, and Twitter, enabling seamless data collection and storage. Through advanced preprocessing techniques, including comprehensive Natural Language Processing (NLP), the application ensures data cleanliness and suitability for subsequent analysis. With pre-trained models, users can perform sophisticated topic modeling and generate comprehensive reports with diverse graph types, facilitating the extraction of meaningful insights. Our evaluation demonstrates the effectiveness of WADCAT in reducing manual effort for data preprocessing and improving topic modeling accuracy. With its user-friendly interface, the application streamlines the entire data collection, preprocessing, and analysis workflow, making it effortless for users to derive valuable insights from diverse data sources. Future work can focus on developing tailored models to cater to specific dataset requirements, enhancing WADCAT’s customization capabilities, expanding its applicability in various domains, and allowing users to incorporate their customized scripts, offering a more flexible approach to Data Collection and Analysis.