Provenance for Longitudinal Analysis in Large Scale Networks
摘要
Concerns related to the veracity and originality of the content on social networks are at an ongoing rise. Considerable work has been done on information spreading, and tools have been built, while approaches with provenance-based analysis are rare. We are of the opinion that provenance-based analysis and visualization tools can make (mis-)information spreading analysis more efficient. Thus, we study provenance, and present a provenance pipeline for data analytics, where users are able to interact with multiple network analysis modules through a graphical user interface, and describe a proof-of-concept system. Although provenance visualization can suffice in capturing all the necessary metadata, integration with other network visualization modules suited to the same data enhanced our results analysis and conclusions. Having designed distinct provenance models, we captured and analysed lineage of information on community dynamics. We tested our proposed prototype with a real-world dataset comprising of more than 10 million filtered tweets, focused on COVID-19 vaccinations, and conducted an analysis on community dynamics with network science metrics and NLP.