Social Network Data
摘要
For a community detection algorithm to be useful, it must work on large real-world data. An important application domain is social networks, where communities represent users with common interests, and identifying communities has potential applications in recommending new connections and in marketing to focused audiences. We use a crawler to extract the user profile data from a particular social network, Instagram. The extraction process is challenging because of the scale required for a reasonable sample; because the sample should be reasonably random, so the graph nucleated by a single, or even a few, nodes does not serve; and because even the simplest analysis of the extracted sample requires computational effort. We show that the Instagram graph is unusual for a real-world social network because it is a mixture of knowledge-dissemination network and a more conventional social network.