Multiple Behavior Patterns in Ad-Related Web Traffic of Humans and Bots
摘要
There are problems in data analysis, in which we should distinguish between two fundamentally different categories (false-true, for-against, bad-good, ….), but in which we do not know the actual structure of the data set considered, and there is no appropriate knowledge of the ground truth to potentially categorize or distinguish the observations. In this particular paper, we are concerned with the data on ad-related web traffic, in which we would like to tell the genuine human-generated traffic from the artificial bot-generated one. Yet, this dichotomy, even if of primary importance from the point of view of the respective “business model”, turns out to be far too simplistic when confronted with the results of the analysis. We show how in the study process the awareness appears of a multiplicity of behavior patterns in the traffic analyzed, contributing also to the achievement of improved results. It is shown how quite simple clustering tools may provide effective identification of such patterns and their characteristics. A cognitive value added is also constituted by the very categorization of the behavior patterns identified and their relative persistence. The lessons learned may be of use for various analyses of a similar nature, where apparently simpler divisions, assumed at the outset, can only be effectively analyzed when perceived in their fuller complexity. It is also possible, given the deeper justification of the results obtained, that the concrete categorizations identified are of a wider significance than just for the case considered.