Discovering a set of spatial features from spatial datasets whose instances frequently appear together in the neighbor area of each other called prevalent co-location pattern (PCP) mining is an important technique in data mining. Currently, most mining algorithms use a single minimum prevalence threshold to filter PCPs. However, the number of instances of each feature in datasets varies greatly, and each feature’s participation ratio in a pattern is also very different. When a single unified threshold is used, valuable patterns will likely be filtered out. To overcome this shortcoming, this paper proposes a method that uses multiple minimum prevalence thresholds, in which each feature is assigned a suitable threshold according to its number of instances and the application of users. The prevalence threshold of each pattern is dynamic according to the different features participating in the pattern. To mine multiple minimum prevalence threshold-based PCPs efficiently, this paper also designs an algorithm based on a participating instance hash table. First, the dataset is scanned once to enumerate all maximal cliques that are sets of neighboring instances, and then these cliques are organized into a hash structure. Next, the mining process uses these keys of the hash structure as the initial candidate set. The participating instances of each candidate are collected by directly executing a query in the hash structure. When the prevalence of this candidate is judged, its direct subsets are generated as new candidates and put into the candidate set. The proposed method is experimentally tested on synthetic and real datasets in various aspects. The experimental results show that the proposed method is effective and efficient when compared with the state-of-the-art algorithms.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Mining Prevalent Co-location Patterns with Multiple Minimum Prevalence Thresholds

  • Vanha Tran,
  • Thiloan Bui,
  • Thaigiang Do,
  • Hoangan Le

摘要

Discovering a set of spatial features from spatial datasets whose instances frequently appear together in the neighbor area of each other called prevalent co-location pattern (PCP) mining is an important technique in data mining. Currently, most mining algorithms use a single minimum prevalence threshold to filter PCPs. However, the number of instances of each feature in datasets varies greatly, and each feature’s participation ratio in a pattern is also very different. When a single unified threshold is used, valuable patterns will likely be filtered out. To overcome this shortcoming, this paper proposes a method that uses multiple minimum prevalence thresholds, in which each feature is assigned a suitable threshold according to its number of instances and the application of users. The prevalence threshold of each pattern is dynamic according to the different features participating in the pattern. To mine multiple minimum prevalence threshold-based PCPs efficiently, this paper also designs an algorithm based on a participating instance hash table. First, the dataset is scanned once to enumerate all maximal cliques that are sets of neighboring instances, and then these cliques are organized into a hash structure. Next, the mining process uses these keys of the hash structure as the initial candidate set. The participating instances of each candidate are collected by directly executing a query in the hash structure. When the prevalence of this candidate is judged, its direct subsets are generated as new candidates and put into the candidate set. The proposed method is experimentally tested on synthetic and real datasets in various aspects. The experimental results show that the proposed method is effective and efficient when compared with the state-of-the-art algorithms.