<p>From spatial data sets, identifying groups of spatial features whose utility participation index (UPI) values are larger than a user-specified utility threshold is called high utility co-location pattern (HUCP) mining that can reveal valuable information and has been applied in many fields. Since UPI does not hold the downward closure property, almost current mining HUCP algorithms design some pruning strategies to remove unnecessary candidates in advance to improve mining performance. However, these algorithms are still powerless when dealing with large and dense data. Moreover, to avoid missing interesting HUCPs, the utility threshold should be set as small as possible. Unfortunately, this setting leads to many HUCPs generated that disturb users’ understanding and requires expensive execution resources. To address these, first, a top-down HUCP mining framework is proposed. Neighboring instances are divided into maximal cliques (MCs), and then, these MCs are organized in a compact hash table structure. The longest keys in the structure correspond to the longest candidates, and the mining process starts with these candidates and performs in a top-down search style. Second, a concise representation, <InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11227_2025_7203_Article_IEq1.gif" Format="GIF" Height="10" Rendition="HTML" Resolution="72" Type="Linedraw" Width="10" /> </InlineMediaObject> <EquationSource Format="TEX">\(\epsilon\)</EquationSource> <EquationSource Format="MATHML"><math> <mi>ϵ</mi> </math></EquationSource> </InlineEquation>-closed HUCPs, is proposed that can effectively compress similar patterns. Finally, comprehensive experiments on diverse data show that the proposed method is effective and efficient.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A common and efficient algorithm for discovering high utility co-location patterns and their concise representation from massive spatial data

  • Thiloan Bui,
  • Vanha Tran,
  • Thaigiang Do,
  • Hoangan Le,
  • Truongminh Ngo

摘要

From spatial data sets, identifying groups of spatial features whose utility participation index (UPI) values are larger than a user-specified utility threshold is called high utility co-location pattern (HUCP) mining that can reveal valuable information and has been applied in many fields. Since UPI does not hold the downward closure property, almost current mining HUCP algorithms design some pruning strategies to remove unnecessary candidates in advance to improve mining performance. However, these algorithms are still powerless when dealing with large and dense data. Moreover, to avoid missing interesting HUCPs, the utility threshold should be set as small as possible. Unfortunately, this setting leads to many HUCPs generated that disturb users’ understanding and requires expensive execution resources. To address these, first, a top-down HUCP mining framework is proposed. Neighboring instances are divided into maximal cliques (MCs), and then, these MCs are organized in a compact hash table structure. The longest keys in the structure correspond to the longest candidates, and the mining process starts with these candidates and performs in a top-down search style. Second, a concise representation, \(\epsilon\) ϵ -closed HUCPs, is proposed that can effectively compress similar patterns. Finally, comprehensive experiments on diverse data show that the proposed method is effective and efficient.