<p>A key task in mining tree-structured data is finding frequent embedded tree patterns, which has two settings: the <i>transactional</i> setting and the <i>per-occurrence</i> setting. In the <i>transactional</i> setting, which is the focus of this paper, the crucial step is to decide whether a tree pattern is subtree homeomorphic to a database tree. Our extensive study on the properties of real-world tree-structured datasets reveals that while many vertices in a database tree may have the same label, no two vertices on the same path are identically labeled. In this paper, we exploit this property and propose a novel and efficient method for deciding whether a tree pattern is subtree homeomorphic to a database tree. Our algorithm is based on a compact data structure called <b>EMET</b>, which stores all information required for subtree homeomorphism. We propose an efficient algorithm to generate <b>EMET</b>s of larger patterns using <b>EMET</b>s of the smaller ones. Based on the proposed subtree homeomorphism method, we introduce <Emphasis FontCategory="SansSerif">TTM</Emphasis>, an effective algorithm for finding frequent tree patterns from rooted ordered trees. We evaluate the efficiency of <Emphasis FontCategory="SansSerif">TTM</Emphasis> on several real-world and synthetic datasets and show that it outperforms well-known existing algorithms by an order of magnitude.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Mining transactional tree databases under homeomorphism

  • Mostafa Haghir Chehreghani,
  • Morteza Haghir Chehreghani

摘要

A key task in mining tree-structured data is finding frequent embedded tree patterns, which has two settings: the transactional setting and the per-occurrence setting. In the transactional setting, which is the focus of this paper, the crucial step is to decide whether a tree pattern is subtree homeomorphic to a database tree. Our extensive study on the properties of real-world tree-structured datasets reveals that while many vertices in a database tree may have the same label, no two vertices on the same path are identically labeled. In this paper, we exploit this property and propose a novel and efficient method for deciding whether a tree pattern is subtree homeomorphic to a database tree. Our algorithm is based on a compact data structure called EMET, which stores all information required for subtree homeomorphism. We propose an efficient algorithm to generate EMETs of larger patterns using EMETs of the smaller ones. Based on the proposed subtree homeomorphism method, we introduce TTM, an effective algorithm for finding frequent tree patterns from rooted ordered trees. We evaluate the efficiency of TTM on several real-world and synthetic datasets and show that it outperforms well-known existing algorithms by an order of magnitude.