<p>This paper presents <Emphasis FontCategory="SansSerif">TSRF-Dist</Emphasis>, a novel distance between time series based on Random Forests (<span>RF</span>s). We extend to the time-series domain concepts and tools of <span>RF</span> distances, a recent class of robust data-dependent distances defined for vectorial representations, thus proposing the <i>first</i> <span>RF</span> distance for time series. The distance is determined by (i) creating an RF to model a set of time series, and (ii) exploiting the trained RF to quantify the similarity between time series. As for the first step, we introduce in this paper the <i>Extremely Randomized Canonical Interval Forest</i> (<span>ERCIF</span>), a novel extension of Canonical Interval Forests that can model time series and can be trained without labels. We then exploit three different schemes, following ideas already employed in the vectorial case. The proposed distance, in different variants, has been thoroughly evaluated with 128 datasets from the <Emphasis FontCategory="NonProportional">UCR Time Series</Emphasis> archive, showing promising results compared with literature alternatives.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

TSRF-Dist: a novel time series distance based on extremely randomized canonical interval forests

  • Alberto Azzari,
  • Manuele Bicego,
  • Carlo Combi,
  • Andrea Cracco,
  • Pietro Sala

摘要

This paper presents TSRF-Dist, a novel distance between time series based on Random Forests (RFs). We extend to the time-series domain concepts and tools of RF distances, a recent class of robust data-dependent distances defined for vectorial representations, thus proposing the first RF distance for time series. The distance is determined by (i) creating an RF to model a set of time series, and (ii) exploiting the trained RF to quantify the similarity between time series. As for the first step, we introduce in this paper the Extremely Randomized Canonical Interval Forest (ERCIF), a novel extension of Canonical Interval Forests that can model time series and can be trained without labels. We then exploit three different schemes, following ideas already employed in the vectorial case. The proposed distance, in different variants, has been thoroughly evaluated with 128 datasets from the UCR Time Series archive, showing promising results compared with literature alternatives.