Use of Riemannian Distance Metric to Verify Topological Similarity of Acoustic and Text Domains
摘要
In the field of speech recognition and natural text processing, it is assumed that the acoustic and textual domains for the same language are similar. In this work, we show this hypothesis is true from the quantitative perspective. For our results to be more pronounced, we applied an unsuper-vised approach to unlabeled data in order to learn its topological structure. In particular, we used variational autoencoders and topological data analysis techniques as follows. First, generative methods based on variational autoencoders were chosen to map datasets into two latent vector spaces. Then, persistent homology methods are used to analyze the topological structure of two spaces. The result of this analysis is the constructed persistence diagrams. Finally, by representing persistence diagrams as persistent images and transferring them to the Hilbert unit sphere, it has been quantitatively shown their similarity using the Riemannian distance metric.