Generation of Space Descriptions Based on Distributional Semantic Models
摘要
The report describes a method for generating descriptions of space in virtual reality based on corpora modeling. We performed an analysis of mathematical methods and machine learning algorithms for processing large arrays of text data and creating text descriptions of objects and scenes in virtual space. A large corpora of descriptions was collected and processed. The report includes keyword extraction and named entity approach. We used pretrained static Word2Vec models to predict lexical substitutions for nouns and adjectives with a locality function in description and compared their predictions with a Word2Vec model trained on our corpora. Textual data preparation was performed by means of Latent Dirichlet Allocation. Nouns and adjectives expressing a locality function were substituted with three models learned on different datasets. The results of the experiments prove that our approach allows us to create texts that meet the requirements of well-formedness, meaningfulness, and coherence.