Deep Semantic Biomedical Document Representation Method
摘要
Textual documents representation is considered as a crucial phase in many text mining tasks. It intends generally to capture the semantic information of the whole document into a representation vector, which could be further utilized in such tasks. In this paper, we have proposed a deep semantic representation method which is based on two types of features. The first is derived from the biomedical widely used Structured Semantic Resource MeSH whereas the second is generated from a deep learning phase. At this end, we chose to combine the two models Word2Vec and Convolutional Neural Network which allow bringing out the semantic relationships existing in large and complex textual documents at high abstraction levels. To evaluate this method, we involved it into a document classification process. The results of experiments, carried out on a sub-set of the OHSUMED collection, show that our method performs well.