Extracting White-Box Knowledge from Word Embedding: Modeling as an Optimization Problem
摘要
Explainability is crucial to building the confidence of the medical team to adopt natural language processing (NLP) techniques. In the majority of recent studies in medical informatics, Deep Learning performed better than other machine learning (ML) techniques for natural language processing (NLP) on medical documents. However, the generated models are black-box models difficult to explain. One of these models is word embedding which allows a representation of text and words as vectors, which makes them more exploitable by machines. This paper proposes a new method to add explainability to word embedding. We propose a modelization as an optimization problem. The first results on the text8 dataset and 5 target words show the local search can obtain explanations with an improvement of cosine similarity by 11% to 30%.