The needs in terms of interpretability and reliability of the results provided by artificial intelligence systems are growing. We address these themes in the specific context of categorical data encoding. In this article, we propose a more reliable technique for categorical data encoding. Our approach is based on information about the probability distributions of the variables across a tabular data. This information is represented in a vectorial manner and the norm of this vector corresponds to the encoding value. Experimental results show that our technique allows to better express the distances or similarities between tabular data. Indeed, in three clustering problems with different datasets, the results of the Kmeans model were considerably closer to reality with our encoding technique than with one-hot encoding. On the other hand, with a KNN model on four classification problems, all binary except one, the performances were slightly in favor of the one-hot technique. But from the point of view of interpretability and reliability, our technique offers a richer base.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Categorical Data Encoding Technique Using Probability Distributions of Variables Across a Tabular Dataset

  • Adama Samaké,
  • Ismaël Koné,
  • Lahsen Boulmane

摘要

The needs in terms of interpretability and reliability of the results provided by artificial intelligence systems are growing. We address these themes in the specific context of categorical data encoding. In this article, we propose a more reliable technique for categorical data encoding. Our approach is based on information about the probability distributions of the variables across a tabular data. This information is represented in a vectorial manner and the norm of this vector corresponds to the encoding value. Experimental results show that our technique allows to better express the distances or similarities between tabular data. Indeed, in three clustering problems with different datasets, the results of the Kmeans model were considerably closer to reality with our encoding technique than with one-hot encoding. On the other hand, with a KNN model on four classification problems, all binary except one, the performances were slightly in favor of the one-hot technique. But from the point of view of interpretability and reliability, our technique offers a richer base.