错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Preserving the Usefulness of Data During the Depersonalisation Process: Problem Statement and Experiment Plan

  • Ilia Trofimychev

摘要

In recent years the amount of data produced by companies and services have grown tremendously. This is a great raw material for such fields of study as data mining, machine learning and artificial intelligence. Nevertheless, before transferring the data for further examination it should be verified that this information is allowed for processing by third-party organizations. One of the conditions that must be taken into account in this situation is that the data should not contain records directly or indirectly pointing to certain people - that is, it should not contain personal data. In case it does, the data should go through the depersonalisation procedure and here is another issue: by altering the data too much during the anonymisation process, the connections within it are disturbed, making the data useless from the scientific point of view. This work aims to develop a mathematical model that evaluates how well data quality is maintained during depersonalisation by addressing following tasks: (1) Setting a mathematical problem; (2) Providing a description of attribute types that can be found in personal data; (3) Describing the apparatus for using metrics and logical groups to assess the similarity of personal data subjects; (4) Featuring difficulties that could arise during depersonalisation process; (5) Proposing a design for an audit of depersonalisation algorithm and a plan for the future experiment.