An Approach to Creating a Synthetic Financial Transactions Dataset Based on an NDA-Protected Dataset
摘要
This paper outlines an experiment in building an obfuscated version of a proprietary financial transactions dataset. As per industry requirements, no data from the original dataset should find its way to third parties, so all the fields, including banks, customers (including geographic locations) and particular transactions, were generated artificially. However, we set our goal to keeping as many distributions and correlations from the original dataset as possible, with adjustable levels of introduced noise in each of the fields: geography, bank-to-bank transaction flows and distributions of volumes/numbers of transactions in various subsections of the dataset. The article could be of use to anyone who may want to produce a publishable dataset, e.g. for the alternative data market, where it’s essential to keep the structure and correlations of a proprietary non-disclosed original dataset.