CHIC: Corporate Document for Visual Question Answering
摘要
The massive use of digital documents due to the substantial trend of paperless initiatives confronted some companies with finding ways to process thousands of documents per day automatically. To achieve this, they use automatic information retrieval (IR) allowing them to extract useful information from large datasets quickly. In order to have effective IR methods, it is first necessary to have an adequate dataset. Although companies have enough data to take into account their needs, there is also a need for a public database to compare contributions between state-of-the-art methods. Some public document datasets already exist like DocVQA and XFUND, but they do not fully satisfy the needs of companies. First, XFUND contains only form documents while the company uses several types of documents (i.e. structured documents like forms but also semi-structured as invoices, and unstructured as emails). DocVQA, for its part, has several types of documents but only 4.5% of them are business documents (i.e. invoice, purchase order, etc.). All of these 4.5% of documents do not meet the diversity of documents that companies may encounter in their daily document flow. In order to extend these limitations, we propose in this paper the CHIC dataset, a visual question-answering public dataset that contains different types of business documents and the information extracted from these documents meets the expectations of companies.