A Balanced Counting Visual Question Answering Dataset
摘要
One of the goals of artificial intelligence is to create a machine that can answer arbitrary questions about an image. This field is known as visual question answering (VQA). A Counting VQA is a subfield that can answer counting questions about an image. The current existing benchmarks for Visual Question Answering especially for the counting problem suffer from bias issues. To solve this issue we developed a synthetic VQA dataset specialized for the counting problem. We have created a generator that automatically combines various 3D models into a checkerboard pattern. The generated image is similar to a Where-is-Waldo problem. Using this generator we created an extensive dataset to help in analysing the VQA general network architecture and the VQAv2 dataset. The main characteristic of this dataset is that it is a balanced dataset; where each object has identical or at least similar odds of showing up any number of times.