Ookpik- A Collection of Out-of-Context Image-Caption Pairs
摘要
The development of AI-based Cheapfakes detection models has been hindered by a significant challenge - the scarcity of real-world datasets. Our work directly tackles this issue by focusing on out-of-context (OOC) image-caption pairs within the Cheapfakes landscape. Here, genuine images come with misleading captions, making it tough to identify them accurately. This study introduces a dataset manually collected for OOC detection. Previous attempts at solving this problem often used strategies like automatically making OOC data or providing datasets made by people. However, these approaches had limitations affecting how practical they were. Our contribution stands out by providing a carefully selected dataset to fill this gap. We tested our dataset using existing OOC detection models, showing it works well. Additionally, we suggest practical ways to use our work to improve Cheapfakes detection in real-world situations.