RealDTT: Towards A Comprehensive Real-World Dataset for Tampered Text Detection
摘要
The swift advancement of text manipulation in AI-generated images and the rise of false document fabrication emphasize the need for effective detection methods applicable in real-world settings. While current forensics research primarily addresses tampered text in natural images, text manipulation in documents presents a more realistic struggle to handle. To address the robustness of current detection methods and datasets, we aim to develop a real-world, large-scale dataset containing manually tampered documents and diverse automatic tampering techniques. Our work distinguishes itself from existing benchmarks through three key features: Manual Tampering: encompassing the simulation of realism and cognition, where human edits are often subtle and contextually coherent. Diverse Generators: rich manipulating types for tampered images ensure the coverage of traditional and advanced tampering techniques. Multilingual and Multiscene Coverage: spanning English and Chinese text across natural scenes and documents, with varied resolutions. We have developed a comprehensive dataset, RealDTT, to evaluate the open-set generalization capabilities of text-tampered detection models. The RealDTT encompasses approximately 300,000 diverse synthetic samples originating from nine distinct generative models. To our knowledge, this represents the most extensive collection of Deepfake model types currently available. Complementing these synthetic samples are 4,012 meticulously manually tampered images. Moreover, leveraging the RealDTT dataset, we propose a robust tampered text detection model, TTDMamba, which fully harnesses the unique strengths of the Mamba architecture and integrates selective scanning, high-frequency feature aggregation, and disentangled semantic axial attention to process global information while maintaining linear complexity. Extensive experiments demonstrate that the proposed TTDMamba exhibits remarkable efficacy.