Digital document images play a critical role in daily life. With the great advances in image editing techniques, document image tampering localization is becoming increasingly important. Most existing methods for document image tampering localization heavily rely on tampered image synthesis and pixel-level annotations for training. However, both acquiring tampered document images and labeling the tampered regions in real-world scenarios are expensive, which motivates us to solve this task by unsupervised learning using authentic images only. In this work, we propose a feature reconstruction network to learn global comprehension of authentic images, and then identify anomalies with high reconstruction errors as tampered pixels. Particularly, we integrate visual features and frequency domain compression artifacts to expose tampering traces from multi-views. We design a random rectangular mask (RRM) strategy that leverages the prior knowledge of text shapes to prevent over-reconstruction of tampered regions. Evaluation on the benchmark dataset DocTamper-FCD/SCD demonstrates that our approach dramatically outperforms other unsupervised baselines, and shows the effectiveness of our method.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Unsupervised Document Image Tampering Localization via Anomaly Detection

  • Yuan Li,
  • Yan-Ming Zhang,
  • Fei Yin,
  • Lin-Lin Huang

摘要

Digital document images play a critical role in daily life. With the great advances in image editing techniques, document image tampering localization is becoming increasingly important. Most existing methods for document image tampering localization heavily rely on tampered image synthesis and pixel-level annotations for training. However, both acquiring tampered document images and labeling the tampered regions in real-world scenarios are expensive, which motivates us to solve this task by unsupervised learning using authentic images only. In this work, we propose a feature reconstruction network to learn global comprehension of authentic images, and then identify anomalies with high reconstruction errors as tampered pixels. Particularly, we integrate visual features and frequency domain compression artifacts to expose tampering traces from multi-views. We design a random rectangular mask (RRM) strategy that leverages the prior knowledge of text shapes to prevent over-reconstruction of tampered regions. Evaluation on the benchmark dataset DocTamper-FCD/SCD demonstrates that our approach dramatically outperforms other unsupervised baselines, and shows the effectiveness of our method.