<p>With the growing popularity of, and easy access to, the Internet, controlling the distribution of sensitive content such as adult images and videos is increasingly important. The abundance of images and videos available on the Internet makes it impossible to distinguish potentially harmful content (adult images or videos) from benign content (non-adult images or videos) through manual monitoring. To address this challenge, This paper presents a hybrid architecture combining Convolutional Neural Networks (CNN) and Transformers to address the challenge of recognizing adult images and videos on the Internet. Due to the vast amount of such content, manual monitoring is impractical. With the growing popularity of, and easy access to, the Internet, controlling the distribution of sensitive content such as adult images and videos is increasingly important. The abundance of images and videos available on the Internet makes it impossible to distinguish potentially harmful content (adult images or videos) from benign content (non-adult images or videos) through manual monitoring. To address this challenge, the proposed method integrates transformer blocks after a convolutional encoder to enhance the representation learning process, focusing on essential features. By examining and selecting the optimal trade-off between the number of layers in the transformer blocks, the proposed method selects essential features rather than minor ones. We show that our model is also effective as a powerful feature extractor for video frames in video classification. Experiments demonstrate that our method outperforms state-of-the-art methods in adult content video classification on the NPDI, Pornography-800, and Pornography-2000 datasets, with minimal computational overhead. Adding transformer blocks on top of the CNN backbone introduces only a few additional parameters to the whole architecture compared to the convolutional backbone, resulting in minimal extra computational cost. Statistical analyses, including non-parametric tests, validate the model’s performance, confirming its robustness.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A hybrid CNN-transformer architecture for adult image and video content recognition on the internet

  • Sasan Karamizadeh,
  • Mahdi Pourmirzaei,
  • Mazdak Zamani,
  • Achyut Shankar

摘要

With the growing popularity of, and easy access to, the Internet, controlling the distribution of sensitive content such as adult images and videos is increasingly important. The abundance of images and videos available on the Internet makes it impossible to distinguish potentially harmful content (adult images or videos) from benign content (non-adult images or videos) through manual monitoring. To address this challenge, This paper presents a hybrid architecture combining Convolutional Neural Networks (CNN) and Transformers to address the challenge of recognizing adult images and videos on the Internet. Due to the vast amount of such content, manual monitoring is impractical. With the growing popularity of, and easy access to, the Internet, controlling the distribution of sensitive content such as adult images and videos is increasingly important. The abundance of images and videos available on the Internet makes it impossible to distinguish potentially harmful content (adult images or videos) from benign content (non-adult images or videos) through manual monitoring. To address this challenge, the proposed method integrates transformer blocks after a convolutional encoder to enhance the representation learning process, focusing on essential features. By examining and selecting the optimal trade-off between the number of layers in the transformer blocks, the proposed method selects essential features rather than minor ones. We show that our model is also effective as a powerful feature extractor for video frames in video classification. Experiments demonstrate that our method outperforms state-of-the-art methods in adult content video classification on the NPDI, Pornography-800, and Pornography-2000 datasets, with minimal computational overhead. Adding transformer blocks on top of the CNN backbone introduces only a few additional parameters to the whole architecture compared to the convolutional backbone, resulting in minimal extra computational cost. Statistical analyses, including non-parametric tests, validate the model’s performance, confirming its robustness.