<p>In recent years, Explainable Artificial Intelligence (XAI) has received increasing attention in the field of computer vision, particularly for its ability to generate interpretable and semantically rich image captions. However, many existing methods fail to effectively capture the contextual relevance between the primary object and its surrounding environment, leading to generic or misleading textual outputs. To address these limitations, this paper presents a novel and efficient framework that integrates the Exponential Parametric Rectified Linear Unit Convolutional Saliency Gradients Neural Network with MaxAbs Long Short-Term Co-Attention Memory (EPReLU-CSGNN-MALSTCAM) and the Sculptor Schwefel Optimization Algorithm (SSOA) to generate high-quality textual explanations from animal images. The methodology involves comprehensive pre-processing techniques such as contrast enhancement using QCCLAHE, which achieves a PSNR of 36.59&#xa0;dB and significantly reduces image distortion. Salient features are extracted using a saliency map that emphasizes informative regions based on object-background interactions. For object detection, the proposed AGSPP-YOLO attains an Average Precision (AP) of 98.78% and a Mean AP of 97.84%, outperforming conventional detectors like YOLO and R-CNN. The feature selection phase, optimized through SSOA, yields a fitness value of 97.24 and a selection time of 2154&#xa0;ms at the fifth iteration, ensuring reduced complexity and improved accuracy. These optimized features are classified using the TL-EPReLU-CSGNN, which achieves a classification accuracy of 99.2325%, along with high precision (99.321%), recall (99.123%), specificity (99.202%), and F-measure (99.624%). For scene recognition and 3D pose estimation, the model uses the Place365 dataset to ensure background relevance is accurately incorporated. In the final stage, the EPReLU-CSGNN-MALSTCAM module generates textual explanations evaluated using BLEU (0.9665), METEOR (0.9698), ROUGE (0.9879), and CIDEr (0.9756) metrics. Comparative analysis with models such as Bi-LSTM, ICTGAN, and semantic scene encoders further demonstrates the superiority of the proposed approach. Overall, the framework significantly enhances the explainability and precision of image captioning systems and provides a solid foundation for future advancements in cross-modal textual generation using large-scale language models.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An Efficient EPReLU-CSGNN-MALSTCAM and SSOA-Based Explainable Artificial Intelligence (XAI) to Generate Textual Explanations

  • B. P. Sheela,
  • H. Girisha,
  • B. Sreepathi

摘要

In recent years, Explainable Artificial Intelligence (XAI) has received increasing attention in the field of computer vision, particularly for its ability to generate interpretable and semantically rich image captions. However, many existing methods fail to effectively capture the contextual relevance between the primary object and its surrounding environment, leading to generic or misleading textual outputs. To address these limitations, this paper presents a novel and efficient framework that integrates the Exponential Parametric Rectified Linear Unit Convolutional Saliency Gradients Neural Network with MaxAbs Long Short-Term Co-Attention Memory (EPReLU-CSGNN-MALSTCAM) and the Sculptor Schwefel Optimization Algorithm (SSOA) to generate high-quality textual explanations from animal images. The methodology involves comprehensive pre-processing techniques such as contrast enhancement using QCCLAHE, which achieves a PSNR of 36.59 dB and significantly reduces image distortion. Salient features are extracted using a saliency map that emphasizes informative regions based on object-background interactions. For object detection, the proposed AGSPP-YOLO attains an Average Precision (AP) of 98.78% and a Mean AP of 97.84%, outperforming conventional detectors like YOLO and R-CNN. The feature selection phase, optimized through SSOA, yields a fitness value of 97.24 and a selection time of 2154 ms at the fifth iteration, ensuring reduced complexity and improved accuracy. These optimized features are classified using the TL-EPReLU-CSGNN, which achieves a classification accuracy of 99.2325%, along with high precision (99.321%), recall (99.123%), specificity (99.202%), and F-measure (99.624%). For scene recognition and 3D pose estimation, the model uses the Place365 dataset to ensure background relevance is accurately incorporated. In the final stage, the EPReLU-CSGNN-MALSTCAM module generates textual explanations evaluated using BLEU (0.9665), METEOR (0.9698), ROUGE (0.9879), and CIDEr (0.9756) metrics. Comparative analysis with models such as Bi-LSTM, ICTGAN, and semantic scene encoders further demonstrates the superiority of the proposed approach. Overall, the framework significantly enhances the explainability and precision of image captioning systems and provides a solid foundation for future advancements in cross-modal textual generation using large-scale language models.