Enhancing Textual Deception Detection: A Fused Handcrafted Feature Approach with Machine Learning Models
摘要
Deception in textual communication presents a persistent and significant challenge across numerous domains, including media, finance, and cybersecurity, where accurate identification of misleading information is critical. This paper introduces a novel approach to deception detection by proposing a machine learning model designed to address the inherent complexities of textual deception. The core of our approach lies in the development of a Fused Handcrafted Features vector, which integrates a diverse set of features: lexical patterns, statistical distributions, syntactic structures, and semantic relationships derived from WordNet-based resources. These features are carefully engineered to capture both surface-level and deep linguistic patterns indicative of deceptive behavior. Our model is further enhanced by its ability to analyze sequential dependencies and identify subtle deceptive cues, especially in longer text segments, where traditional methods often struggle. Unlike conventional ML techniques that primarily rely on shallow feature representations, our model leverages advanced architectures to process the nuanced characteristics of deceptive texts. By fusing these handcrafted features with sophisticated ML algorithms, we achieve a robust framework capable of generalizing across diverse datasets. The proposed approach is rigorously evaluated on three widely recognized benchmark datasets, demonstrating its superior performance compared to traditional methods. Our model achieves accuracy improvements ranging from 2% to 7.5%, underscoring its effectiveness in handling various textual deception scenarios. These results highlight the potential of integrating handcrafted linguistic features with machine learning models to address the multifaceted nature of deception. This research not only establishes the efficacy of our proposed Fused Handcrafted Features based model but also opens new avenues for further exploration in real-world applications. By providing a comprehensive analysis of deception through a multi-faceted feature representation, this work lays a solid foundation for future advancements in deception detection technologies, particularly in applications where trust and integrity in textual communication are paramount.