Review of Various Approaches for Authorship Identification in Digital Forensics
摘要
Authorship identification involves extracting and analyzing the author writing styles. Digital Forensics along with cyber investigations employ writing style to identify the author and their traits. Authorship identification was done on lengthy and short texts in English, Arabic, Chinese, and Greek. However, it emphasizes a unique and difficult scenario: identifying whether two writings published in distinct discourse types are produced by the same person. We provide sets of texts that use four different discourse styles, including essays, emails, text messages, and business notes, based on a new corpus of English texts. The cross-discourse-type ownership verification assignment is highly challenging due to the disparities in communication intent, target audience, and formality level. This paper evaluates various aspects of authorship identification and provides a thorough analysis of the assessment findings. This study also explores the language proficiency and problems in authorship identification tasks. A number of significant authorship identification domain studies were assessed for data, characteristics, techniques, and outcomes. After reviewing the research, we conclude that the outcomes of authorship identification task depend primarily on the specified stylometric characteristics and dataset used. The beneficial qualities also vary by the language type.