Authorship Attribution for Assamese Language Documents: Initial Results
摘要
Impact of Digital India in the creation of electronic content on the web is primarily acknowledged. However, critically observing, we also realize the problems, especially in the cases of identifying the creator of content. Also, there can be issues of false annotation or even plagiarism. Every person has a unique style of writing. This characteristic can be explored to further train a system to identify the accurate author of a content, also popularly known as the field of authorship attribution. This paper attempts to showcase the initial experimental results of author identification done on a manually collected and annotated assamese literary corpus. Assamese being a low-resourced language, the applications of NLP like authorship identification has not been explored till date. And this reported work can be marked as the first attempt of bringing to the world the research scopes of authorship attribution in assamese language.