The volume of text documents on the web is growing dramatically as data comes in a variety of forms. In Natural Language Processing (NLP), the genre of a literary work presents a challenging foundation for classifying ambiguous text, such as short poems. It is imperative to analyse a variety of features (requiring different pre-processing) that capture multiple aspects of the text. This work experiments with several prominent classifiers to classify short English poems by genre, primarily guided by a hybrid dimensionality reduction strategy and a feature bucket technique to develop deep insights into the strengths of paired and individual feature sets. The contributions of poetic feature combinations and the superior performance of Logistic Regression (LR), Linear Discriminant Analysis classifier (LDA), and Support Vector Classifier (SVC) are presented. The highest mean balanced accuracy (ACC) of 72.61% and a weighted F-score ( \(F_1\) ) of 72.97% were achieved by LR trained with lexical features, outperforming the mixture of all features.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Exploring the Strength of Extensive Features in Short Poem Genre Classification Using Advanced Feature Engineering

  • B. Lavanya,
  • R. Sowmiya

摘要

The volume of text documents on the web is growing dramatically as data comes in a variety of forms. In Natural Language Processing (NLP), the genre of a literary work presents a challenging foundation for classifying ambiguous text, such as short poems. It is imperative to analyse a variety of features (requiring different pre-processing) that capture multiple aspects of the text. This work experiments with several prominent classifiers to classify short English poems by genre, primarily guided by a hybrid dimensionality reduction strategy and a feature bucket technique to develop deep insights into the strengths of paired and individual feature sets. The contributions of poetic feature combinations and the superior performance of Logistic Regression (LR), Linear Discriminant Analysis classifier (LDA), and Support Vector Classifier (SVC) are presented. The highest mean balanced accuracy (ACC) of 72.61% and a weighted F-score ( \(F_1\) ) of 72.97% were achieved by LR trained with lexical features, outperforming the mixture of all features.