错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Extraction and Processing of Web Content for Corpus Creation: A Systematic Literature Review

  • Jair Alfredo Flores Luna,
  • Miguel Hidalgo Reyes,
  • Virginia Lagunes Barradas

摘要

The processes and methods of text extraction and pre-processing for corpus generation are not widely documented, especially when it comes to Spanish texts. The majority of the documents that collect this information are in English and focus on research carried out in the United States or in Asian countries. The aim of this systematic literature review is to know the state of the art of the technologies and methods used for the extraction of text from web platforms and the pre-processing to generate a specific corpus. Thanks to this review, the issues defined by the research questions have been addressed and an area of opportunity has been identified for the development of new projects focused on the extraction of web information and the creation of corpora to perform analysis in the text mining and natural language processing fields.