错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

The Aranea Corpora Family: Ten+ Years of Processing Web-Crawled Data

  • Vladimír Benko

摘要

Aranea is a project to create a family of web-crawled corpora for languages taught at Slovak Universities. Since 2013, more than two dozen languages have been added to the project: they are often represented with Gigaword+ size corpora and sometimes have subcorpora for their territorial varieties. Our paper summarizes the development of the Aranea project in the past decade. We describe the step-by-step optimization of the processing pipeline, highlight existing issues, and discuss the linguistic rationale behind some engineering decisions associated with the idiosyncrasies of individual languages.