Robustness of Corpus-Based Typological Strategies for Dependency Parsing
摘要
This chapter presents a comparison of the corpus-based typological classification of ten European Union languages obtained using parallel corpora, with one generated in a less controlled scenario, with non-parallel automatically annotated data. First, we described the specific pipeline that was created to extract and annotate multilingual data from the Arquivo.pt 2019 European Parliamentary Elections collection. Two new corpora for all EU languages were generated and made publicly and freely available: one composed of raw texts extracted from this collection and the other with syntactic annotation obtained automatically. Then, we presented an overview of different quantitative typological approaches developed for dependency parsing improvement and selected the most optimised ones to conduct our comparative analysis. Finally, we compared both scenarios using the same corpus-based strategy and showed that the classification obtained using the data provided by the Arquivo.pt dataset provides valuable linguistic information for this type of study, presenting similarities when compared to the classification based on parallel corpora. However, considering the dissimilarities observed, further analysis is required before validating this new method.