Exploring the Leipzig Corpora Collection in the LSP Classroom: A Data-Driven Approach
摘要
The collection and analysis of large text corpora in different languages assume a fundamental role in an increasingly complex globalized world, particularly when dealing with the study of Languages for Special Purposes (LSPs). The access to large volume corpora allows the study and description of linguistic patterns and phenomena. The starting point for this study is the Leipzig Corpora Collection (LCC), which provides access to monolingual dictionaries in several languages, created from newspaper texts or web pages. The LCC provides word frequency information, sample sentences, as well as word co-occurrences, which can be represented by a word graph (Goldhahn et al., p. 759 [1]). The study proposes to analyze a small sample of lexical items from the area of environment and climate change policies in both the German and Portuguese monolingual corpora and compare how these concepts are used and in which contexts they appear in the analyzed journalistic texts. The lexical items were retrieved from word graphs generated from the LCC platform. The study falls within the broader scope of a data-driven learning approach.