Entwicklung eines mittelhochdeutschen Sentiment-Wörterbuchs aus korpushermeneutischer Perspektive
摘要
The first part outlines problems and possible solutions for medieval corpus research. The inadequate availability of digital texts is still a central issue. Although the Mittelhochdeutsche Begriffsdatenbank and some exemplary edition projects are welcome approaches, a non-negotiable obligation to publish the relevant data digitally is overdue for publicly funded edition projects. In addition, medieval studies is at a disadvantage compared to corpus research on newer languages in several ways which are partly due to the non-standardized Middle High German writing and partly to the fact that there are fewer or less powerful tools and resources from the field of automatic language processing for Middle High German.
The second part of the article takes up one such desideratum as an example: Sentiment Analysis. Here, too, medieval studies has some catching up to do, which is why the first Middle High German sentiment dictionary »SentiMhd« is presented. Using automatic methods for sentiment analysis, it is possible to examine large corpora or even sections of text with regard to positive or negative sentiments. Problems are discussed that arise with ambiguous, context-dependent or negated lemmas when different annotators characterize the same text with regard to its sentiment content. Such ambiguities push automatic analysis to its limits in terms of hermeneutics. In addition to the evaluation of SentiMhd using a manually annotated corpus, initial analyses of Hartmann’s Iwein will be presented. Thanks to annotated figure references, sentiment words can also be collected in the context of figure references. The sentiment model indicates that there are rather two low points in Iwein (missed appointment and Lunete’s incarceration), whereas research usually sees only one crisis between the first and second part of the novel. After an evaluation of male and female main characters, maids and opponents, an exemplary macro-view of four different text types concludes the study.