Is Text Normalization Relevant for Classifying Medieval Charters?
摘要
This study examines the impact of historical text normalization on the classification of medieval charters, focusing specifically on document dating and locating. Using a data set of Middle High German charters from the digital archive Monasterium.net, we evaluate various classifiers, including traditional and transformer-based models, both with and without normalization. Our results show that normalization minimally improves locating tasks but reduces accuracy in dating, suggesting that original texts contain crucial features that normalization may obscure. We find that support vector machines and gradient boosting outperform other models, which questions the efficiency of transformers for this use case. The results suggest adopting a selective approach to historical text normalization, highlighting the importance of preserving certain textual characteristics critical for document classification.