Estimating the Number of Annotations Required to Detect Content Types in Historical Newspapers
摘要
Distant reading and distant viewing (complementing close reading/viewing), allow humanities researchers to process larger volumes of source material, effectively increasing the resolution of our view on the human condition. Recently, document layouting methods based on neural networks have shown promise, but it is still a challenge for pre-trained models to perform well when applied to completely novel digitised sources without any fine-tuning, or in cases where a departure from the original model’s classification grammar is a hard requirement. In this paper, we present a new annotated dataset of 423 newspaper pages and the baseline performance of a fine-tuned model on the dataset spanning 46 years. We evaluate the minimal amount of data required to fine-tune a YOLOv8n model to classify advertisements and obituaries in the newspaper Slovenski narod, issued at the turn of the 20th century. We compare the performance of progressively smaller annotation datasets to determine a region of diminishing returns. We show that the increase in performance for every additional annotated image fine-tuning a YOLOv8n model tapers out after about 250 annotated labels regardless of the number of classes trained.