错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Retrieving and Analyzing Translations of American Newspaper Comics with Visual Evidence

  • Jacob Murel,
  • David A. Smith

摘要

Research on image classification and text translation for comics have transpired largely independent of one another. Machine translation tools focus on comics’ text features, thereby largely ignoring comics’ heavily visual dimension. Image classification applications for comics focus primarily on genre and artist attribution. This paper bridges the gap between these areas by investigating image classification model accuracy for identifying translations of American newspaper comic strips. How might machine learning algorithms leverage comics’ distinguishing visual features in order to identify pre-existing translations? To what extent do textual differences affect classification model accuracy in identifying otherwise identical comics? Using a dataset of \(\tilde{1}\) 8,000 English and Spanish comics, we generate embeddings from three CNNs and a Vision Transformer. We generate additional embeddings from binarized images and images with text redacted using an OCR model. We compute the cosine distance between given pairs of comics and evaluate its accuracy at retrieving translations. The best models rank the true translation first for 97% of queries, falling to 94% when the language is not known.