<p>The digitization of historical documents plays a vital role in preserving cultural heritage and ensuring wide access to valuable information. One of the major challenges in this process is the separation of individual articles from historical newspaper images, a key step for effective text analysis and information retrieval. In this study, we introduce a novel method called <Emphasis Type="BoldItalic">S</Emphasis>emantic <b>T</b>extual-cues leveraged <b>R</b>ule-based approach for <b>A</b>rticle <b>S</b>eparation (STRAS) in historical newspapers. STRAS leverages textual information by extracting text region embeddings from scanned images and their corresponding <i>PAGE</i> format files. Text regions with similar contextual embeddings are grouped, and articles are separated according to a defined rule-set. The approach is evaluated on <i>French</i> and <i>Finnish</i> newspapers from the <InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(19^{th}\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <msup> <mn>19</mn> <mrow> <mi mathvariant="italic">th</mi> </mrow> </msup> </mrow> </math></EquationSource> </InlineEquation> and early <InlineEquation ID="IEq2"> <EquationSource Format="TEX">\(20^{th}\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <msup> <mn>20</mn> <mrow> <mi mathvariant="italic">th</mi> </mrow> </msup> </mrow> </math></EquationSource> </InlineEquation> centuries. Additionally, we propose new metrics for the article separation task: <i>article error rate</i> (AER), <i>article coverage score</i> (ACS), and <i>proper predicted article</i> (PPA). Our evaluation includes several embedding models, such as <i>skip-gram</i> (sgSTRAS), <i>continuous-bag-of-words</i> (cbowSTRAS), <i>FastText</i> (ftSTRAS), and the <i>pre-trained SpaCy</i> model (preSTRAS). Furthermore, we compare these methods with a transfer learning-based model (TLAS) that employs visual features for article separation. The results indicate that the sgSTRAS model achieves the highest <i>mean ACS</i> scores of 0.8343 and 0.8611 on the <i>French</i> and <i>Finnish</i> datasets, respectively, outperforming other models. Our findings highlight the value of semantic textual features and emphasize the importance of the embedding method in improving the performance of article segmentation. To our knowledge, this is the first work that employs a rule-based semantic textual similarity approach for article separation in historical newspapers, filling a gap in existing research and opening up possibilities for future studies.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

STRAS: a semantic textual-cues leveraged Rule-based approach for article separation in historical newspapers

  • Nancy Girdhar,
  • Mickaël Coustaty,
  • Antoine Doucet

摘要

The digitization of historical documents plays a vital role in preserving cultural heritage and ensuring wide access to valuable information. One of the major challenges in this process is the separation of individual articles from historical newspaper images, a key step for effective text analysis and information retrieval. In this study, we introduce a novel method called Semantic Textual-cues leveraged Rule-based approach for Article Separation (STRAS) in historical newspapers. STRAS leverages textual information by extracting text region embeddings from scanned images and their corresponding PAGE format files. Text regions with similar contextual embeddings are grouped, and articles are separated according to a defined rule-set. The approach is evaluated on French and Finnish newspapers from the \(19^{th}\) 19 th and early \(20^{th}\) 20 th centuries. Additionally, we propose new metrics for the article separation task: article error rate (AER), article coverage score (ACS), and proper predicted article (PPA). Our evaluation includes several embedding models, such as skip-gram (sgSTRAS), continuous-bag-of-words (cbowSTRAS), FastText (ftSTRAS), and the pre-trained SpaCy model (preSTRAS). Furthermore, we compare these methods with a transfer learning-based model (TLAS) that employs visual features for article separation. The results indicate that the sgSTRAS model achieves the highest mean ACS scores of 0.8343 and 0.8611 on the French and Finnish datasets, respectively, outperforming other models. Our findings highlight the value of semantic textual features and emphasize the importance of the embedding method in improving the performance of article segmentation. To our knowledge, this is the first work that employs a rule-based semantic textual similarity approach for article separation in historical newspapers, filling a gap in existing research and opening up possibilities for future studies.