<p>Information Extraction aims to analyze and extract relevant information in document samples. For visual documents, such as academic and commercial forms, key-value pair extraction is capable of extracting and grouping the requested information automatically. The state of the art presents multimodal models capable of analyzing document’s image, text and layout. To study and use these models in downstream tasks, labeled datasets must be available for fine-tuning of models. However, publicly available data is scarce, even more so for the Portuguese language. In this context, this work aims to provide a dataset of academic forms, named UFLA-FORMS, for the task of Information Extraction. The dataset is composed of 200 manually labeled samples, containing 7710 entities with 4442 relationships pairs between them. The labeling obtained a Kappa coefficient of agreement of 0.918 for the labeling of entities and 0.909 for the relationships attributed between them. The dataset was experimentally evaluated through cross-validation with hyper-parameter search in Named Entity Recognition and Relation Extraction tasks, obtaining, respectively, an average <InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="10579_2024_9802_Article_IEq1.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="18" /> </InlineMediaObject> <EquationSource Format="TEX">\(F_1\)</EquationSource> <EquationSource Format="MATHML"><math> <msub> <mi>F</mi> <mn>1</mn> </msub> </math></EquationSource> </InlineEquation> of 0.921 and 0.846.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

UFLA-FORMS: an academic forms dataset for information extraction in the Portuguese language

  • Victor Gonçalves Lima,
  • Denilson Alves Pereira

摘要

Information Extraction aims to analyze and extract relevant information in document samples. For visual documents, such as academic and commercial forms, key-value pair extraction is capable of extracting and grouping the requested information automatically. The state of the art presents multimodal models capable of analyzing document’s image, text and layout. To study and use these models in downstream tasks, labeled datasets must be available for fine-tuning of models. However, publicly available data is scarce, even more so for the Portuguese language. In this context, this work aims to provide a dataset of academic forms, named UFLA-FORMS, for the task of Information Extraction. The dataset is composed of 200 manually labeled samples, containing 7710 entities with 4442 relationships pairs between them. The labeling obtained a Kappa coefficient of agreement of 0.918 for the labeling of entities and 0.909 for the relationships attributed between them. The dataset was experimentally evaluated through cross-validation with hyper-parameter search in Named Entity Recognition and Relation Extraction tasks, obtaining, respectively, an average \(F_1\) F 1 of 0.921 and 0.846.