<p>Automatic retrieval of formal business process models from their natural language descriptions is a well-established way to facilitate the time- and cost-intensive modeling procedure. Yet, a lack of data usable for developing and training new retrieval methods is impeding progress in this field of research. This issue can be overcome by either using methods less reliant on high-quality data, such as large language models, or creating bigger datasets. The latter is often preferable in the context of business process modeling, especially when internal workflows of organizations have to be treated confidentially. It is the more data-intensive solution, though, which is costly. Data augmentation techniques aim to improve both quality and quantity of existing datasets, by deliberate perturbations resulting in new, synthetic data. In this article, we present a collection of data augmentation techniques, which are specifically selected for the task of improving data quality in the context of process information extraction. We show why data augmentation techniques from the wider field of natural language processing are often not applicable to process information extraction, and how the resulting data differ in terms of linguistic variety, structure, and feature space coverage. In our experiments, data augmentation results in an absolute improvement in the <InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(F_1\)</EquationSource> <EquationSource Format="MATHML"><math> <msub> <mi>F</mi> <mn>1</mn> </msub> </math></EquationSource> </InlineEquation> measure of <InlineEquation ID="IEq2"> <EquationSource Format="TEX">\(5.7\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>5.7</mn> <mo>%</mo> </mrow> </math></EquationSource> </InlineEquation> for extracting process-relevant entities from text and <InlineEquation ID="IEq3"> <EquationSource Format="TEX">\(4.5\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>4.5</mn> <mo>%</mo> </mrow> </math></EquationSource> </InlineEquation> for extracting relations between those entities. We make all code available at <a href="https://github.com/JulianNeuberger/pet-data-augmentation">https://github.com/JulianNeuberger/pet-data-augmentation</a> and results for our experiments at <a href="https://zenodo.org/doi/10.5281/zenodo.10941423">https://zenodo.org/doi/10.5281/zenodo.10941423</a>.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Repeat, reorder, rephrase: data augmentation for process information extraction

  • Julian Neuberger,
  • Lars Ackermann,
  • Stefan Jablonski

摘要

Automatic retrieval of formal business process models from their natural language descriptions is a well-established way to facilitate the time- and cost-intensive modeling procedure. Yet, a lack of data usable for developing and training new retrieval methods is impeding progress in this field of research. This issue can be overcome by either using methods less reliant on high-quality data, such as large language models, or creating bigger datasets. The latter is often preferable in the context of business process modeling, especially when internal workflows of organizations have to be treated confidentially. It is the more data-intensive solution, though, which is costly. Data augmentation techniques aim to improve both quality and quantity of existing datasets, by deliberate perturbations resulting in new, synthetic data. In this article, we present a collection of data augmentation techniques, which are specifically selected for the task of improving data quality in the context of process information extraction. We show why data augmentation techniques from the wider field of natural language processing are often not applicable to process information extraction, and how the resulting data differ in terms of linguistic variety, structure, and feature space coverage. In our experiments, data augmentation results in an absolute improvement in the \(F_1\) F 1 measure of \(5.7\%\) 5.7 % for extracting process-relevant entities from text and \(4.5\%\) 4.5 % for extracting relations between those entities. We make all code available at https://github.com/JulianNeuberger/pet-data-augmentation and results for our experiments at https://zenodo.org/doi/10.5281/zenodo.10941423.