<p>The <b>Po</b>tsdam <b>Te</b>xtbook <b>C</b>orpus (PoTeC) is a naturalistic eye-tracking-while-reading corpus containing data from 75 participants reading 12 scientific texts. PoTeC is the first naturalistic eye-tracking-while-reading corpus that contains eye-movements from domain experts as well as novices in a within-participant manipulation: It is based on a 2<InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="13428_2024_2536_Article_IEq1.gif" Format="GIF" Height="13" Rendition="HTML" Resolution="72" Type="Linedraw" Width="19" /> </InlineMediaObject> <EquationSource Format="TEX">\(\times \)</EquationSource> <EquationSource Format="MATHML"><math> <mo>×</mo> </math></EquationSource> </InlineEquation>2<InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="13428_2024_2536_Article_IEq1.gif" Format="GIF" Height="13" Rendition="HTML" Resolution="72" Type="Linedraw" Width="19" /> </InlineMediaObject> <EquationSource Format="TEX">\(\times \)</EquationSource> <EquationSource Format="MATHML"><math> <mo>×</mo> </math></EquationSource> </InlineEquation>2 fully crossed factorial design, which includes the participants’ level of studies and the participants’ discipline of studies as between-subjects factors and the text domain as a within-subjects factor. The participants’ reading comprehension was assessed by a series of text comprehension questions and their domain knowledge was tested by text-independent background questions for each of the texts. The materials are annotated for a variety of linguistic features at different levels. We envision PoTeC to be used for a wide range of studies including but not limited to analyses of expert and non-expert reading strategies. The corpus and all the accompanying data <i>at all stages of the preprocessing pipeline</i> and <i>all</i> code used to preprocess the data is made available via GitHub: <a href="https://github.com/DiLi-Lab/PoTeC">https://github.com/DiLi-Lab/PoTeC</a> and OSF: <a href="https://osf.io/dn5hp/">https://osf.io/dn5hp/</a>. The data is furthermore integrated into the open-source package <Emphasis FontCategory="NonProportional">pymovements</Emphasis>, which can be used in Python and R: <a href="https://github.com/aeye-lab/pymovements">https://github.com/aeye-lab/pymovements</a>.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

PoTeC: A German naturalistic eye-tracking-while-reading corpus

  • Deborah N. Jakobi,
  • Thomas Kern,
  • David R. Reich,
  • Patrick Haller,
  • Lena A. Jäger

摘要

The Potsdam Textbook Corpus (PoTeC) is a naturalistic eye-tracking-while-reading corpus containing data from 75 participants reading 12 scientific texts. PoTeC is the first naturalistic eye-tracking-while-reading corpus that contains eye-movements from domain experts as well as novices in a within-participant manipulation: It is based on a 2 \(\times \) × 2 \(\times \) × 2 fully crossed factorial design, which includes the participants’ level of studies and the participants’ discipline of studies as between-subjects factors and the text domain as a within-subjects factor. The participants’ reading comprehension was assessed by a series of text comprehension questions and their domain knowledge was tested by text-independent background questions for each of the texts. The materials are annotated for a variety of linguistic features at different levels. We envision PoTeC to be used for a wide range of studies including but not limited to analyses of expert and non-expert reading strategies. The corpus and all the accompanying data at all stages of the preprocessing pipeline and all code used to preprocess the data is made available via GitHub: https://github.com/DiLi-Lab/PoTeC and OSF: https://osf.io/dn5hp/. The data is furthermore integrated into the open-source package pymovements, which can be used in Python and R: https://github.com/aeye-lab/pymovements.