错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

CONCORD: enhancing COVID-19 research with weak-supervision based numerical claim extraction

  • Dhwanil Shah,
  • Krish Shah,
  • Manan Jagani,
  • Agam Shah,
  • Bhaskar Chaudhury

摘要

The COVID-19 Numerical Claims Open Research Dataset (CONCORD) is a comprehensive, open-source dataset that extracts numerical claims from academic papers on COVID-19 research. A weak-supervision model is employed for this extraction, taking advantage of its white-box, explainable nature and reduced computational and annotation costs compared to transformer-based models. This model uses labelling functions such as pattern matching, external knowledge bases, phrase matching, and third-party models to generate labels, with an aggregator function handling contradictory labels. Evaluated against established baselines, the model achieved a weighted F1-score of 0.932 and a micro F1-score of 0.930. While transformer-based models achieve comparable results, the explainability of weak-supervision offers distinct advantages. Additionally, generative LLMs were tested to understand their effectiveness in extracting numerical claims, highlighting the impact of prompt engineering on performance. CONCORD contains approximately 200,000 numerical claims from over 57,000 COVID-19 research articles, serving as a valuable resource for tracking developments in COVID-19 research. This dataset, coupled with the weak-supervision approach, provides researchers with a significant tool for advancing COVID-19 research and showcases the potential of these methodologies in the broader biomedical field.