The automated extraction of chemical-induced disease (CID) relationships from biomedical literature has emerged as a critical task in advancing biomedical research and healthcare applications. This task is particularly challenging at the document level, where relevant information often spans multiple sentences and requires understanding complex contextual dependencies. Current approaches typically fall into two categories: graph-based methods that model document structure but struggle with semantic understanding, and transformer-based methods that excel at local context but face challenges in capturing long-range dependencies. In this work, we present a CID relation extraction model that combines graph and transformer-base models, specifically extracting node embedding utilizing pre-trained transformer-based language models as an encoder and using an enhanced graph convolutional neural network to learn more effective node embedding. Through comprehensive experiments on the BioCreative V Chemical Disease Relation (CDR) dataset, we demonstrate that our approach achieves substantial improvements over state-of-the-art methods, showing particular strength in handling complex cases where relationship evidence is dispersed across the document. The empirical results validate our hypothesis that combining transformer-based semantic modeling with structure-aware graph representation learning creates a more robust framework for document-level relation extraction.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Integrating Graph and Transformer-Based Models for Enhanced Chemical-Disease Relation Extraction in Document-Level Contexts

  • Ngoc-Huyen Ngo,
  • Anh-Duc Nguyen,
  • Quynh-Trang Pham Thi,
  • Thanh Hai Dang

摘要

The automated extraction of chemical-induced disease (CID) relationships from biomedical literature has emerged as a critical task in advancing biomedical research and healthcare applications. This task is particularly challenging at the document level, where relevant information often spans multiple sentences and requires understanding complex contextual dependencies. Current approaches typically fall into two categories: graph-based methods that model document structure but struggle with semantic understanding, and transformer-based methods that excel at local context but face challenges in capturing long-range dependencies. In this work, we present a CID relation extraction model that combines graph and transformer-base models, specifically extracting node embedding utilizing pre-trained transformer-based language models as an encoder and using an enhanced graph convolutional neural network to learn more effective node embedding. Through comprehensive experiments on the BioCreative V Chemical Disease Relation (CDR) dataset, we demonstrate that our approach achieves substantial improvements over state-of-the-art methods, showing particular strength in handling complex cases where relationship evidence is dispersed across the document. The empirical results validate our hypothesis that combining transformer-based semantic modeling with structure-aware graph representation learning creates a more robust framework for document-level relation extraction.