Enhancing Biomedical Event Extraction with Error Data Detection: A Novel Approach for Improved Classification Performance
摘要
Bio-event extraction is vital for understanding biomolecular interactions and aiding applications like drug repurposing and pathway curation. Yet, its effectiveness is hindered by the natural language’s complexity, making robust methodologies essential. Our research introduces a novel strategy for enhancing supervised machine learning in biomedical text mining by reducing mislabeled instances in training data. We developed a method to identify and filter out errors in unlabeled data through vector representation of sample pairs, improving classifier accuracy. Tested in the BioNLP Shared Task GENIA, our approach effectively enhances biomedical event extraction, presenting a significant leap forward by constructing more reliable prediction models.