Boundary Extraction for Unseen Entities Prediction Using Bidirectional Long Short-Term Memory
摘要
Named Entities Recognition (NER), which is a task to identify and classify named entities within text, has gained significant popularity in recent years. This task often requires pre-labeled data and large datasets to achieve high accuracy. A key challenge in NER is extracting unseen entities that have not been previously labeled or are new terms, posing difficulties in adapting to rapidly changing data and the emergence of new vocabularies. In this article, we propose a model utilizing Bidirectional Long Short-Term Memory (Bi-LSTM) for boundary extraction. The model aims to enhance the detection and labeling of unseen entities by leveraging Bi-LSTM's strengths in capturing contextual information and accurately identifying entities boundaries. The model demonstrates high potential in identifying unseen names without pre-labeling. Our study demonstrates the model's showcasing its effectiveness in discovering unlabeled words and unseen entities in Thai language datasets.