Genealogy extraction from ancient texts is a challenging but significant task for genealogical research due to its historical data. To explore each family history, one has to read large amounts of text to find the depth of the relationships between characters. Recent advancements in Natural Language Processing (NLP), especially for information extraction and visualization, have become significantly more accurate, convenient, and interactive due to Artificial Intelligence (AI) integration, enabling them to handle complex language tasks with ever-increasing accuracy through continuous learning from data. This paper first explores the application of modern information extraction (IE) for automated genealogy extraction from ancient texts by incorporating the advanced large language modeling capabilities through OpenAI’s GPT models that are integrated into the LangChain framework. This approach not only identifies genealogical information but also interprets and contextualizes relationships using advanced natural language understanding capabilities. The LangChain framework facilitates a seamless connection between OpenAI’s GPT (generative pre-trained transformer) models and the Neo4j graph database, enabling the construction of dynamic and context-aware family trees. Then this paper compares the strengths and limitations of the automatic extraction of genealogical data from unstructured ancient texts using GPT models.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Comparative Study on Automatic Extraction of Genealogical Data from Ancient Text by Using Large Language Models and Open AI

  • Violet Chella Christy Alex Xavier,
  • Md Mahbub Ul Alam,
  • Doina Logofătu

摘要

Genealogy extraction from ancient texts is a challenging but significant task for genealogical research due to its historical data. To explore each family history, one has to read large amounts of text to find the depth of the relationships between characters. Recent advancements in Natural Language Processing (NLP), especially for information extraction and visualization, have become significantly more accurate, convenient, and interactive due to Artificial Intelligence (AI) integration, enabling them to handle complex language tasks with ever-increasing accuracy through continuous learning from data. This paper first explores the application of modern information extraction (IE) for automated genealogy extraction from ancient texts by incorporating the advanced large language modeling capabilities through OpenAI’s GPT models that are integrated into the LangChain framework. This approach not only identifies genealogical information but also interprets and contextualizes relationships using advanced natural language understanding capabilities. The LangChain framework facilitates a seamless connection between OpenAI’s GPT (generative pre-trained transformer) models and the Neo4j graph database, enabling the construction of dynamic and context-aware family trees. Then this paper compares the strengths and limitations of the automatic extraction of genealogical data from unstructured ancient texts using GPT models.