<p>Archival collections, such as those at the Library of Congress (LoC), often suffer from fragmented metadata and disconnected discovery tools, making research navigation challenging. These collections—comprising catalog records, finding aids, and digital surrogates—frequently lack uniform semantic links, forcing researchers to piece together information from disparate sources. Current discovery systems are limited in their ability to support intuitive or exploratory inquiries, requiring users to possess significant technical expertise or rely on librarian assistance. To address these challenges, this study introduces FolkRAG, a proof-of-concept system that leverages retrieval-augmented generation (RAG). RAG is an advanced natural language processing framework that combines large language models with external information retrieval systems, enhancing access to archival materials through natural language queries and citation-supported responses. FolkRAG is designed to query public data from the American Folklife Center at the LoC, optimizing vector store parameters and integrating advanced retrieval strategies, including the core principles of sentence window retrieval, hypothetical document embedding, and re-ranking. By bridging intellectual and semantic gaps in archival metadata, FolkRAG provides researchers with a cohesive interface for discovering catalog records, finding aids, and digital objects. This approach not only upholds key values such as reliability, transparency, and contextual accuracy but also redefines archival search as a more intuitive and exploratory experience. By showcasing how metadata integration from diverse data sources is used to construct a vector database that powers advanced RAG systems, FolkRAG demonstrates the potential to improve access to cultural heritage materials for both novice and expert users without sacrificing archival practices.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

FolkRAG: a retrieval-augmented generation system for cultural heritage materials

  • Paul Kelly,
  • Jonathan Schild,
  • Amir Jafari

摘要

Archival collections, such as those at the Library of Congress (LoC), often suffer from fragmented metadata and disconnected discovery tools, making research navigation challenging. These collections—comprising catalog records, finding aids, and digital surrogates—frequently lack uniform semantic links, forcing researchers to piece together information from disparate sources. Current discovery systems are limited in their ability to support intuitive or exploratory inquiries, requiring users to possess significant technical expertise or rely on librarian assistance. To address these challenges, this study introduces FolkRAG, a proof-of-concept system that leverages retrieval-augmented generation (RAG). RAG is an advanced natural language processing framework that combines large language models with external information retrieval systems, enhancing access to archival materials through natural language queries and citation-supported responses. FolkRAG is designed to query public data from the American Folklife Center at the LoC, optimizing vector store parameters and integrating advanced retrieval strategies, including the core principles of sentence window retrieval, hypothetical document embedding, and re-ranking. By bridging intellectual and semantic gaps in archival metadata, FolkRAG provides researchers with a cohesive interface for discovering catalog records, finding aids, and digital objects. This approach not only upholds key values such as reliability, transparency, and contextual accuracy but also redefines archival search as a more intuitive and exploratory experience. By showcasing how metadata integration from diverse data sources is used to construct a vector database that powers advanced RAG systems, FolkRAG demonstrates the potential to improve access to cultural heritage materials for both novice and expert users without sacrificing archival practices.