Implementing a Retrieval Augmented Generation pipeline by training semantic search models represents a significant advancement in enhancing the accuracy and relevance of question-answering systems. This study focuses on the Arabic language, an area often underrepresented in AI research despite its global significance. Our primary contribution is the development and evaluation of Matryoshka models, state-of-the-art Arabic embeddings designed to optimize semantic understanding and generative capabilities. Utilizing publicly available ARCD and Xtreme datasets as benchmarks, we assess the models’ performance using the average Recall@k metric. Our findings indicate that the Matryoshka variants consistently outperform their base model counterparts, demonstrating superior accuracy and efficiency in retrieving contextually relevant information. This research underscores the potential of Matryoshka models in advancing Arabic language processing and highlights the critical role of training semantic search models in RAG systems.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhanced Arabic Retrieval Augmented Generation Using Nested Embedding Models

  • Omer Nacar,
  • Serry Sibaee,
  • Anis Koubaa

摘要

Implementing a Retrieval Augmented Generation pipeline by training semantic search models represents a significant advancement in enhancing the accuracy and relevance of question-answering systems. This study focuses on the Arabic language, an area often underrepresented in AI research despite its global significance. Our primary contribution is the development and evaluation of Matryoshka models, state-of-the-art Arabic embeddings designed to optimize semantic understanding and generative capabilities. Utilizing publicly available ARCD and Xtreme datasets as benchmarks, we assess the models’ performance using the average Recall@k metric. Our findings indicate that the Matryoshka variants consistently outperform their base model counterparts, demonstrating superior accuracy and efficiency in retrieving contextually relevant information. This research underscores the potential of Matryoshka models in advancing Arabic language processing and highlights the critical role of training semantic search models in RAG systems.