Enhanced Arabic Retrieval Augmented Generation Using Nested Embedding Models
摘要
Implementing a Retrieval Augmented Generation pipeline by training semantic search models represents a significant advancement in enhancing the accuracy and relevance of question-answering systems. This study focuses on the Arabic language, an area often underrepresented in AI research despite its global significance. Our primary contribution is the development and evaluation of Matryoshka models, state-of-the-art Arabic embeddings designed to optimize semantic understanding and generative capabilities. Utilizing publicly available ARCD and Xtreme datasets as benchmarks, we assess the models’ performance using the average Recall@k metric. Our findings indicate that the Matryoshka variants consistently outperform their base model counterparts, demonstrating superior accuracy and efficiency in retrieving contextually relevant information. This research underscores the potential of Matryoshka models in advancing Arabic language processing and highlights the critical role of training semantic search models in RAG systems.