Performance evaluation of tokenizers in large language models for the Assamese language
摘要
This study evaluates the performance of tokenizers in Large Language Models (LLMs) for the Assamese language, a low-resource language from Northeast India. Tokenization plays a pivotal role in LLMs, influencing their efficiency and adaptability. The research compares tokenizers using Normalized Sequence Length (NSL), and token count. The findings reveal that the SUTRA tokenizer outperforms others with the lowest NSL and token count, highlighting its suitability for Assamese. The study identifies significant gaps in Assamese NLP, emphasizing the need for robust tokenizer evaluations to advance research in low-resource languages.