The evaluation of language models is crucial in determining their effectiveness across various NLP tasks. This study investigates the performance of four prominent language models: Turkcell-LLM-7b, Trendyol-LLM-7b, Gemma-7b, and Gemma2-2b. Using a comprehensive set of evaluation metrics, including ROUGE, BLEU, BERTScore, semantic similarity, and cosine similarity, we analyzed their ability to generate high-quality responses. Our research is motivated by the need to understand how these models perform in diverse linguistic contexts and tasks, aiming to bridge the gap between lexical overlap and semantic understanding. The dataset consists of 1,446 question-answer pairs related to Turkish Rent Law, with an average of 9.03 words per question and 11.27 words per answer. The evaluation reveals distinct strengths for each model, with Gemma-7b excelling in ROUGE metrics and Turkcell-LLM-7b showing superior semantic alignment through BERTScore. Furthermore, Trendyol-LLM-7b demonstrated competitive precision in BLEU evaluations, while Gemma2-2b showcased robust performance in cosine similarity assessments. The word count analysis indicates significant differences in response length among the models, with Turkcell-LLM-7b generating the most detailed answers (51,302 words in total, averaging 35.48 words per answer), whereas Gemma-7b and Gemma2-2b produced more concise responses (12.73 and 13.15 words per answer, respectively). These findings underscore the importance of using varied evaluation metrics to capture the multifaceted nature of language generation quality while also highlighting the impact of response length on model performance in legal NLP applications.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Computational Intelligence in Legal NLP: Evaluating Language Models on Turkish Legal Texts

  • Berken Cam,
  • Mehmet Ali Kömürcü,
  • Merve Kaya,
  • Ahmet Emre Ergün,
  • Tuğba Çelikten,
  • Aytuğ Onan

摘要

The evaluation of language models is crucial in determining their effectiveness across various NLP tasks. This study investigates the performance of four prominent language models: Turkcell-LLM-7b, Trendyol-LLM-7b, Gemma-7b, and Gemma2-2b. Using a comprehensive set of evaluation metrics, including ROUGE, BLEU, BERTScore, semantic similarity, and cosine similarity, we analyzed their ability to generate high-quality responses. Our research is motivated by the need to understand how these models perform in diverse linguistic contexts and tasks, aiming to bridge the gap between lexical overlap and semantic understanding. The dataset consists of 1,446 question-answer pairs related to Turkish Rent Law, with an average of 9.03 words per question and 11.27 words per answer. The evaluation reveals distinct strengths for each model, with Gemma-7b excelling in ROUGE metrics and Turkcell-LLM-7b showing superior semantic alignment through BERTScore. Furthermore, Trendyol-LLM-7b demonstrated competitive precision in BLEU evaluations, while Gemma2-2b showcased robust performance in cosine similarity assessments. The word count analysis indicates significant differences in response length among the models, with Turkcell-LLM-7b generating the most detailed answers (51,302 words in total, averaging 35.48 words per answer), whereas Gemma-7b and Gemma2-2b produced more concise responses (12.73 and 13.15 words per answer, respectively). These findings underscore the importance of using varied evaluation metrics to capture the multifaceted nature of language generation quality while also highlighting the impact of response length on model performance in legal NLP applications.