Integrated Evaluation Metrics for Large Language Models (MAILLM)
摘要
Large Language Models (LLMs) have established themselves as one of the most transformative technologies in modern artificial intelligence (AI). This study aimed to evaluate the performance of these models using a comprehensive set of metrics to determine their effectiveness and consistency in solving complex questions. The methodology involved preparing 11 multiple-choice questions from the 2024 admission exam of the Technological Institute of Aeronautics (ITA), which were individually administered to the GPT-4 and Gemini models. Each model underwent three rounds of testing, and the responses were analyzed based on metrics of Accuracy, Precision, Response Consistency, and Consistent Errors, in addition to the integrated MAILLM metric developed specifically for this study. The results indicated that GPT-4 exhibited greater variability in responses, although it demonstrated superior performance in some specific tests. In contrast, the Gemini model proved to be more stable but had a lower average accuracy. The MAILLM metric stood out by providing a more comprehensive and integrative evaluation of the different performance aspects of the models, highlighting both strengths and weaknesses. In conclusion, despite significant advancements, both models have areas that require improvement. Future research should explore the applicability and effectiveness of the MAILLM metric in different knowledge contexts, validating its robustness and utility in evaluating language models.