In this study, the authors explore a new approach in the development and optimisation of a Transformer-based large language model deployed in a cloud environment. The system uses a memory-augmented Transformer architecture to retrieve and integrate domain-specific information, producing contextually accurate and lexically coherent responses. To evaluate performance, token-level metrics (accuracy, precision, recall, F1-score) and sequence-level metrics (ROUGE) are used to show the model’s generation quality. Compared to an earlier study on Llama 2 fine-tuning, where low F1 Score and limited textual alignment were frequent, the current improved version of our model acquired significant performance, achieving high accuracy, precision, recall, and F1 Score but also robust similarity metrics.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Development of a Transformer-Based Large Language Model Architecture in Cloud

  • Theodor-Radu Grumeza,
  • Thomas-Andrei Lazăr,
  • Alexandra-Emilia Forti

摘要

In this study, the authors explore a new approach in the development and optimisation of a Transformer-based large language model deployed in a cloud environment. The system uses a memory-augmented Transformer architecture to retrieve and integrate domain-specific information, producing contextually accurate and lexically coherent responses. To evaluate performance, token-level metrics (accuracy, precision, recall, F1-score) and sequence-level metrics (ROUGE) are used to show the model’s generation quality. Compared to an earlier study on Llama 2 fine-tuning, where low F1 Score and limited textual alignment were frequent, the current improved version of our model acquired significant performance, achieving high accuracy, precision, recall, and F1 Score but also robust similarity metrics.