错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Limitations and Benefits of the ChatGPT for Python Programmers and Its Tools for Evaluation

  • Ricardo Arias,
  • Grecia Martinez,
  • Didier Cáceres,
  • Eduardo Garces

摘要

The artificial intelligence called ChatGPT exhibits an outstanding feature, which is its ability to automatically generate code; in this case, the scope is with programming in Python. This paper focuses on evaluating the quality of the code generated and its limitations by ChatGPT compared to the code created by an advanced human programmer. To achieve this purpose, a systematic research is carried out with an exhaustive analysis of the artificial intelligence with a new tool and survey for the evaluation, in order to clarify the capabilities and constraints with its context in the field of artificial intelligence. Through a rigorous systematic review process, following PRISMA guidelines, a total of 6879 relevant publications are initially identified. After applying inclusion and exclusion criteria, the sample was reduced to 165 publications, and after eliminating irrelevant publications, a set of 15 quality articles that fit the study objectives was finally selected. The results of the articles selected in the research reveal an in-depth evaluation of ChatGPT, the understanding of the integration of Artificial Intelligence in education, the analysis of the motivations behind the use of generative chatbots, the challenges and paradigms that arise from the use of AIs in programming, the perception of AI with human attributes, the evaluation of metrics for software quality, the detection of irregularities in the code, the categorization of clones and the exploration of software quality and errors in the code. The PICOC methodology included the use of software related to quality testing, which allows achieving an astonishing 95.66% accuracy in the evaluation of the code generated, both by the programmer and by ChatGPT. This resulted in a 30.28% decrease in human errors and 96.2% effectiveness in evaluating the quality of the generated code, the evaluation considered 3 surveys for Python programmers in 1000 of advanced programmers from October to December 2023. In summary, this paper provides a complete view of the success factors in the comparison of both codes, which in turn leads to greater efficiency in software development, significantly reducing the time required by programmers with its limitation with ChatGPT.