Context: <p>In collaborative software development, the peer code review process proves beneficial only when the reviewers provide useful comments.</p> Objective: <p>This paper investigates the usefulness of Code Review Comments (<i>CR comments</i>) through textual feature-based and featureless approaches.</p> Method: <p>We select three available datasets from both open-source and commercial projects. Additionally, we introduce new features from software and non-software domains. Moreover, we experiment with the presence of jargon, voice, and codes in <i>CR Comments</i> and classify the usefulness of <i>CR Comments</i> through featurization, bag-of-words, and transfer learning techniques.</p> Results: <p>Our models outperform the baseline by achieving state-of-the-art performance. Furthermore, the result demonstrates that the commercial gigantic LLM, GPT-4o, and non-commercial naive featureless approach, Bag-of-Word with TF-IDF, are more effective for predicting the usefulness of <i>CR Comments</i>.</p> Conclusion: <p>The significant improvement in predicting usefulness solely from <i>CR Comments</i> escalates research on this task. Our analyses portray the similarities and differences of domains, projects, datasets, models, and features for predicting the usefulness of <i>CR Comments</i>.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Hold on! is my feedback useful? evaluating the usefulness of code review comments

  • Sharif Ahmed,
  • Nasir U. Eisty

摘要

Context:

In collaborative software development, the peer code review process proves beneficial only when the reviewers provide useful comments.

Objective:

This paper investigates the usefulness of Code Review Comments (CR comments) through textual feature-based and featureless approaches.

Method:

We select three available datasets from both open-source and commercial projects. Additionally, we introduce new features from software and non-software domains. Moreover, we experiment with the presence of jargon, voice, and codes in CR Comments and classify the usefulness of CR Comments through featurization, bag-of-words, and transfer learning techniques.

Results:

Our models outperform the baseline by achieving state-of-the-art performance. Furthermore, the result demonstrates that the commercial gigantic LLM, GPT-4o, and non-commercial naive featureless approach, Bag-of-Word with TF-IDF, are more effective for predicting the usefulness of CR Comments.

Conclusion:

The significant improvement in predicting usefulness solely from CR Comments escalates research on this task. Our analyses portray the similarities and differences of domains, projects, datasets, models, and features for predicting the usefulness of CR Comments.