Machine learning technique for generation of human readable rules to detect software code smells in open-source software
摘要
Defects entering software systems due to bad programming practice during evolution and maintenance are termed code smells. Smells impacts software at design, architectural and implementation level. These flaws report depreciation in quality of software thereby wasting human effort and time. The serious outcomes of these defects motivated researchers to study code smells. This study is conducted to determine the most daunting smell of architectural, design and implementation category of code smell. We considered 60 metrics of Apache Tomcat to analyse the quality aspects of the software. Twelve Machine Learning (ML) algorithms are applied on instances of architectural smells, implementation smells and design smells extracted from source code of 15 versions of open-source software. J48 machine learning algorithm gave 98% accurate result in code smell detection. The results are validated against tenfold cross validation and the statistical parameters: Kappa statistics, Recall, Precision, etc. used for analysing the results. Human readable classification rules are generated and validated using tenfold cross validation. Results are analysed using several Precision, Recall, F-Measure, TP Rate, FP Rate, ROC Area, Kappa statistics and accuracy performance metrics. The study also focusses on architectural category which has yet not been fully explored.