Unveiling Data Clumps: A Detailed Longitudinal Analysis on Software Quality Across Public Repositories
摘要
As software evolves, groups of variables often recur, forming data clumps. This study extends our previous work by applying our detection tool to 16 public repositories covering 2696 tagged versions. We mined over 6 million additional data clumps and updated the publicly available dataset. The analysis confirms that data clumps persist over time and tend to grow into larger clusters, which hinders refactoring. Variables such as name and key dominate across projects, whereas credential variables occur rarely. These results consolidate earlier observations and provide a stronger foundation for studying the connection between data clumps and faults.