Stack Overflow is a vital platform for developers to exchange knowledge, yet its extensive content creates challenges such as information overload and inconsistent quality. While previous research has focused mainly on assessing content quality, effectively identifying high-value and low-value content remains essential for improving information retrieval. Information Foraging Theory (IFT) provides a useful framework for understanding developers’ information-seeking behaviors in information-rich environments like StackOverflow. Building on prior research applying IFT to Q&A websites, this study introduces an IFT-enhanced agglomerative hierarchical clustering approach, classifying 15,000 StackOverflow questions and 15,000 answers into high-value, low-value, high-cost, and low-cost clusters. Results indicate that high-value content is characterized by higher upvotes, greater author reputation, and increased engagement, whereas high-cost content tends to have longer descriptions and additional metadata cues. These findings demonstrate that IFT-driven clustering effectively distinguishes content clusters based on developer-perceived value and cost, offering insights for refining recommendation systems and optimizing information retrieval.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Unveiling Value-Cost Dynamics in StackOverflow with IFT-Enhanced Clustering

  • Abim Sedhain,
  • Sruti Srinivasa Ragavan,
  • Brett McKinney,
  • Shahnewaz Leon,
  • Sandeep Kaur Kuttal

摘要

Stack Overflow is a vital platform for developers to exchange knowledge, yet its extensive content creates challenges such as information overload and inconsistent quality. While previous research has focused mainly on assessing content quality, effectively identifying high-value and low-value content remains essential for improving information retrieval. Information Foraging Theory (IFT) provides a useful framework for understanding developers’ information-seeking behaviors in information-rich environments like StackOverflow. Building on prior research applying IFT to Q&A websites, this study introduces an IFT-enhanced agglomerative hierarchical clustering approach, classifying 15,000 StackOverflow questions and 15,000 answers into high-value, low-value, high-cost, and low-cost clusters. Results indicate that high-value content is characterized by higher upvotes, greater author reputation, and increased engagement, whereas high-cost content tends to have longer descriptions and additional metadata cues. These findings demonstrate that IFT-driven clustering effectively distinguishes content clusters based on developer-perceived value and cost, offering insights for refining recommendation systems and optimizing information retrieval.