This study proposes an innovative approach to designing short videos for science communication tailored to aging populations. By leveraging web scraping and text mining techniques, key elements such as content style, font selection, visuals, background music, color schemes, and video length are extracted and clustered using constraint-based k-means clustering. Domain experts contribute predefined centroids, enhancing the clustering process's relevance and accuracy. The methodology involves text pre-processing, vectorization using Term Frequency, and evaluation via recall analysis, achieving values between 0.71 and 0.75. Clusters focusing on visuals, color schemes, and video length show the highest recall, reflecting their well-defined nature, while broader topics like content style perform slightly lower. The results highlight the efficiency and scalability of combining automated methods with expert input, providing a robust framework for creating engaging and accessible science communication content. Future work includes integrating semantic embedding techniques to improve clustering outcomes and address broader categories more effectively.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Clustering-Based Approach for Identifying Key Information to Develop Short Video Prototypes in Science Communication for Aging Populations

  • Cong Mo,
  • Jantima Polpinij,
  • Khachakrit Liamthaisong,
  • Chunpong Chan,
  • Bancha Luaphol

摘要

This study proposes an innovative approach to designing short videos for science communication tailored to aging populations. By leveraging web scraping and text mining techniques, key elements such as content style, font selection, visuals, background music, color schemes, and video length are extracted and clustered using constraint-based k-means clustering. Domain experts contribute predefined centroids, enhancing the clustering process's relevance and accuracy. The methodology involves text pre-processing, vectorization using Term Frequency, and evaluation via recall analysis, achieving values between 0.71 and 0.75. Clusters focusing on visuals, color schemes, and video length show the highest recall, reflecting their well-defined nature, while broader topics like content style perform slightly lower. The results highlight the efficiency and scalability of combining automated methods with expert input, providing a robust framework for creating engaging and accessible science communication content. Future work includes integrating semantic embedding techniques to improve clustering outcomes and address broader categories more effectively.