The agricultural resources in the Philippines are essential for national food security and economic development with coffee being at its center. Moreover, recent data released by the Philippine Statistics Authority (PSA) show an increase in coffee production although there has been a worrying decline in production in Caraga region which grows over two thousand five hundred growers and has a huge area of land planted to coffee. The FarmVista project addressed this challenge through a data-driven approach by applying Principal Component Analysis (PCA) and various machine learning algorithms to classify and analyze coffee yield in Caraga. The study utilized a comprehensive dataset, the Coffee Farmers Enumerated Data, encompassing socio-demographic details, farming practices, and other influential factors. Gradient Boosting achieved the highest accuracy of 98.69%, with Random Forest closely following at 95.63%. These results highlight the effectiveness of advanced analytics and machine learning in improving coffee yield classification. By uncovering key patterns and factors affecting yield quality, this study provides valuable insights to optimize the coffee value chain in Caraga and addresses the region’s production challenges.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Principal Component Analysis and Machine Learning for Classification of Coffee Yield

  • Vicente A. Pitogo,
  • Cristopher C. Abalorio,
  • Rolyn C. Daguil,
  • Ryan O. Cuarez,
  • Sandra T. Solis,
  • Rex G. Parro

摘要

The agricultural resources in the Philippines are essential for national food security and economic development with coffee being at its center. Moreover, recent data released by the Philippine Statistics Authority (PSA) show an increase in coffee production although there has been a worrying decline in production in Caraga region which grows over two thousand five hundred growers and has a huge area of land planted to coffee. The FarmVista project addressed this challenge through a data-driven approach by applying Principal Component Analysis (PCA) and various machine learning algorithms to classify and analyze coffee yield in Caraga. The study utilized a comprehensive dataset, the Coffee Farmers Enumerated Data, encompassing socio-demographic details, farming practices, and other influential factors. Gradient Boosting achieved the highest accuracy of 98.69%, with Random Forest closely following at 95.63%. These results highlight the effectiveness of advanced analytics and machine learning in improving coffee yield classification. By uncovering key patterns and factors affecting yield quality, this study provides valuable insights to optimize the coffee value chain in Caraga and addresses the region’s production challenges.