<p>The two most commonly used filter-based feature selection techniques for numerical data are Pearson correlation for identifying linear relationships and Spearman’s rank correlation for capturing monotonic, non-linear associations. However, these methods have limitations, as they rely on rank-based analysis and may not fully capture complex dependencies, making them insufficient for comprehensive feature selection. This study presents a novel approach to feature selection for independent variables, based on the principle that ’features closest to the target variables are significant features required for prediction’. This method, termed Concentric Dilation Correlation (CDC), aims to enhance feature relevance assessment. The CDC method is particularly advantageous due to its user-friendliness and ability to account for interactions between variables. The CDC is employed in the present study to identify the parameters influencing Kerala’s crop productivity (coconut and paddy). The features selected by the CDC are compared against those identified using Stepwise Regression (SR), Random Forest (RF), and Principal Component Analysis (PCA). Skill metrics, including Mean Absolute Error, Root Mean Square Error, Nash-Sutcliffe Efficiency, and Index of Agreement, were used to evaluate the model’s efficiency and performance. Additionally, a random dataset was created using Monte Carlo (MC) simulations in order to validate the CDC. The CDC’s RMSE was 0.93, whereas the PCA, RF, and SR values were 1.26, 0.95, and 1.46, respectively for the testing data. Furthermore, the results show that the CDC technique is effective in improving agricultural productivity predictive modeling, outperforming SR, RF, and PCA across all skill measures in identifying significant features for crop productivity prediction.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A novel concentric-dilation-correlation (CDC) based feature selection technique for paddy and coconut productivity over Kerala

  • Chalissery Mincy Thomas,
  • Archana Nair,
  • K. Somasundaram

摘要

The two most commonly used filter-based feature selection techniques for numerical data are Pearson correlation for identifying linear relationships and Spearman’s rank correlation for capturing monotonic, non-linear associations. However, these methods have limitations, as they rely on rank-based analysis and may not fully capture complex dependencies, making them insufficient for comprehensive feature selection. This study presents a novel approach to feature selection for independent variables, based on the principle that ’features closest to the target variables are significant features required for prediction’. This method, termed Concentric Dilation Correlation (CDC), aims to enhance feature relevance assessment. The CDC method is particularly advantageous due to its user-friendliness and ability to account for interactions between variables. The CDC is employed in the present study to identify the parameters influencing Kerala’s crop productivity (coconut and paddy). The features selected by the CDC are compared against those identified using Stepwise Regression (SR), Random Forest (RF), and Principal Component Analysis (PCA). Skill metrics, including Mean Absolute Error, Root Mean Square Error, Nash-Sutcliffe Efficiency, and Index of Agreement, were used to evaluate the model’s efficiency and performance. Additionally, a random dataset was created using Monte Carlo (MC) simulations in order to validate the CDC. The CDC’s RMSE was 0.93, whereas the PCA, RF, and SR values were 1.26, 0.95, and 1.46, respectively for the testing data. Furthermore, the results show that the CDC technique is effective in improving agricultural productivity predictive modeling, outperforming SR, RF, and PCA across all skill measures in identifying significant features for crop productivity prediction.