错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An Application of Support Vector Machine, Random Forest, and Related Machine Learning Algorithms on California Wildfire Data

  • Joshua Ologbonyo,
  • Roger B. Sidje

摘要

Computerized weather predictions have allowed the National Weather Service to identify meteorological conditions that give rise to thunderstorms and tornadoes, and therefore issue advance warnings to the population. This suggests that wildfires could similarly benefit from mathematical simulations and data analysis. In this study, we extract the subset of data related to California in the U.S. wildfire data from 1992 to 2018, with the aim of gaining insights into the causes of wildfires and their sizes (acres burned). We perform this using support vector machine (SVM), random forest (RF) and related machine learning algorithms. In addition to seeking to predict a fire size from its given ignition point, our study also sought to predict the fire duration using ensemble regression methods. Of the multi-output methods considered, we observed that random forest regression significantly outperformed multi-output support vector regression (MSVR). Looking into the causes of fires, we noted that human activities such as arson, debris burning, smoking, and accidental in-home ignitions were responsible for numerous fire incidents and accounted for significant burned acres. Machine learning methodologies that build on mathematical modeling and computational approaches can improve the knowledge of fire locations, sizes and durations, which in turn can inform fire-fighting strategies (both the allocation of equipment and human resources) as well as guide future re-evaluation of decision and policy making. Informed public policies help preparedness for the yearly fires that affect large swaths of the U.S., notably California, Texas and Florida. While our study explores California wildfire data, we use methodologies that readily extend where similar datasets are available.