<p>The real estate market involves large sums of money and is linked to the entire supply chain of the construction industry. The lack of standardization in property prices and the high fluctuations in the sector present opportunities for the use of machine learning (ML) in this type of problem. This systematic literature review analyzed research trends and identified research gaps in the field of ML applied to real state pricing. The results elucidated that the most frequently used models were Random Forest, Gradient Boosting Machine and XGBoost, Linear Regression, and Artificial Neural Networks. The “location”, “condominium fee”, and “property area” were the most employed features in each of the main categories: locational, financial, and physical, respectively. The main data sources for the ML models were real estate websites, which can present significant bias. A clear relationship was not observed between model quality metrics and the amount of data, nor the algorithm employed. The performance of the algorithms varied according to the database characteristics. The knowledge gaps identified are related to the impact of the time on property sale value, as well as the implementation of a correction factor that follows the sector's temporality and improves future predictions.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Systematic Review of the use of Machine Learning in the Prediction of House Pricing

  • Romário Parreira Pita,
  • Aldo Ribeiro de Carvalho,
  • Rafaela Miranda Barbosa,
  • Alexandre Abrahão Cury,
  • Julia Castro Mendes

摘要

The real estate market involves large sums of money and is linked to the entire supply chain of the construction industry. The lack of standardization in property prices and the high fluctuations in the sector present opportunities for the use of machine learning (ML) in this type of problem. This systematic literature review analyzed research trends and identified research gaps in the field of ML applied to real state pricing. The results elucidated that the most frequently used models were Random Forest, Gradient Boosting Machine and XGBoost, Linear Regression, and Artificial Neural Networks. The “location”, “condominium fee”, and “property area” were the most employed features in each of the main categories: locational, financial, and physical, respectively. The main data sources for the ML models were real estate websites, which can present significant bias. A clear relationship was not observed between model quality metrics and the amount of data, nor the algorithm employed. The performance of the algorithms varied according to the database characteristics. The knowledge gaps identified are related to the impact of the time on property sale value, as well as the implementation of a correction factor that follows the sector's temporality and improves future predictions.