Alternative Model Specifications for Big Datasets
摘要
Spatial econometrics is currently experiencing the Big DataBig Data revolution in terms of V: the Volume, the Velocity and the Variety with which spatial data are collected. Indeed, spatial data employed traditionally in the spatial econometric modeling, can be very large, the information being more and more available at a very fine resolution at the level of census tracts, local markets, town blocks, regular grids (or any other small partition of the territory) or even at the level of the single, spatially located, individual. As a consequence, the procedures discussed in this book can become in some cases computational prohibitive because of the vast quantity of data to be analysed. To overcome the computational problems raised by the treatment of large volumes of data, several procedures have been introduced in the literature. However, some of the most recent studies have followed a different route and concentrated, instead, on the definition of some alternative specifications that depart from the paradigms presented in the previous chapters of this book. These contributions have some common characteristics: they are theoretically very simple, they produce closed-form solutions and they improve dramatically the numerical performances. In this chapter we will review some of these alternative specifications presenting, in particular, the matrix exponential spatial specification, the unilateral approximation and the bivariate coding technique. While the bulk of the chapter is focused on the problems raised by the treatment of large volumes of data, the concluding Sect. 5.5 is devoted, instead, to consider the problems connected with the velocity with which, in many applications, spatial data can be acquired. Finally the chapter presents the computer codes in the languages R, STATA and Python needed for the implementation of the techniques presented.