Benchmarking STAC Ecosystem Server Backends in Dataset Creation Applications
摘要
Due to its variety, velocity, and volume, efficiently handling and disseminating large volumes of Earth Observation data constitutes a significant challenge. Ongoing efforts to standardize how geospatial data can be described, indexed, and searched have matured, and rich software ecosystems have been developed around specifications and standards such as the Spatio-Temporal Asset Catalog (STAC) specification. With the advent of cloud orchestration frameworks such as Kubernetes, cloud-native approaches for deploying data catalogs, processing services and storage have become paramount for efficiently handling geospatial data. This paper aims to evaluate and compare the performance of multiple STAC server backends, namely those that rely on PgSTAC, Elasticsearch and MongoDB for indexing items and collections metadata by simulating high user loads using the Locust framework for heavy, real-life scenarios involved in building datasets for Machine Learning.