Due to its variety, velocity, and volume, efficiently handling and disseminating large volumes of Earth Observation data constitutes a significant challenge. Ongoing efforts to standardize how geospatial data can be described, indexed, and searched have matured, and rich software ecosystems have been developed around specifications and standards such as the Spatio-Temporal Asset Catalog (STAC) specification. With the advent of cloud orchestration frameworks such as Kubernetes, cloud-native approaches for deploying data catalogs, processing services and storage have become paramount for efficiently handling geospatial data. This paper aims to evaluate and compare the performance of multiple STAC server backends, namely those that rely on PgSTAC, Elasticsearch and MongoDB for indexing items and collections metadata by simulating high user loads using the Locust framework for heavy, real-life scenarios involved in building datasets for Machine Learning.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Benchmarking STAC Ecosystem Server Backends in Dataset Creation Applications

  • Alexandru Munteanu,
  • Gabriel Iuhasz,
  • Silviu Panica

摘要

Due to its variety, velocity, and volume, efficiently handling and disseminating large volumes of Earth Observation data constitutes a significant challenge. Ongoing efforts to standardize how geospatial data can be described, indexed, and searched have matured, and rich software ecosystems have been developed around specifications and standards such as the Spatio-Temporal Asset Catalog (STAC) specification. With the advent of cloud orchestration frameworks such as Kubernetes, cloud-native approaches for deploying data catalogs, processing services and storage have become paramount for efficiently handling geospatial data. This paper aims to evaluate and compare the performance of multiple STAC server backends, namely those that rely on PgSTAC, Elasticsearch and MongoDB for indexing items and collections metadata by simulating high user loads using the Locust framework for heavy, real-life scenarios involved in building datasets for Machine Learning.