Unveiling the Potential of Synthetic Data in Sports Science: A Comparative Study of Generative Methods
摘要
In the pursuit of understanding fatigue in optimizing sports training, acquiring sufficient biological data poses a significant challenge, especially in scenarios involving invasive data collection. To address this limitation, we explored the idea of generating synthetic time-series data from a constrained dataset of five athletes, including daily metrics such as sleep quality, mood, training load (Foster load), and an indicator of the intrinsic antioxidant state (O2score) obtained through a blood sample. We compared diverse synthetic data generation algorithms, including classical approaches like k-medoids and deep learning approaches like Variational Auto-Encoders (VAE), generative adversarial networks (TimeGAN) and Autoregressive Denoising Diffusion Models (TimeGrad). To evaluate the quality of the generated data we focused on assessing both, the fidelity of synthetic data by comparing diverse measures of similarity of data distributions, and the utility of these data by measuring the generalization capabilities of models trained on them. In the comparative analysis of synthetic data generation methods, TimeGan emerges with promising results, effectively balancing utility and fidelity to the original data. This study contributes not only to the methodological landscape of synthetic data generation but also offers valuable insights into the utility and limitations of synthetic datasets and illustrates potential applications in enhancing sports data analysis, particularly in scenarios where invasive data collection is impractical.