We provide an analysis of the squared Wasserstein-2 ( \(W_2\) ) distance between two probability distributions associated with two stochastic differential equations (SDEs). Based on this analysis, we propose using squared \(W_2\) distance-based loss functions to train parametrized neural networks in order to reconstruct SDEs from noisy data. Specifically, we propose minimizing a time-decoupled squared \(W_2\) distance loss function. To demonstrate the practicality of our Wasserstein distance-based loss functions, we performed numerical experiments that demonstrate the efficiency of our method in learning SDEs that arise across a number of applications.