Optimizing Temporal and Spectral Interval Selection for Imagined Speech Decoding from EEG
摘要
Decoding speech imagination to directly control machines is a novel human-computer interaction approach and holds potential as a neural speech prosthetic. The aim of this work is to test whether any temporal or spectral specificity exists in speech imagination. This work further seeks to identify the optimal time window and frequency band for decoding imagined speech using a scalp electroencephalogram (EEG). Two open-access datasets were included and segmented into six frequency bands (delta, theta, alpha, beta, gamma, and full band) with three-time windows (-500–0 ms, 0–2,000 ms, and -500–2,000 ms). Dataset 1 was a five-word/phrase imagery classification experiment, while dataset 2 was a four-syllable imagery classification experiment. Machine learning was used to realize the classification and evaluation of different temporal and spectral intervals. Feature extraction and selection were implemented by independent component analysis (ICA) and mutual information (MI). The classification was realized using a support vector machine (SVM) with a radial basis function kernel. Results confirmed that the gamma band provides the highest accuracy and that the -500–0 ms pre-onset standby period contains speech-related information and is distinguishable. With a -500–2,000 ms time window and a gamma band-pass filter, the average one-versus-rest balanced accuracy achieved was 56.9% ± 9.9% and 64.7% ± 13.2% on the training and test sets, respectively, in dataset 1, while 93.5% ± 5.7% and 82.8% ± 17.3% in dataset 2.