Detecting Confusing Drug Names Based on the Phonetic Characteristics of Mel-Frequency Cepstral Coefficient and Evolutionary Computation
摘要
Errors in drug administration due to confusing names, known as Look-Alike, Sound-Alike (LASA) errors, are common and can lead to serious consequences such as overdose toxicity, adverse effects, and even death. In this work, the incorporation of Mel-Frequency Cepstral Coefficients (MFCC) as a phonetic measure using the audio signal directly is proposed. MFCC are a feature extraction technique designed to capture the sound human perception. Additionally, techniques such as Genetic Algorithm (GA) and Differential Evolution (DE) were employed to optimize the capabilities of MFCC for identifying LASA drugs in Spanish. Using a ground truth list of 586 Spanish drugs names with 784 drug pairs reported as LASA errors, the audios were generated and the MFCCs were calculated and summarized into their descriptive statistics, for each drug pair, then the Euclidean distance between statistics of the MFCC per drug pair was calculated. The GA selected 41 most relevant MFCC statistics from a total of 80, which improved LASA pair identification performance, reaching a maximum F1-Score of 0.4030. Subsequently, DE weightily optimized these 41 features, further increasing the identification power to reach an F1-Score of 0.5313. \( DistMFC{C}_{17_{STD}} \) , \( DistMFC{C}_{20_{MEAN}} \) , and \( DistMFC{C}_{8_{STD}} \) showed better performance in describing LASA drugs, while all 41 features showed significant difference between LASA and non-confused drugs. MFCCs are useful for identifying confusing LASA drug pairs in Spanish, outperforming phonetic metrics such as Soundex or Phonix. One of the advantages of using MFCCs instead of conventional measures is their ability to generate a tie-free ranking by employing similarity measures with decimal values.