Multi-environment Audio Dataset Using RPi-Based Sound Logger
摘要
Rapid urbanization has led us being surrounded by many machines in our homes, workplaces, and neighborhood. These machines create assorted noises that have disturbing ramifications on the mental and physical health of human beings. A proper analysis of these sounds is essential for medical professionals, especially ENT specialists to study the noise effects on patients. Though AI (Artificial Intelligence) and ML (Machine Learning) technologies have presented complex algorithms to aid research and development in sound analysis and categorization, the paucity of appropriately labeled sound datasets limits the extent of the accuracy of these models. To bridge this gap a database is proposed with a collection of audio recordings of systems operating in ten diverse environments ranging from domestic appliances, automobiles, HVAC systems, industrial machines, construction equipment, etc. encountered by a person on a day-to-day basis. A portable sound logging device is developed using Raspberry Pi 3B+ and INMP441 MEMS microphone breakout board. I2Smic, PyAudio, Librosa, and several python libraries are used in system development. Labeling or annotation of the collected dataset is done manually by experts. All audio files are sampled at 44.1 kHz, single-channel, 16-bit audio wav file format. The database can be used for speech recognition, acoustic classification, audio anomaly detection, and the development of diagnostic systems in otology studies and allied medical research.