Design and Analysis of a Volatile Organic Compounds Mixture Dataset for Indoor Air Quality Monitoring
摘要
Machine learning models are widely utilized to improve the precision and sensitivity of intelligent electronic noses designed to detect volatile organic compound (VOC) mixtures. It is well known that the quality of a dataset, including its accuracy, volume, and variety, is crucial for achieving a good machine-learning model. However, recent studies on VOCs mixtures have typically used limited datasets, often consisting of only two gases. Moreover, many publicly available datasets lack quality control assessments of collected data. In this paper, a procedure for collecting a VOCs mixture dataset (including Ethanol, Acetone, and Methanol) is presented. The study then focused on designing suitable array sensors using the Nano SnO2 structure to ensure the accuracy and sensitivity to the targeted gas mixture. Finally, the dataset was analyzed to assess its quality using statistical tools, such as Westgard’s rules, and to understand its characteristics. The dataset contained 43,856 rows and 17 columns comprising 135 VOCs mixtures. This dataset serves as a robust basis for further studies on indoor air quality monitoring, which is crucial for pollution control and long-term health protection, thus significantly advancing sustainable development efforts.