Data Science: Machine Learning and Multivariate Analysis in Learning Styles
摘要
Data science is responsible for the analysis, interpretation and prediction of simple data to generate significant knowledge, it applies to any area that produces data, for example: sales, finance, production, health, education, etc. Data science in education uses machine learning techniques and multivariate analysis. Currently, Educational Data Mining or EDM is spoken, which unites the areas of education, Big data, and data science to improve learning. In this study, data science models are analyzed and compared to identify patterns of behavior, similarity, and anomalous data from students to generate new unknown knowledge and characterize the profile of students according to learning styles such as those proposed by Felder and Silverman in addition to relating these styles to development by competencies. The results show that the most efficient methods are: Clustering and its k-means algorithm with which group characteristics are obtained, decision trees with its ID3 algorithm that through the gain of information a better classification is obtained and the PCA mathematical model or Principal Component Analysis that by its properties of variability analysis and dimension reduction allows to obtain more information from data with noise or outliers. A characterization of the data to be processed is also carried out, classifying it into profile data, class data and test data. These analyzed models will be implemented in a next phase in students of the Software Development career of the ISMAC Institute, which allows obtaining promising results to predict student learning styles and improve said learning process.