Integration of the Peruvian Citizen’s Public Information by Applying Web Scraping Under SCRUM Methodology
摘要
The Digital Platform of the Peruvian State is mainly composed of seven websites. To obtain complete information about a citizen, information must be extracted from each website and integrated manually, which can take more than 3 min. The objective is to centralize the public information coming from the seven websites through a single web platform by applying web scraping. The methodology to implement the web scraping technique, the Selenium tool was used to simulate the information query process by a user entering an ID number, and the web platform was developed based on the Scrum methodology divided into three Sprints. As a result, users can visualize with a simple query the public information of a citizen stored and available on different websites, and the average time of information search of the citizen was reduced from 136 to 24 s. In conclusion, it can be affirmed that the use of web scraping can extract from different governmental websites the information of a citizen with a simple query in a fast and complete way.