A self-training framework using pre-trained language models to detect and classify technology adoption in protest event analysis
摘要
This paper presents a self-training computational framework that extends protest event analysis (PEA) beyond extracting event characteristics to capturing the contextual granularity of protest. Using the adoption of digital technology as an illustrative case, we show how specific subdimensions of protest can be systematically detected from large-scale news data. Existing PEA approaches typically rely on conventional machine learning techniques focused on event identification, which limits their capacity to capture contextual nuance, and depend on manual annotation. We conceptualize the adoption of digital technology as a distinct subdimension of protest repertoire and operationalize it through the Technology in Movement (TiM) framework, an end-to-end methodological pipeline designed to reconstruct protest datasets, enrich contextual information, and identify technology-mediated protest dynamics. Methodologically, this framework shifts protest event classification from simple named-entity recognition to a syntactic graph-based pattern-matching problem in which graph convolutional networks (GCNs) model how technologies are embedded in protest actions and relations. By integrating self-training with pre-trained language models and syntactic GCNs, the framework enables context-sensitive learning while reducing reliance on manual annotation. The findings enrich the field of computational PEA through the development of a transferable framework for reconstructing protest databases and a demonstration of the analytical value of LLM–GCN integration.