On the recoverability of cognitive levels in educational text: a multi-task transformer analysis
摘要
Artificial intelligence systems have demonstrated strong performance in educational reasoning and classification tasks; however, the extent to which transformer-based representations encode structured cognitive abstractions remains insufficiently understood. This study investigates the recoverability of cognitive levels from educational question–answer text using a Cognitive-Conditioned Multi-Task Transformer (CC-MTT) framework. To ensure methodological rigor, a leakage-controlled dataset reconstruction pipeline was developed using question-level partitioning and majority-vote label stabilization. The proposed framework jointly models reasoning and cognitive classification, enabling analysis of predictive performance, representational structure, and stability across multiple random seed initializations. Experimental results show that fine-grained Bloom taxonomy levels remain only weakly recoverable from text-only representations, achieving performance close to random baseline under leakage-controlled evaluation. In contrast, hierarchical abstraction (higher-order vs. lower-order thinking) consistently achieved higher predictive recoverability across multiple runs, although quantitative geometry analysis revealed limited cluster separability in the learned embedding space. Statistical analysis further indicated a consistent trend toward improved binary abstraction performance compared with fine-grained categorization, despite the absence of strong statistical significance under Wilcoxon testing. These findings suggest that coarse-grained cognitive abstraction may be partially encoded in transformer representations, while fine-grained pedagogical categories remain structurally difficult to recover from textual representations alone. The study provides methodological and analytical insights for the development of more reliable AI-driven educational assessment systems.