PreSL: Document Level Pre-training of Scientific Literature via Self-supervised Learning
摘要
Every year, numerous research papers are published, introducing a vast number of scientific ideas and posing significant challenges for researchers to follow and utilize relevant studies. Additionally, scientific literature, as a unique type of text, not only contains complex semantic content but also features strong academic network associations, such as citation information. These associations are crucial for accurately modeling scientific texts. Overall, scientific literature presents three major challenges: large volume, complex semantics, and intricate associations. To address these challenges, we introduce PreSL, a pre-training model specifically designed for handling scientific literature. PreSL leverages multi-task learning and self-supervised learning techniques to effectively handle the large volume, complex semantics, and intricate associations inherent in scientific texts. The effectiveness of our proposed model has been demonstrated through citation prediction and visual analysis.