<p>This paper introduces the Sanskrit Sembank (<span>ssb</span>), a comprehensive lexical semantic resource integrated within the Digital Corpus of Sanskrit. The <span>ssb</span> combines lexicographic data from Sanskrit dictionaries with synsets derived from the Princeton WordNet, and offers lexical semantic annotations of more than 600,000 words in context across both Vedic and Classical Sanskrit texts. We discuss the methodological challenges in adapting WordNet’s conceptual framework to the vocabulary of Sanskrit, particularly in religious and scientific domains. We further present evaluation results from word sense disambiguation experiments, achieving F1 scores of up to 86.7% without using an LLM. The paper also examines the complementary relationship between the <span>ssb</span> and the Sanskrit WordNet project, highlighting their distinct and complementary approaches to lexical semantic annotation.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

The Sanskrit Sembank

  • Oliver Hellwig,
  • Erica Biagetti

摘要

This paper introduces the Sanskrit Sembank (ssb), a comprehensive lexical semantic resource integrated within the Digital Corpus of Sanskrit. The ssb combines lexicographic data from Sanskrit dictionaries with synsets derived from the Princeton WordNet, and offers lexical semantic annotations of more than 600,000 words in context across both Vedic and Classical Sanskrit texts. We discuss the methodological challenges in adapting WordNet’s conceptual framework to the vocabulary of Sanskrit, particularly in religious and scientific domains. We further present evaluation results from word sense disambiguation experiments, achieving F1 scores of up to 86.7% without using an LLM. The paper also examines the complementary relationship between the ssb and the Sanskrit WordNet project, highlighting their distinct and complementary approaches to lexical semantic annotation.