错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Generative Byte-Level Models for Restoring Spaces, Punctuation, and Capitalization in Multiple Languages

  • Laurence Dyer,
  • Anthony Hughes,
  • Burcu Can

摘要

Restoration of textual features such as spaces, punctuation, and capitalization is a fundamental NLP task that has applications in post-processing of automatic speech recognition (ASR) outputs, hashtags on social media, and other types of unprocessed or noisy data. In this work, we build upon the token-free approach adopted in previous work on this topic by proposing a generative method that leverages pre-trained byte-level transformer models. The newly proposed models are shown to outperform previously proposed models across a range of languages—English, Japanese, and Gujarati—from distinct language families and possessing distinct linguistic features. Overall F-score for the English dataset is 96%, showing that our new models outperform not only previous token-free models (F-score 90%) but also a high-performing pipeline that employs word embeddings (F-score 95%). Furthermore, the newly proposed model is found to be more effective than previous models at restoring mid-token features, correctly reflecting 66% of tokens containing mid-token features in the English dataset compared to 47% in previous work.