First, most architecture-producing contextual embeddings rely on the self-attention operation with a time/space complexity being quadratic with respect to the context length. Second, embedding long contexts on one point in the embedding space limits the information we can extract/retrieve from. Third, Because of the polysemic nature of words (consequently their embeddings), higher level (e.g. a sentence) should require a specific representation.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Hierarchical Prefixes for Long Document Representations

  • Iskandar Boucharenc

摘要

First, most architecture-producing contextual embeddings rely on the self-attention operation with a time/space complexity being quadratic with respect to the context length. Second, embedding long contexts on one point in the embedding space limits the information we can extract/retrieve from. Third, Because of the polysemic nature of words (consequently their embeddings), higher level (e.g. a sentence) should require a specific representation.