Hierarchical Prefixes for Long Document Representations
摘要
First, most architecture-producing contextual embeddings rely on the self-attention operation with a time/space complexity being quadratic with respect to the context length. Second, embedding long contexts on one point in the embedding space limits the information we can extract/retrieve from. Third, Because of the polysemic nature of words (consequently their embeddings), higher level (e.g. a sentence) should require a specific representation.