We introduce Jina Embeddings V3, a 570-million-parameter text embedding model that excels in long-context (up to 8192 tokens) and multilingual text retrieval tasks. The model incorporates task-specific Low-Rank Adaptation (LoRA) modules for high-quality embeddings specialized for retrieval, clustering, classification, and text matching. On the MTEB benchmark, Jina Embeddings V3 outperforms other embedding models of similar size.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Jina Embeddings V3: Multilingual Text Encoder with Low-Rank Adaptations

  • Saba Sturua,
  • Isabelle Mohr,
  • Mohammad Kalim Akram,
  • Michael Günther,
  • Bo Wang,
  • Markus Krimmel,
  • Feng Wang,
  • Georgios Mastrapas,
  • Andreas Koukounas,
  • Nan Wang,
  • Han Xiao

摘要

We introduce Jina Embeddings V3, a 570-million-parameter text embedding model that excels in long-context (up to 8192 tokens) and multilingual text retrieval tasks. The model incorporates task-specific Low-Rank Adaptation (LoRA) modules for high-quality embeddings specialized for retrieval, clustering, classification, and text matching. On the MTEB benchmark, Jina Embeddings V3 outperforms other embedding models of similar size.