stephantulkens/pubmedqa-query-mxbai-pooled
收藏官方服务:
资源简介:
Embedpress: mixedbread large on the PubmedQA queries dataset 是一个使用 Mixedbread AI 的 mixedbread-ai/mxbai-embed-large-v1 模型嵌入的 PubmedQA 数据集的查询部分。每个文档取前 510 个标记(模型的最大长度 -2 个特殊标记),进行嵌入,未使用任何指令。由于该模型使用 Matryoshka Representation Learning 进行训练,这些嵌入可以安全地截断。这些数据主要适用于大规模知识蒸馏。
Embedpress: mixedbread large on the PubmedQA queries dataset is the query portion of the PubmedQA dataset, embedded with Mixedbread AIs mixedbread-ai/mxbai-embed-large-v1. For each document, we take the first 510 tokens (the models max length -2 special tokens), and embed it, not using any instructions. Because the model was trained using Matryoshka Representation Learning, these embeddings can safely be truncated. These are mainly useful for large-scale knowledge distillation.
提供机构:
stephantulkens


