遇见数据集

stephantulkens/pubmedqa-query-mxbai-pooled

收藏
Hugging Face2025-10-13 更新2025-10-25 收录
官方服务:

资源简介:

Embedpress: mixedbread large on the PubmedQA queries dataset 是一个使用 Mixedbread AI 的 mixedbread-ai/mxbai-embed-large-v1 模型嵌入的 PubmedQA 数据集的查询部分。每个文档取前 510 个标记(模型的最大长度 -2 个特殊标记),进行嵌入,未使用任何指令。由于该模型使用 Matryoshka Representation Learning 进行训练,这些嵌入可以安全地截断。这些数据主要适用于大规模知识蒸馏。

Embedpress: mixedbread large on the PubmedQA queries dataset is the query portion of the PubmedQA dataset, embedded with Mixedbread AIs mixedbread-ai/mxbai-embed-large-v1. For each document, we take the first 510 tokens (the models max length -2 special tokens), and embed it, not using any instructions. Because the model was trained using Matryoshka Representation Learning, these embeddings can safely be truncated. These are mainly useful for large-scale knowledge distillation.

提供机构:
stephantulkens
二维码
社区交流群
二维码
科研交流群
商业服务