遇见数据集

jansowa/en-pl-qwen3-embeddings-metricx-scores

收藏
Hugging Face2026-05-02 更新2026-05-31 收录
官方服务:

资源简介:

该数据集包含英语-波兰语平行句对,并带有预计算的英语句子嵌入,这些嵌入由Qwen/Qwen3-Embedding-4B模型生成。源数据使用google/metricx-24-hybrid-large-v2p6模型进行评分,仅保留评分<=7.0的样本,因此这是一个经过质量过滤的高质量子集,而非完整的原始平行语料库。数据集共有32,925,668条样本,每行包含四个列:english(英语句子)、non_english(对应的波兰语句子)、label(英语句子的归一化嵌入,由Qwen/Qwen3-Embedding-4B生成)和metricx_pred(英语-波兰语对的评分,越低越好)。数据集来源于多个平行句子子集,如sentence-transformers/parallel-sentences-ccmatrix、europarl等,主要用于多语言句子嵌入蒸馏任务,特别是训练波兰语或多语言学生编码器以在对齐的英语-波兰语句对上模仿英语教师嵌入。嵌入以bfloat16格式编码为uint16存储,需解码使用。源语料库可能包含噪声、重复或不完美的翻译,尽管经过过滤。

This dataset contains English-Polish parallel sentence pairs with precomputed English sentence embeddings from Qwen/Qwen3-Embedding-4B. The source data was scored with google/metricx-24-hybrid-large-v2p6, and only examples with metricx_pred <= 7.0 were retained, making it a quality-filtered subset rather than the full original parallel corpus. The dataset includes 32,925,668 examples, with columns: english (English sentence), non_english (corresponding Polish sentence), label (normalized embedding of the English sentence generated with Qwen/Qwen3-Embedding-4B), and metricx_pred (score for the English-Polish pair, where lower is better). It is sourced from multiple parallel sentence subsets such as sentence-transformers/parallel-sentences-ccmatrix, europarl, etc., and is intended for multilingual sentence embedding distillation, particularly training Polish or multilingual student encoders to mimic English teacher embeddings on aligned English-Polish sentence pairs. Embeddings are stored as bfloat16 values packed into uint16 for efficient storage. The source corpora may contain noisy, duplicated, or imperfect translations despite filtering.

提供机构:
jansowa
二维码
社区交流群
二维码
科研交流群
商业服务