遇见数据集

mahiyama/mqa-ja

收藏
Hugging Face2026-05-21 更新2026-05-31 收录
官方服务:

资源简介:

该数据集基于日语QA数据集hpprc/mqa-ja,通过Hard Negative Mining和Cross-Encoder蒸馏评分技术处理,用于日语检索模型学习。提供三种格式:pairs、triplets和n-tuples,以及n-tuples的蒸馏评分排序top-K子集(100k、250k、500k、1m)。可用于密集检索器、Cross-Encoder、SPLADE等日语检索模型的学习和KL散度蒸馏。

This dataset is based on the Japanese QA dataset hpprc/mqa-ja, enhanced with Hard Negative Mining and Cross-Encoder distillation scoring for Japanese retrieval learning. It provides three formats: pairs, triplets, and n-tuples, along with top-K subsets (100k, 250k, 500k, 1m) of n-tuples sorted by learning value using distillation scores. It can be used for training Japanese retrieval models such as Dense Retriever, Cross-Encoder, and SPLADE, as well as for KL divergence distillation.

提供机构:
mahiyama
二维码
社区交流群
二维码
科研交流群
商业服务