遇见数据集

cp500/multilingual-automotive-sparse

收藏
Hugging Face2026-04-27 更新2026-05-03 收录
官方服务:

资源简介:

Multilingual Automotive Sparse Retrieval Corpus是一个合成的多语言训练语料库,专门用于微调基于SPLADE家族的多语言稀疏检索模型,内容涵盖汽车、供应链和地缘政治领域。该数据集旨在教授模型跨语言对齐能力,例如使日语查询能够检索到正确的英文段落,反之亦然。它包含三种语言(英语、日语、韩语)的数据,文件包括原始概念记录、训练三元组和评估对。数据通过Anthropic Claude Haiku 4.5合成生成,基于2500个概念种子,并涵盖多个类别如能源商品、制药生物技术等。数据集适用于微调稀疏检索模型、跨语言检索研究以及汽车智能领域的FLOPS正则化实验,但所有内容均为合成,可能不适用于事实性问答,且领域较窄。

The Multilingual Automotive Sparse Retrieval Corpus is a synthetic training corpus designed for fine-tuning multilingual sparse retrieval models (SPLADE-family) on automotive, supply-chain, and geopolitics content. It aims to teach models cross-lingual alignment, enabling queries in one language (e.g., Japanese) to retrieve relevant passages in another (e.g., English), and vice versa. The dataset includes data in three languages (English, Japanese, Korean) and consists of files for raw concept records, training triplets, and evaluation pairs. Generated using Anthropic Claude Haiku 4.5 based on 2500 concept seeds, it covers categories such as energy commodities, pharma biotech, and more. Intended uses include fine-tuning sparse retrieval models, cross-lingual retrieval research, and FLOPS-regularization experiments in the automotive-intelligence domain, though all content is synthetic and domain-specific.

提供机构:
cp500
二维码
社区交流群
二维码
科研交流群
商业服务