数据链接:
官方服务:
资源简介:
Dutch Information Retrieval dataset build with the wikIR tool.
应用场景:
提供机构:
FREJ Jibril创建时间:
2020-01-03
相关数据集
M3IT
# Dataset Card for M3IT Project Page: [M3IT](https://m3-it.github.io/) ## Dataset Description - **Homepage: https://huggingface.co/datasets/MMInstruction/M3IT** - **Repository: https://huggingfa
魔搭社区2026-07-08 更新870
Dataset for "Cross-lingual Knowledge Projection Using Machine Translation and Target-side Knowledge Base Completion"
This dataset contains 17,068 facts (triples) of Japanese commonsense knowledge obtained by translating English counterparts. Data Credit This work includes data from ConceptNet 5, which w
DataCite Commons2020-08-28 更新70
bifrost-translation-source-classifier-dataset
Bifrost翻译源分类器数据集是一个用于训练翻译源分类器的数据集,包含英文文本及其原始翻译来源语言的标签,以及作为控制类的原生英文文本。所有文本均为英文,标签指示文本的原始翻译来源语言(或en表示原生英文)。该数据集旨在帮助分类器学习检测原始语言的文化和风格痕迹。数据集包含180种语言,每种语言有10,000个训练样本、1,000个验证样本和1,000个测试样本,总计2,160,000个样本。数
Hugging Face2026-04-27 更新30
FAUST 0.5
Syntactic (including deep-syntactic - tectogrammatical) annotation of user-generated noisy sentences. The annotation was made on Czech-English and English-Czech Faust Dev/Test...
B2FIND90
yywwrr/mmarco_dutch_500k
--- dataset_info: features: - name: text dtype: string splits: - name: train num_bytes: 185513219 num_examples: 480000 - name: dev num_bytes: 73612 num_examples: 200 -
Hugging Face2025-12-15 更新30



