官方服务:
资源简介:
Database
应用场景:
相关数据集
Polish-Kashubian parallel translation corpus
The data set contains about 120,000 Polish words and sentences and their translations into Kashubian. It was created using two types of sources. The first one is the online dictionaries: kaszebe.org
DataCite Commons2024-09-30 更新130
Translation-EnKo/arxiv-translation-result-6.9k-0909
--- dataset_info: features: - name: eng dtype: string - name: gukbap dtype: string - name: exaone_S dtype: string - name: exaone_L dtype: string splits: - name: train
Hugging Face2024-09-09 更新60
SIMPITIKI_GITHUB_意大利语文本简化语料库数据
本数据集为意大利语文本简化语料库SIMPITIKI,包含两组简化文本对:第一组通过半自动方式从意大利语维基百科获取,第二组从行政领域文档中逐句手动标注。数据集仅含一个XML格式文件,无训练测试、数据标签或原始处理数据的划分。 GitHub仓库dhfbk/simpitiki
海数据40
Parallel corpus EN-SL RSDO4 1.0
The RSDO4 parallel corpus of English-Slovene and Slovene-English translation pairs was collected as part of work package 4 of the Slovene in the Digital Environment project. It...
B2FIND60
frontier-science-multilingual
该数据集包含德语(deu)和法语(fra)两个版本,每个版本包含159个测试样本。数据集的主要字段包括:id(唯一标识符)、benchmark(基准名称)、subset(子集名称)、subject(主题)、task_group_id(任务组ID)、problem(问题描述)、answer(答案)、flag_for_review(是否需要审核标记)、review_reason(审核原因)、targe
Hugging Face2026-03-25 更新50



