官方服务:
资源简介:
Training data for ai
应用场景:
创建时间:
2024-08-04
相关数据集
google/smol
Smol数据集专注于翻译任务,包括多种语言对,数据被分为训练集。该数据集在CC BY 4.0许可证下发布,大小在10,000到100,000条记录之间。README文件列出了各种语言代码,表示涉及的翻译语言对。数据集在不同的配置中设置,每个配置都有特定的语言对,并提供了每个配置的数据文件路径。
Hugging Face2025-03-03 更新180
Fulfulde–French and Fulfulde–English Parallel Corpus (Adamawa Fulani) for Machine Translation
This dataset consists of parallel sentence pairs in Fulfulde–French and Fulfulde–English, developed for training and evaluating machine translation systems and other natural language processing (NLP)
Zenodo2026-04-23 更新50
saillab/alpaca-igbo-cleaned
--- language: - ig pretty_name: Igbo alpaca-52k size_categories: - 100K<n<1M --- This repository contains the dataset used for the TaCo paper. Please refer to the paper for more details: [OpenRevie
Hugging Face2024-09-20 更新40
550万组土耳其语-英文平行语料数据
550万组土耳其语-英语平行互译语料,数据存储格式为txt文档,内容覆盖多个领域。已进行数据清洗脱敏质检,可作为文本类数据分析的基础语料库,用于机器翻译等领域。
国家数据集管理服务平台2026-04-28 更新50
Fulfulde–French and Fulfulde–English Parallel Corpus (Adamawa Fulani) for Machine Translation
This dataset consists of parallel sentence pairs in Fulfulde–French and Fulfulde–English, developed for training and evaluating machine translation systems and other natural language processing (NLP)
Zenodo2026-04-23 更新40



