数据链接:
官方服务:
资源简介:
Data and code for project investigating bias in machine translations
应用场景:
创建时间:
2023-07-06
相关数据集
reddit_dataset_82
该数据集是Bittensor Subnet 13去中心化网络的一部分,包含预处理的Reddit数据。数据由网络矿工持续更新,提供Reddit内容的实时流,适用于各种分析和机器学习任务。数据集包括文本、标签、数据类型、社区名称、日期时间、用户名编码和URL编码等字段。主要语言为英语,但可以是多语言的。该数据集在MIT许可下发布,并受Reddit使用条款的约束。
Hugging Face2024-12-05 更新320
jamesagilesoda/ko-corpus-cleaned-12653878
--- dataset_info: features: - name: text dtype: string splits: - name: clean num_bytes: 100325969043 num_examples: 12653878 - name: noisy num_bytes: 144185494007 num_exam
Hugging Face2024-01-25 更新60
AlvaroCentellas/typst-instruct
--- dataset_info: features: - name: prompt dtype: string - name: completion dtype: string splits: - name: train num_bytes: 3848595 num_examples: 1006 download_size: 1783309
Hugging Face2025-12-11 更新40
RyanZZZZZ/w5_dev_all_input_bhc
--- dataset_info: features: - name: input dtype: string - name: label dtype: string splits: - name: train num_bytes: 220273513 num_examples: 14719 download_size: 117067450
Hugging Face2024-04-12 更新70
vlsp-2023-vllm/mmlu
--- configs: - config_name: default data_files: - split: validation path: data/validation-* - split: dev path: data/dev-* - split: test path: data/test-* dataset_info: features:
Hugging Face2023-09-30 更新60



