遇见数据集

ali-issa/eng_and_arab_unused_aya_datasets

收藏
Hugging Face2025-01-27 更新2025-04-12 收录
官方服务:

资源简介:

Aya数据集是一个包含多种语言的数据集,最初包含19个子数据集。由于阿拉伯语和英语数据集行数不匹配,有3个子数据集被移除,包括CNN Daily Mail、XLEL-WD和SODA。该数据集的更新日期为2025年1月27日。

The Aya dataset is a multilingual dataset that initially contained 19 subsets. Due to mismatched row counts between the Arabic and English datasets, three subsets, including CNN Daily Mail, XLEL-WD, and SODA, were removed. The dataset was last updated on January 27, 2025.

提供机构:
ali-issa
二维码
社区交流群
二维码
科研交流群
商业服务