官方服务:
资源简介:
Source: Mark the Evangelist
来源:福音书作者马可(Mark the Evangelist)
应用场景:
相关数据集
SAMER阿拉伯语文本简化语料库
SAMER阿拉伯语文本简化语料库是由纽约大学阿布扎比分校计算语言学建模实验室创建的,旨在为学龄学习者提供文本简化资源。该数据集包含从15本公开可用的阿拉伯语小说中选取的约15.9万字文本,这些小说大多在1865年至1955年间出版。数据集不仅包括文档和单词级别的可读性标注,还为每个文本提供了两种简化版本,针对不同可读性水平的学习者。创建过程中遵循严格的指导原则以确保标注质量。该数据集的应用领域包括
arXiv2024-04-29 更新300
deokhk/ar_wiki_sentences_1000000
--- dataset_info: features: - name: sentence dtype: string splits: - name: train num_bytes: 213292535 num_examples: 1000000 - name: dev num_bytes: 214085 num_examples: 10
Hugging Face2024-01-12 更新90
Arabic News Headlines
This dataset is offered as .csv and is part of 3 files which are:- File 1: has all 1699 arabic news headlines colllected with the corresponding emotion classification that 3 annotators agreed on with
IEEE2020-11-27 更新80
irds/clueweb09_ar
`clueweb09/ar`数据集由`ir-datasets`包提供,包含29,192,662个文档。这些文档包括文档ID、URL、日期、HTTP头、正文和正文内容类型等信息。用户可以使用`datasets`库中的`load_dataset`函数加载该数据集,该函数将下载数据集或提供访问指南,并将数据转换为🤗 Dataset格式。
Hugging Face2023-01-05 更新150
asas-ai/APGC_v2
--- task_categories: - text-classification language: - ar license: cc-by-nc-4.0 pretty_name: APGC_v2 size_categories: - 1K<n<10K tags: - Gender Identification --- # Dataset Card for "APGC_v2: The Arab
Hugging Face2024-05-05 更新90



