遇见数据集

Multi-Label Arabic Dataset

收藏
Mendeley Data2026-04-18 收录
官方服务:

资源简介:

The dataset is a collection of hierarchical multi-label Arabic texts, related to the Islamic field. It consists of 26,470 instances distributed over 578 labels ordered in a hierarchy. After ranking the features using (BR-Chi-Square) feature selection method, a different number of the high-ranking features are selected for evaluation purposes which are 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000 features. The processed version of the dataset with all aforementioned features sets is available in the ARFF file format suitable for MULAN multi-label classification tool, along with the XML file format that defines the hierarchical structure of the labels.

本数据集为一批与伊斯兰领域相关的层级式多标签(hierarchical multi-label)阿拉伯语文本集合,共包含26470条样本,分布于578个按层级结构组织的标签之下。我们采用BR-Chi-Square特征选择方法对特征进行排序后,选取了不同数量的高排名特征用于模型评估,选取的特征数量分别为1000、2000、3000、4000、5000、6000、7000与8000。本数据集经过处理后的版本(包含上述全部特征子集),可通过适配MULAN多标签分类工具的ARFF文件格式获取,同时附带用于定义标签层级结构的XML文件。

创建时间:
2020-09-02
二维码
社区交流群
二维码
科研交流群
商业服务