Multi-Label Arabic Dataset
收藏资源简介:
The dataset is a collection of hierarchical multi-label Arabic texts, related to the Islamic field. It consists of 26,470 instances distributed over 578 labels ordered in a hierarchy. After ranking the features using (BR-Chi-Square) feature selection method, a different number of the high-ranking features are selected for evaluation purposes which are 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000 features. The processed version of the dataset with all aforementioned features sets is available in the ARFF file format suitable for MULAN multi-label classification tool, along with the XML file format that defines the hierarchical structure of the labels.
本数据集为层级式多标签(hierarchical multi-label)阿拉伯语文本集合,关联伊斯兰研究领域。数据集共包含26470条样本,分布于578个层级化排列的标签类别中。研究人员采用(BR-Chi-Square)特征选择方法对特征进行排序后,选取不同数量的高排名特征用于模型评估,选取的特征数量分别为1000、2000、3000、4000、5000、6000、7000、8000。本数据集的预处理版本(包含上述全部特征子集)以ARFF文件格式存储,可适配MULAN多标签分类工具;同时配套提供用于定义标签层级结构的XML文件格式。



