遇见数据集

dbpedia-14

收藏
OpenXLab2026-04-18 收录
官方服务:

资源简介:

The DBpedia ontology classification dataset is constructed by picking 14 non-overlapping classes from DBpedia 2014. They are listed in classes.txt. From each of thse 14 ontology classes, we randomly choose 40,000 training samples and 5,000 testing samples. Therefore, the total size of the training dataset is 560,000 and testing dataset 70,000. There are 3 columns in the dataset (same for train and test splits), corresponding to class index (1 to 14), title and content. The title and content are escaped using double quotes ("), and any internal double quote is escaped by 2 double quotes (""). There are no new lines in title or content.

提供机构:
OpenDataLab
创建时间:
2023-12-07
二维码
社区交流群
二维码
科研交流群
商业服务