遇见数据集

Replication Data for: Experiments with Learning Graphical Models on Text

收藏
DataONE2018-04-16 更新2024-06-25 收录
官方服务:

资源简介:

This dataset contains the binarised document collections in .arff format to replicate the Experiments with Learning Graphical Models on Text. Every zip file contains a train a test folder. In each folder, there also are two set of files corresponding to the version with 500 words or 2000 words vocabulary. For every set, there are several files corresponding to the different training/testing splits and to the 5 repetitions per experiment. The test folders also contain files that start with WORD2PREDICT, which contain the words in the test set to prediction for the prediction task.

本数据集包含以属性-关系文件格式(.arff)存储的二值化文档集,用于复现《基于文本的图模型学习实验》(Experiments with Learning Graphical Models on Text)中的相关实验。每个压缩包均包含训练集与测试集文件夹,每个文件夹内均设有两组文件,分别对应词汇量为500词与2000词的数据集版本。针对每组词汇量设置,存在若干文件,分别对应不同的训练/测试划分方案以及每项实验的5次重复。测试集文件夹中还包含以WORD2PREDICT开头的文件,这类文件存储了测试集中用于预测任务的目标词汇。

创建时间:
2023-11-22
二维码
社区交流群
二维码
科研交流群
商业服务