DiscoFuse
收藏资源简介:
DiscoFuse 是通过在两个语料库上应用基于规则的拆分方法创建的 - 体育文章从网络和维基百科爬取。详细见论文 数据集生成过程和评估的描述。 DiscoFuse 有两部分,分别来自体育文章和维基百科的 44,177,443 和 16,642,323 个示例。 对于每个部分,提供随机拆分来训练(98% 的示例)、开发(1%)和测试(1%)集。此外,由于原始数据分布高度倾斜(详见论文),因此还提供了每个部分的平衡版本。
DiscoFuse was developed by applying a rule-based splitting method to two corpora: sports articles crawled from the web and Wikipedia. For detailed accounts of the dataset generation process and evaluation procedures, please refer to the accompanying research paper. DiscoFuse contains two subsets, with 44,177,443 and 16,642,323 examples sourced from the crawled web sports articles and Wikipedia, respectively. For each subset, random splits are implemented to partition the data into training (98% of the subset's examples), development (1%) and test (1%) sets. Additionally, given the highly skewed distribution of the original dataset (further details are provided in the paper), balanced versions of each subset are also made available.




