CLIPPER
收藏资源简介:
CLIPPER数据集是由马里兰大学帕克分校和麻省理工学院的研究人员创建的,包含19K条关于公共领域小说书籍的合成索赔。该数据集通过两阶段的压缩方法生成:首先将书籍压缩成章节概要和书籍摘要,然后基于这些压缩表征生成真实/虚假的索赔和相应的思维链。数据集旨在用于叙事索赔验证任务,以解决长文本情境下的推理问题。
The CLIPPER dataset was created by researchers from the University of Maryland, College Park and the Massachusetts Institute of Technology. It contains 19K synthetic claims about public-domain fictional books. This dataset is generated via a two-stage compression pipeline: first, books are compressed into chapter summaries and book abstracts, then authentic and fake claims along with their corresponding chain-of-thoughts are generated based on these compressed representations. The dataset is intended for the narrative claim verification task, aiming to address reasoning problems in long-text scenarios.




