遇见数据集

ComplexDataLab/chai-veracity-dry-run-20260525-dbscan-5c1

收藏
Hugging Face2026-05-26 更新2026-05-31 收录
官方服务:

资源简介:

该数据集是一个包含多个日期配置(从2024年4月27日至2024年5月20日)的声明(claim)数据集,主要用于自然语言处理任务,如文本分类、聚类或信息检索。每个数据条目包含以下特征:id(唯一标识符)、_batch和_batch_1(批次标识)、claim(声明文本内容)、cluster_id(聚类ID,用于将相似声明分组)、date(日期信息)、diverse_sample_ids(多样本ID列表,可能关联其他样本)、post_count(帖子数量,表示该声明相关的帖子数)和original_texts(原始文本列表,提供声明的原始上下文)。数据集按日期分为多个子集,每个子集仅包含训练集(train split),总共有238个示例,数据量较小,适合小规模分析或模型测试。数据可能来源于社交媒体或在线论坛,旨在支持声明验证、话题追踪或内容分析等应用。

This dataset is a claim dataset covering multiple date configurations from April 27, 2024 to May 20, 2024, primarily intended for natural language processing (NLP) tasks such as text classification, clustering, or information retrieval. Each data entry includes the following features: id (unique identifier), _batch and _batch_1 (batch identifiers), claim (the text content of the claim), cluster_id (cluster ID used to group similar claims), date (date information), diverse_sample_ids (list of diverse sample IDs that may be associated with other samples), post_count (number of posts, indicating the count of posts related to this claim), and original_texts (list of original texts that provide the original context of the claim). The dataset is divided into multiple subsets by date, with each subset only containing the training split. It has a total of 238 examples, making it a small-scale dataset suitable for small-scale analysis or model testing. The data may originate from social media or online forums, and is designed to support applications such as claim verification, topic tracking, or content analysis.

提供机构:
ComplexDataLab
二维码
社区交流群
二维码
科研交流群
商业服务