ComplexDataLab/chai-veracity-dry-run-20260521-dbscan-diverse
收藏资源简介:
该数据集是一个包含多个日期配置(从20260427到20260519)的文本数据集合,主要用于处理与声明(claim)相关的信息。每个数据条目包括唯一标识符(id)、声明内容(claim)、聚类标识符(cluster_id)、日期(date)、多样化样本ID列表(diverse_sample_ids)、帖子数量(post_count)以及原始文本列表(original_texts)。数据集可能用于自然语言处理任务,如声明分类、聚类分析或文本生成,但具体用途未在README中明确说明。数据以训练集(train)形式提供,总示例数约为1102条(默认配置),其他每日配置的示例数从30到68不等。
This dataset is a collection of text data with multiple date configurations (from 20260427 to 20260519), primarily focused on handling claim-related information. Each data entry includes a unique identifier (id), claim content (claim), cluster identifier (cluster_id), date (date), a list of diverse sample IDs (diverse_sample_ids), post count (post_count), and a list of original texts (original_texts). The dataset may be used for natural language processing tasks such as claim classification, cluster analysis, or text generation, but its specific purpose is not explicitly stated in the README. The data is provided in a training set (train) format, with a total of approximately 1102 examples in the default configuration, and other daily configurations ranging from 30 to 68 examples each.



