Truthfulness Stance Detection on Claim-Tweet (TSD-CT)
收藏资源简介:
Under the final threshold, TSD-CT contains 5,331 finalized claim-tweet pairs (including 269 screening and training pairs) covering 2,201 unique factual claims. The label distribution is as follows: 2,104 (39.47%) are labeled as Positive, 882 (16.57%) as Neutral/No Stance, 883 (16.54%) as Negative, 309 (5.80%) as Different Topics, and 1,153 (21.62\%) as Problematic. The claim veracity is also diverse, with 722 claims labeled as false (32.80%), 353 as pants-fire (16.04%), 340 as barely-true (15.45%), 292 as half-true (13.27%), 287 as mostly-true (13.04%), and 207 as true (9.40%). On average, each claim–tweet pair contains 34.61 tokens and 261.76 characters. Additionally, 2,169 pairs (40.71%) include at least one hyperlink, indicating external reference or contextual support. Among all the 845 topics, the dataset is dominated by discussions on coronavirus (1,573 pairs; 29.51%) and public health (870; 16.32%), followed by Donald Trump (583; 10.94%), elections (438; 8.22%), economy (390; 7.32%), health care (329; 6.17%), crime (295; 5.54%), government regulation (255; 4.78%), drugs (240; 4.50\%), science (239; 4.49\%) and so on. Fields/Columns id: Type: Integer or String Description: Unique identifier for each claim or record. claim_author: Type: String Description: The author of the claim. claim: Type: String Description: The text of the claim being analyzed. tweet: Type: String (always "REDACTED") Description: Placeholder for the tweet text, which has been redacted for privacy. screening: Type: String or Boolean Description: Indicates whether the claim has been screened. answered: Type: Boolean Description: Indicates whether the claim has been answered or fact-checked. tweet_url_title: Type: String Description: Title or description of the tweet's URL, if applicable. claim_timestamp: Type: DateTime Description: Timestamp when the claim was made. tweet_timestamp: Type: DateTime Description: Timestamp when the associated tweet was posted. tweet_id: Type: String or Integer Description: Unique identifier for the tweet. tweet_userhandle: Type: String Description: Twitter handle of the user who posted the tweet. retweet_count: Type: Integer Description: Number of retweets for the tweet. reply_count: Type: Integer Description: Number of replies to the tweet. like_count: Type: Integer Description: Number of likes for the tweet. quote_count: Type: Integer Description: Number of quote tweets for the tweet. claim_source: Type: String Description: Source of the claim (e.g., news outlet, individual, etc.). claim_verdict: Type: String Description: Verdict of the claim (e.g., true, false, misleading). factcheck_timestamp: Type: DateTime Description: Timestamp when the claim was fact-checked. claim_review_summary: Type: String Description: Summary of the claim review. claim_review: Type: String Description: Detailed review of the claim. factcheck_url: Type: String Description: URL to the fact-checking article or source. claim_tags: Type: List of Strings Description: Tags or categories associated with the claim. claimbuster_score: Type: Float Description: Score assigned by ClaimBuster, indicating the claim's importance or likelihood of being fact-checked. pair_id: Type: String or Integer Description: Identifier for paired records (e.g., claim and fact-check). factcheck_author_url: Type: String Description: URL to the profile of the fact-checking author. factcheck_post_time: Type: DateTime Description: Time when the fact-checking post was published. factcheck_author_info: Type: String Description: Information about the fact-checking author. subset: Type: String Description: Subset or category of the dataset (e.g., training, testing, validation). annotator_agreement: Type: Float or String Description: Claim-tweet pair label. One of Positive (1), Neutral/No Stance (0), Negative (-1), Different Topics (2), Problematic (3).
最终筛选定稿后的TSD-CT数据集共包含5331条经过确认的声明-推文对(其中包含269条筛查与训练对),涵盖2201条独特的事实性声明。该数据集的标签分布如下:2104条(占比39.47%)被标记为正向(Positive),882条(16.57%)为中性/无立场(Neutral/No Stance),883条(16.54%)为负向(Negative),309条(5.80%)为话题不符(Different Topics),1153条(21.62%)为存在问题(Problematic)。声明真实性的分布同样多元,其中722条(32.80%)被标记为false(虚假),353条为pants-fire(彻头彻尾虚假),340条为barely-true(几乎不实),292条为half-true(半真半假),287条为mostly-true(大部分真实),207条为true(真实)。平均每条声明-推文对包含34.61个词元(Token)与261.76个字符。此外,2169条推文对(占比40.71%)包含至少一个超链接,表明其带有外部参考或上下文支撑。在全部845个话题中,数据集的讨论主题以冠状病毒相关(1573条,占比29.51%)与公共卫生(870条,占比16.32%)为主,其次为唐纳德·特朗普(583条,占比10.94%)、选举(438条,占比8.22%)、经济(390条,占比7.32%)、医疗保健(329条,占比6.17%)、犯罪(295条,占比5.54%)、政府监管(255条,占比4.78%)、药品(240条,占比4.50%)与科学(239条,占比4.49%)等。 字段/列 id: 类型:整数或字符串 描述:每条声明或记录的唯一标识符。 claim_author: 类型:字符串 描述:声明的发布者。 claim: 类型:字符串 描述:待分析的声明文本。 tweet: 类型:字符串(固定为"REDACTED") 描述:推文文本占位符,因隐私保护已做脱敏处理。 screening: 类型:字符串或布尔值 描述:标识该声明是否已完成筛查。 answered: 类型:布尔值 描述:标识该声明是否已得到回应或事实核查。 tweet_url_title: 类型:字符串 描述:推文链接的标题或描述(如适用)。 claim_timestamp: 类型:日期时间(DateTime) 描述:声明发布的时间戳。 tweet_timestamp: 类型:日期时间(DateTime) 描述:关联推文发布的时间戳。 tweet_id: 类型:字符串或整数 描述:推文的唯一标识符。 tweet_userhandle: 类型:字符串 描述:发布该推文的用户的Twitter账号名。 retweet_count: 类型:整数 描述:该推文的转发次数。 reply_count: 类型:整数 描述:该推文的回复次数。 like_count: 类型:整数 描述:该推文的点赞次数。 quote_count: 类型:整数 描述:该推文的引用转发次数。 claim_source: 类型:字符串 描述:声明的来源(例如新闻媒体、个人等)。 claim_verdict: 类型:字符串 描述:声明的核查结论(例如true、false、misleading等)。 factcheck_timestamp: 类型:日期时间(DateTime) 描述:声明被完成事实核查的时间戳。 claim_review_summary: 类型:字符串 描述:声明核查结果的摘要。 claim_review: 类型:字符串 描述:声明的详细核查内容。 factcheck_url: 类型:字符串 描述:指向事实核查文章或来源的链接。 claim_tags: 类型:字符串列表 描述:与声明相关的标签或分类。 claimbuster_score: 类型:浮点数 描述:由ClaimBuster给出的评分,用于衡量声明的重要性或被事实核查的可能性。 pair_id: 类型:字符串或整数 描述:配对记录的标识符(例如声明与对应的事实核查内容)。 factcheck_author_url: 类型:字符串 描述:指向事实核查作者个人主页的链接。 factcheck_post_time: 类型:日期时间(DateTime) 描述:事实核查帖子的发布时间。 factcheck_author_info: 类型:字符串 描述:事实核查作者的相关信息。 subset: 类型:字符串 描述:数据集的子集或分类(例如训练集、测试集、验证集)。 annotator_agreement: 类型:浮点数或字符串 描述:声明-推文对的标签,可选值为:正向(Positive,1)、中性/无立场(Neutral/No Stance,0)、负向(Negative,-1)、话题不符(Different Topics,2)、存在问题(Problematic,3)。



