timaeus/lang5_probes
收藏资源简介:
该数据集包含多个CSV文件,每个文件代表一个特定的探测任务。每个CSV文件包含`prompt`(提示)、`prompt_len`(提示长度)和`target`(目标)列。目标通常是0/1整数,除非另有说明。大多数数据集是平衡的(50/50)。数据集涵盖多个领域,包括历史人物性别判断(如“Margaret of Clisson”)、历史人物是否美国人、新闻标题是否关于奥巴马、科学问题答案是否正确、Reddit帖子是否短文本、行为是否合理、文本是否以“high school”结尾、电影评论情感分析、短信是否为垃圾邮件、美国地点是否属于洛杉矶时区、算术问题正确答案是否为A、新闻是否属于政治类、医学摘要是否关于消化系统疾病、推文是否表达悲伤情绪、代码是否为Python等。
The dataset consists of multiple CSV files, each representing a specific probe task. Each CSV file contains `prompt`, `prompt_len`, and `target` columns. Targets are typically 0/1 integers unless noted otherwise. Most datasets are balanced (50/50). The datasets cover various domains, including determining the gender of historical figures (e.g., Margaret of Clisson), whether a historical figure is American, whether a news headline is about Obama, whether a science questions answer is correct, whether a Reddit post is short, whether an action is justified, whether a text ends with the high school bigram, sentiment analysis of movie reviews, whether an SMS is spam, whether a US location is in the Los Angeles timezone, whether the correct answer to an arithmetic question is A, whether a news article is about politics, whether a medical abstract is about digestive system diseases, whether a tweet expresses sadness, and whether a code snippet is in Python.




