BSC-LT/cobie_sst2
收藏资源简介:
该数据集是基于原始SST-2数据集修改而来,用于评估大语言模型(LLM)的认知偏差。数据集包含25,000个实例,其余实例用作少样本示例。每个实例都带有所有可能的不平衡4样本分布。为了增加任务的复杂性,还在第一个和最后两个示例之间引入了一个额外的中性示例。数据集的主要字段包括原始句子ID、测试句子、情感标签、少样本分布、示例句子及其情感标签等。数据集支持情感分类任务,并由巴塞罗那超级计算中心语言技术部门策划。
This dataset is a modification of the original SST-2 dataset for LLM cognitive bias evaluation. The dataset contains 25,000 instances, with the remaining ones serving as few-shot examples. Each instance is prompted with all possible unbalanced 4-shot distributions. To increase the original task complexity, an additional neutral example is introduced between the first and last two examples. The main fields of the dataset include the original sentence ID, test sentence, sentiment label, few-shot distribution, example sentences and their sentiment labels, etc. The dataset supports sentiment classification tasks and is curated by the Language Technologies Unit at the Barcelona Supercomputing Center.



