Anno-lexical
收藏资源简介:
Anno-lexical数据集由日本东京国立信息学研究所和德国哥廷根大学联合创建,是一个用于媒体偏见分类的大型数据集,包含48330条合成标注的句子。数据集通过使用大型语言模型(LLMs)自动化标注过程,减少了人工标注的成本和时间,同时保持了数据的高质量。数据集的创建过程包括选择LLMs、使用少样本上下文学习进行标注,并通过多数投票确定最终标签。该数据集主要应用于媒体偏见检测领域,旨在解决传统标注方法成本高、质量不稳定的问题,提升媒体偏见分类器的性能。
The Anno-lexical dataset was co-created by the National Institute of Informatics in Tokyo, Japan, and the University of Göttingen in Germany. It is a large-scale dataset for media bias classification, containing 48,330 synthetically annotated sentences. The dataset automates the annotation process using Large Language Models (LLMs), which reduces the cost and time required for manual annotation while maintaining high data quality. The development process of the dataset includes selecting appropriate LLMs, conducting annotation via few-shot in-context learning, and determining the final labels through majority voting. This dataset is primarily applied in the field of media bias detection, aiming to solve the problems of high cost and unstable quality of traditional annotation methods, and improve the performance of media bias classifiers.

- 1The Promises and Pitfalls of LLM Annotations in Dataset Labeling: a Case Study on Media Bias Detection日本东京国立信息学研究所,德国哥廷根大学 · 2024年



