annotated dataset
收藏资源简介:
该数据集由赫尔辛基大学研究团队创建,旨在支持基于检索增强生成的接地讽刺内容研究。数据集包含100条人工标注的讽刺性词典定义,每条定义均基于芬兰广播公司Yle的英文新闻内容生成,并由六位标注者从幽默性和政治相关性维度进行评分。数据通过自动化流程采集,包括网络爬取、时间戳过滤、情感分析和主题建模,最终利用RAG框架生成定义。该数据集主要应用于自然语言生成和计算幽默领域,用于评估模型在特定文化背景下生成具有政治意义的讽刺内容的能力,并探索LLM作为评估工具的可靠性。
This dataset was created by a research team at the University of Helsinki to support research on grounded sarcastic content leveraging retrieval-augmented generation (RAG). It comprises 100 manually annotated sarcastic dictionary definitions, each generated from English news content sourced from Finnish public broadcaster Yle, and scored by six annotators across two dimensions: humor and political relevance. The data was collected through an automated pipeline including web crawling, timestamp filtering, sentiment analysis, and topic modeling, with the final definitions generated using a RAG framework. This dataset is primarily applied in the domains of natural language generation (NLG) and computational humor, to evaluate models' capability of generating politically salient sarcastic content within specific cultural contexts, and to explore the reliability of large language models (LLMs) as evaluation tools.

- 1Grounded Satirical Generation with RAG赫尔辛基大学 · 2026年



