遇见数据集

Media Coding Dataset for News Content Analysis

收藏
Zenodo2025-06-29 更新2026-05-26 收录
官方服务:

资源简介:

This dataset accompanies the study Beyond Manual Media Coding: Evaluating Large Language Models and Agents for News Content Analysis. It provides a reproducible benchmark for evaluating automated content analysis methods against human-annotated ground truth. The dataset includes: articles.csvContains the 200 news articles collected for this study, each with: id: unique identifier url: source URL of the original article content: full text of the news article codebook.jsonA structured JSON file defining the 26-question analysis codebook used for annotation.Each question entry specifies: questionId: question ID (e.g., Q1) prompt: annotation question text questionAnswerType: type (SINGLE_CHOICE or MULTI_CHOICE) eligibleQuestionAnswers: list of possible tags / codes annotations.jsonContains the complete human annotation data.For each article id, it provides the list of responses to all 26 codebook questions as determined by an expert annotator, establishing the ground truth labels. Intended use Designed for research popuses including natural language understanding, content classification, and LLM evaluation. Please request access with your academic email.

提供机构:
Zenodo
创建时间:
2025-06-29
二维码
社区交流群
二维码
科研交流群
商业服务