遇见数据集

Supplementary Materials and Dataset for: "LLM-Inferred Narrative Frames in Geopolitical Conflict Reporting"

收藏
Zenodo2025-06-12 更新2026-05-26 收录
官方服务:

资源简介:

This supplementary material contains both tabular data and visualizations derived from a computational media framing analysis conducted using zero-shot large language models (LLMs). The files support the analyses presented in the corresponding study and are intended for transparency, exploratory review, and reproducibility. All frame and entity associations were generated via automated inference and do not reflect human-coded labels or verified facts. framing_predictions.csv Description: This file contains article-level narrative frame predictions, produced through a two-stage classification pipeline using facebook/bart-large-mnli followed by google/flan-t5-large. BART provided initial zero-shot frame scores, and FLAN served as a semantic filter to retain only prominent or well-justified frames. Fields: Column Name Description source Name of the news outlet or source published_date Date the article was published content Full main text of the article predicted_frames Comma-separated list of top BART-predicted frames frame_scores Dictionary of model-predicted frame scores ({frame: score}) from BART validated_frames Final list of frames retained after semantic validation by FLAN validation_flags Raw responses from FLAN model for each evaluated frame entity_level_framing_predictions.csv Description: This file contains sentence-level associations between named entities and narrative frames, filtered through a multi-step semantic validation pipeline. Each entry represents a model-inferred connection between an entity and a frame, based on the sentence context and validated through five FLAN-T5 prompts. Fields: Column Name Description article_id Identifier linking to the original article entity Named entity mentioned in the sentence sentence Sentence in which the entity appears frame Narrative frame assigned to the sentence containing the entity score Model-assigned frame confidence score (from BART) Supplementary Figure S1 Filename: Supplementary Figure S1.png Description: Entity–frame network showing model-inferred associations between a broader set of regionally or politically significant actors (e.g., UNRWA, Palestinian Authority, IDF, Amnesty International) and narrative frames. Nodes represent named entities; edges reflect sentence-level co-occurrences with specific frames. Only entities matching a predefined geopolitical keyword set were included. This graph supports a deeper view of how large language models detect thematic emphases around institutional actors. Supplementary Figure S2 Filename: Supplementary Figure S2.png Description: Detailed entity–frame association graph displaying validated entity–frame pairs extracted from the dataset. Due to the high volume of unique named entities, this figure serves as a comprehensive visualization of LLM-inferred sentence-level framing across all actors. It reflects raw interpretive patterns detected by the model and is intended for transparency and exploratory reference, rather than close reading. Interpretation and Responsible Use All annotations and inferences in this dataset reflect the interpretation of large language models applied to surface-level textual content. These outputs should not be treated as verified facts or editorial claims. No manual coding or human judgment was involved in the labeling process, except for the construction of prompt templates. The materials are provided to support transparency and reproducibility in automated framing analysis, especially within the context of politically sensitive topics. They are not recommended for downstream supervised learning tasks or for use as benchmark ground truth in other framing-related studies. Licensing and Use This supplementary dataset and accompanying visualizations are released under the Creative Commons Attribution-NonCommercial 4.0 International License (CC BY-NC 4.0). They are intended solely for academic, non-commercial use. All article content remains the copyright of the original publishers. No private or personally identifiable information has been included, and journalist bylines were collected only as part of publicly accessible metadata.

本补充材料包含基于零样本大语言模型(LLM)开展的计算式媒体框架分析所生成的表格数据与可视化成果。本数据集文件可辅助对应研究中的分析工作,旨在提升研究透明度、支持探索性评审与结果可复现性。所有框架与实体关联均通过自动化推理生成,不代表人工编码标签或已验证的事实。 framing_predictions.csv 描述: 本文件包含文章级叙事框架预测结果,通过基于facebook/bart-large-mnli与google/flan-t5-large的两阶段分类流水线生成。其中BART提供初始零样本框架得分,FLAN则作为语义过滤器,仅保留突出或合理性充足的框架。 字段: 列名 | 描述 source | 新闻机构或数据源名称 published_date | 文章发布日期 content | 文章完整正文 predicted_frames | 以逗号分隔的BART预测顶级框架列表 frame_scores | BART生成的模型预测框架得分字典(格式为{frame: score}) validated_frames | 经FLAN语义验证后保留的最终框架列表 validation_flags | FLAN模型针对每个评估框架返回的原始响应 entity_level_framing_predictions.csv 描述: 本文件包含命名实体与叙事框架间的句子级关联,经多步语义验证流水线过滤后生成。每条记录代表基于句子上下文由模型推理得到的实体与框架间关联,并通过5条FLAN-T5提示词完成验证。 字段: 列名 | 描述 article_id | 关联至原始文章的标识符 entity | 句子中提及的命名实体 sentence | 实体所在的句子 frame | 分配给包含该实体的句子的叙事框架 score | 模型(BART)赋予的框架置信度得分 Supplementary Figure S1 文件名:Supplementary Figure S1.png 描述: 实体-框架网络图谱,展示了一批具有区域或政治影响力的主体(如UNRWA、巴勒斯坦权力机构、IDF、大赦国际)与叙事框架间由模型推理得到的关联。节点代表命名实体,边则体现其与特定框架的句子级共现关系。本图谱仅纳入符合预定义地缘政治关键词集合的实体,可帮助深入理解大语言模型如何识别围绕制度性主体的主题侧重。 Supplementary Figure S2 文件名:Supplementary Figure S2.png 描述: 详细实体-框架关联图谱,展示了从本数据集中提取的经验证的实体-框架配对结果。由于唯一命名实体数量较多,本图全面可视化了大语言模型针对所有主体推理得到的句子级框架分析结果。其反映了模型检测到的原始解读模式,仅用于提升研究透明度与探索性参考,而非用于精细研读。 解读与负责任使用 所有标注与推理结果均为大语言模型针对表层文本内容的解读产物,不应被视为已验证的事实或社论主张。除提示词模板构建环节外,标注过程未涉及人工编码或人类判断。 本补充材料旨在助力自动化框架分析的透明度与可复现性,尤其适用于政治敏感议题场景。本材料不建议用于下游监督学习任务,亦不可作为其他框架相关研究的基准真实标签。 许可与使用条款 本补充数据集及配套可视化成果采用知识共享署名-非商业性使用4.0国际许可协议(CC BY-NC 4.0)发布,仅可用于学术非商业用途。 所有文章内容的版权仍归原出版机构所有。本数据集未包含任何私人或可识别个人身份的信息,记者署名仅作为公开可获取的元数据的一部分进行收集。

提供机构:
Zenodo
创建时间:
2025-06-12
二维码
社区交流群
二维码
科研交流群
商业服务