遇见数据集

Supplementary Materials and Dataset for: "LLM-Inferred Narrative Frames in Geopolitical Conflict Reporting"

收藏
Zenodo2025-10-19 更新2026-05-26 收录
官方服务:

资源简介:

This supplementary material contains both tabular data and visualizations derived from a computational media framing analysis conducted using zero-shot large language models (LLMs). The files support the analyses presented in the corresponding study and are intended for transparency, exploratory review, and reproducibility. All frame and entity associations were generated via automated inference and do not reflect human-coded labels or verified facts. framing_predictions.csv Description: This file contains article-level narrative frame predictions, produced through a two-stage classification pipeline using facebook/bart-large-mnli followed by google/flan-t5-large. BART provided initial zero-shot frame scores, and FLAN served as a semantic filter to retain only prominent or well-justified frames. Note:To comply with copyright restrictions, full article texts have been replaced with standardized placeholders referencing the original source URLs. All analytical fields remain unchanged. Fields: Column Name Description source Name of the news outlet or source published_date Date the article was published url Original article URL content Placeholder indicating omission of full text due to copyright resolved_url Canonical or redirected URL used for analysis predicted_frames Comma-separated list of top BART-predicted frames frame_scores Dictionary of model-predicted frame scores ({frame: score}) from BART validated_frames Final list of frames retained after semantic validation by FLAN validation_flags Raw responses from FLAN model for each evaluated frame entity_level_framing_predictions.csv Description: This file contains sentence-level associations between named entities and narrative frames, filtered through a multi-step semantic validation pipeline. Each entry represents a model-inferred connection between an entity and a frame, based on the sentence context and validated through five FLAN-T5 prompts. Note:To ensure copyright compliance, sentence texts have been replaced with placeholders while preserving all entity–frame relationships and model scores. Fields: Column Name Description article_id Identifier linking to the original article entity Named entity mentioned in the sentence sentence Placeholder indicating omission of sentence text frame Narrative frame assigned to the sentence containing the entity score Model-assigned frame confidence score (from BART) High-Resolution Figures (Figures 1–5) The following high-resolution figures correspond to visualizations presented in the main paper. Due to visual complexity or density, these figures are included separately for clarity and legibility. All filenames are prefixed with highres_ for easier identification. Filename Description highres_Figure_1.pdf Comparison of average frame scores across sources (two-stage inference: BART and FLAN-T5) highres_Figure_2.pdf Temporal distribution of predicted narrative frames (May 2–6, 2025) highres_Figure_3.pdf Frame usage frequency across top news sources highres_Figure_4.png Entity-frame associations for prominent geopolitical actors highres_Figure_5.png Heatmap of co-occurrence patterns between narrative frames at the sentence level Supplementary Figure S1 Filename: Supplementary Figure S1.png Description: Entity–frame network showing model-inferred associations between a broader set of regionally or politically significant actors (e.g., UNRWA, Palestinian Authority, IDF, Amnesty International) and narrative frames. Nodes represent named entities; edges reflect sentence-level co-occurrences with specific frames. Only entities matching a predefined geopolitical keyword set were included. This graph supports a deeper view of how large language models detect thematic emphases around institutional actors. Supplementary Figure S2 Filename: Supplementary Figure S2.png Description: Detailed entity–frame association graph displaying validated entity–frame pairs extracted from the dataset. Due to the high volume of unique named entities, this figure serves as a comprehensive visualization of LLM-inferred sentence-level framing across all actors. It reflects raw interpretive patterns detected by the model and is intended for transparency and exploratory reference, rather than close reading. Interpretation and Responsible Use All annotations and inferences in this dataset reflect the interpretation of large language models applied to surface-level textual content. These outputs should not be treated as verified facts or editorial claims. No manual coding or human judgment was involved in the labeling process, except for the construction of prompt templates. The materials are provided to support transparency and reproducibility in automated framing analysis, especially within the context of politically sensitive topics. They are not recommended for downstream supervised learning tasks or for use as benchmark ground truth in other framing-related studies. Licensing and Use Copyright © 2025 Arpish R. Solanki. This supplementary dataset and accompanying visualizations are released under the Creative Commons Attribution-NonCommercial 4.0 International License (CC BY-NC 4.0). They are intended solely for academic, non-commercial use.The license applies only to derived annotations, metadata, and visualizations.Original article content remains the copyright of the respective publishers.

提供机构:
Zenodo
创建时间:
2025-10-19
二维码
社区交流群
二维码
科研交流群
商业服务