遇见数据集

Geosocial Media's Perspective on Energy: A Text Classification Approach using Natural Language Processing

收藏
Zenodo2025-04-09 更新2026-05-26 收录
官方服务:

资源简介:

This repository contains the code, datasets and supplementary materials used in the study "Geosocial Media’s Perspective on Energy: A Text Classification Approach using Natural Language Processing". The aim of the study was to examine public opinion on fossil fuels, nuclear energy, and renewable sources (solar and wind) using Twitter ("X") data and natural language processing techniques. The following materials are included: Labeled Tweet datasets ("labeled_datasets.zip"): Manually labeled samples for each energy source, used to train and validate NLP models. Tweets are labeled as "in favor", "against", "neither", or "irrelevant" based on their stance towards the given energy topic. Labeled validation datasets ("labeled_validationsets.zip"): Subsets of the annotated tweets reserved for model evaluation. These were used to benchmark BERTweet and GPT model performance in stance detection. Interactive HTML maps ("html visualisations.zip"): Visual representations of tweet relevance and stance over time and space, including associated word clouds that highlight frequently used terms. These tools allow users to explore how public opinion has evolved across regions and events. geo_dicts.zip and world-administrative-boundaries.zip: These files contain necessary geographic data for location tagging and administrative boundary assignment. code.ipynb: The Jupyter notebook that demonstrates the analysis workflow. NB: The restricted-access code to the full Twitter datasets is also needed to execute code.ipynb in its entirety. You can request access via this Zenodo link.

本仓库包含研究《地缘社交媒体视域下的能源议题:基于自然语言处理(Natural Language Processing,NLP)的文本分类方法》所使用的代码、数据集与补充材料。本研究旨在利用推特(Twitter,现更名为X)数据与自然语言处理技术,探究公众针对化石燃料、核能及太阳能、风能两类可再生能源的舆论态度。本次收录的材料如下: 1. 带标注的推特数据集(labeled_datasets.zip):针对各类能源主题的人工标注样本,用于训练与验证自然语言处理模型。推文将根据其对应能源主题的立场被标注为“支持”“反对”“中立”或“不相关”四类。 2. 带标注的验证数据集(labeled_validationsets.zip):从已标注推文中抽取的专属子集,用于模型评估。本数据集被用于基准测试BERTweet与GPT模型在立场检测任务中的性能表现。 3. 交互式HTML可视化地图(html visualisations.zip):用于可视化不同时空维度下推文的相关性与立场分布,附带展示高频用词的词云。该工具支持用户探究不同地区与事件背景下公众舆论的演变轨迹。 4. geo_dicts.zip与world-administrative-boundaries.zip:此类文件包含用于地理位置标记与行政边界分配的必要地理数据。 5. code.ipynb:用于完整演示本研究分析流程的Jupyter笔记本文件。 注:完整推特数据集的受限访问代码是完整运行code.ipynb的必要前提。用户可通过此Zenodo链接申请访问权限。

提供机构:
Zenodo
创建时间:
2025-04-09
二维码
社区交流群
二维码
科研交流群
商业服务