遇见数据集

Geosocial Media's Perspective on Energy: A Text Classification Approach using Natural Language Processing

收藏
Zenodo2025-04-09 更新2026-05-26 收录
官方服务:

资源简介:

This repository contains the code, datasets and supplementary materials used in the study "Geosocial Media’s Perspective on Energy: A Text Classification Approach using Natural Language Processing". The aim of the study was to examine public opinion on fossil fuels, nuclear energy, and renewable sources (solar and wind) using Twitter ("X") data and natural language processing techniques. The following materials are included: Labeled Tweet datasets ("labeled_datasets.zip"): Manually labeled samples for each energy source, used to train and validate NLP models. Tweets are labeled as "in favor", "against", "neither", or "irrelevant" based on their stance towards the given energy topic. Labeled validation datasets ("labeled_validationsets.zip"): Subsets of the annotated tweets reserved for model evaluation. These were used to benchmark BERTweet and GPT model performance in stance detection. Interactive HTML maps ("html visualisations.zip"): Visual representations of tweet relevance and stance over time and space, including associated word clouds that highlight frequently used terms. These tools allow users to explore how public opinion has evolved across regions and events. geo_dicts.zip and world-administrative-boundaries.zip: These files contain necessary geographic data for location tagging and administrative boundary assignment. code.ipynb: The Jupyter notebook that demonstrates the analysis workflow. NB: The restricted-access code to the full Twitter datasets is also needed to execute code.ipynb in its entirety. You can request access via this Zenodo link.

提供机构:
Zenodo
创建时间:
2025-04-09
二维码
社区交流群
二维码
科研交流群
商业服务