遇见数据集

Web-scraped corpus of URV teaching guides for SDG-oriented text mining

收藏
Zenodo2025-11-19 更新2026-05-26 收录
官方服务:

资源简介:

This dataset contains a longitudinal corpus of teaching guides (course syllabi) from the Rovira i Virgili University (URV), collected through web scraping of the official online catalogue of courses. The main purpose of this repository is to support text mining analyses aimed at detecting references to the United Nations Sustainable Development Goals (SDGs) and related concepts (e.g. using tools such as text2sdg or similar approaches). The repository is organised as a series of scraping runs, each stored in a separate subfolder named with the pattern `scraping_urv-guides_YYYYMMDD`, where `YYYYMMDD` indicates the date of the web scraping in reverse date format. Each folder includes: - The original scraped content (for example, CSV, JSON or JSONL files),- A `run-metadata_YYYYMMDD.json` file describing the configuration and context of that specific scraping run. The top-level `README.md` file and `dataset-metadata.json` file document the overall structure, intended use and basic methodological choices of the project. The dataset is designed to be extensible over time: new versions of this Zenodo record are expected to add additional scraping runs (new subfolders) while preserving the previous ones, allowing for temporal comparisons and reproducible research. This corpus is intended for research and teaching purposes only. Users must comply with URV policies and applicable legal requirements regarding the use of teaching materials and web-scraped content. Any publication making use of this dataset should acknowledge the original data source (URV) and cite this Zenodo record.

提供机构:
Zenodo
创建时间:
2025-11-19
二维码
社区交流群
二维码
科研交流群
商业服务