Web-scraped corpus of URV teaching guides for SDG-oriented text mining
收藏资源简介:
This dataset contains a published scraping batch of teaching guides (course syllabi) from the Universitat Rovira i Virgili (URV), prepared to support SDG-oriented text mining, curriculum analysis, and reproducible downstream processing. This release corresponds to the batch `scraped-20251127`, collected on `2025-11-27`. At the time of collection, URV teaching guides were distributed across two coexisting systems, Docnet and GUIdO. This batch combines both sources in order to maximise coverage and preserve the publication context of that transition period. The Zenodo upload includes:- a ZIP archive named `urv_teaching_guides-scraped-20251127.zip`- `README.md`- `LICENSE.txt`- `dataset-metadata.json` The ZIP archive contains the folder `urv_teaching_guides/scraped-20251127/` with the following files:- `1_centres_list.csv`- `2_programmes_list.csv`- `3_course_details_list.csv`- `4_docnet_course_urls.csv`- `5_docnet_course_info.csv`- `5_guido_course_info.csv`- `scraping-meta.yml` The file `scraping-meta.yml` stores batch-level provenance information, including scraping date, authorship, source organisation, scope, and the Docnet/GUIdO platform context for this release. This archival package is the published data input used by the URV SDGs workflow, in which `urv-sdgs-tracker` prepares derived outputs that are then published through `urv-sdgs-api` and consumed by `urv-sdgs-dashboard`. New data releases should be published as new Zenodo versions, keeping the concept DOI stable while assigning a version-specific DOI to each published batch.



