Thai SDG-labelled Text Corpus (v1.1)
收藏资源简介:
Thai SDG-labelled Text Corpus (v1.1) The Thai SDG-labelled Text Corpus is a publicly available dataset for Sustainable Development Goal (SDG) text classification in Thai. The corpus contains 6,800 records evenly balanced across all 17 SDGs (400 records per goal). It was derived from the OSDG Community Dataset (OSDG-CD), Version 2024.04 (Zenodo DOI: 10.5281/zenodo.11441197), translated into Thai through a machine-translation and human post-editing workflow, and augmented with expert-reviewed synthetic data for SDG 17 (Partnerships for the Goals), which is not included in the original OSDG-CD. The dataset is intended to support Thai-language and multilingual natural language processing (NLP) research, including SDG classification, cross-lingual learning, information retrieval, and policy document analysis. Version 1.1 update Version 1.1 Updated documentation and metadata to correct the author surname and update the source OSDG-CD citation to Version 2024.04 (DOI: 10.5281/zenodo.11441197). Corrected the description of expert_agreement_score in §3 and the inter-annotator agreement reporting in §5. The field represents raw percent (observed) agreement and is not Cohen's kappa. A separately computed, chance-corrected Cohen's kappa (pooled κ = 0.751; by verdict subgroup: Expert_Verified_Pass κ = 0.847, Expert_Reviewed_Acceptable κ = 0.590, Expert_Flagged_Minor_Nuance κ = 0.537) is now reported alongside the original percent-agreement figures. The underlying CSV and XLSX data files are unchanged; only the documentation has been corrected. Dataset Construction Records for SDG 1–16 were translated from English into Thai through a machine-translation and human post-editing workflow, while SDG 17 records were newly generated and reviewed by experts. Each record includes automatic quality-control indicators based on LaBSE semantic similarity and length-ratio checks. A stratified sample of 1,360 records (80 per SDG) additionally includes human-expert quality assessments and inter-annotator agreement scores. Dataset Contents The dataset is distributed in CSV and XLSX formats and contains 18 documented metadata fields, including: Source provenance SDG labels Translation metadata Semantic similarity scores Quality-control indicators Expert-review outcomes Potential Applications Thai-language SDG text classification Multilingual NLP benchmarking Research funding and policy document analysis Cross-lingual transfer learning Sustainable development analytics Source Dataset and Attribution This derivative dataset was created from the OSDG Community Dataset (OSDG-CD), Version 2024.04. The source dataset should be cited as: OSDG, UNDP IICPSD SDG AI Lab, & PPMI. (2024). OSDG Community Dataset (OSDG-CD) (Version 2024.04) [Dataset]. Zenodo. https://doi.org/10.5281/zenodo.11441197 License This derivative dataset is released under the CC BY 4.0 license. The original OSDG Community Dataset is also distributed under CC BY 4.0, and appropriate attribution has been retained in accordance with the license terms.



