遇见数据集

OSDG-CD Multi-Label SDG Classification Benchmark (5K Sample): LLM and Keyword-Based Predictions

收藏
Zenodo2026-06-19 更新2026-06-21 收录
官方服务:

资源简介:

A 5,000-entry benchmark subset of the OSDG Community Dataset (OSDG-CD), drawn as a uniform random sample with no filtering, augmented with multi-label Sustainable Development Goal (SDG) classifications from five state-of-the-art Large Language Models and two keyword-based classifiers. Unlike the original OSDG-CD, which assigns a single SDG label per excerpt, this benchmark provides multi-label predictions: each classifier may assign zero, one, or several SDG labels (1–17) to a given text. This supports evaluation of multi-label agreement, coverage, and overlap between LLM-based and keyword-based approaches. Added columns LLM predictions (multi-label, integer SDGs 1–17): deepseek/deepseek-v3.2, anthropic/claude-sonnet-4.6, google/gemini-3-flash-preview, openai/gpt-5.4, qwen/qwen3.6-plus Keyword-based predictions: elsevier_2025 (Elsevier 2025 SDG mapping) and partaourides2023thematic (consolidated thematic keyword set synthesizing the Aurora, Auckland, and Elsevier families). Provenance Derivative of the OSDG Community Dataset (OSDG-CD) v2022.10, released under CC BY 4.0: OSDG, UNDP IICPSD SDG AI Lab, & PPMI (2022), https://doi.org/10.5281/zenodo.7136826. All original columns (text_id, doi, text, sdg, labels_negative, labels_positive, agreement) are preserved verbatim. Citation Please cite both this dataset and the original OSDG-CD. See README.md for the full schema, classifier details, and example record.

提供机构:
Zenodo
创建时间:
2026-06-19
二维码
社区交流群
二维码
科研交流群
商业服务