GSAP-Benchmark-KG: A Knowledge Graph of Machine Learning Entities and Relations Extracted from Scholarly Publications
收藏资源简介:
GSAP-Benchmark-KG links the machine-learning models, datasets, methods and tasks mentioneds in 100 ML publications through 18 GSAP-ERE relation types, such as which model was trained or evaluared on which dataset. It can be used to study how ML resources are used and reused across papers, to trace every fact back to the sentence it was annotated in, and as a manually annotated benchmark for building knowledge graphs from scholarly text. The graph provides a queryable RDF representation of the entity and relation mentions annotated in GSAP-ERE, designed to make the machine-learning literature it covers both machine-actionable and provenance-traceable.Scholarly concepts such as models, datasets, methods, tasks, model architectures, and data sources thare mapped to deduplicated concept nodes and typed edges. Every concept-level fact is backed by full mention-level provenance, including the exact text span, sentence, document, and (anonymized) annotator it came from, so the graph supports both aggregate queries over the whole corpus and drill-down into a single fact’s textual evidence. The graph is built with a documented, reproducible pipeline and validated against its own ontology with SHACL before every release.



