遇见数据集

KastorText2KG_V2

收藏
Zenodo2026-03-06 更新2026-05-26 收录
官方服务:

资源简介:

KastorKG is the distilled knowledge graph produced by the https://github.com/datalogism/Kastor pipeline. It is built from DBpedia by filtering triples through SHACL shapes and verifying their grounding in Wikipedia abstracts (Wikicheck step). The KB is stored as a set of named graphs in Apache Jena TDB, queryable via a Corese SPARQL endpoint. --- Statistics - ~141 000 wiki-validated entities across 18 DBpedia classes: Airport, Artist, Athlete, Building, CelestialBody, City, ComicsCharacter, Company, Film, Food, MeanOfTransportation, Monument, MusicalWork, Politician, Scientist, SportsTeam, University, WrittenWork - Relations are drawn exclusively from the DBpedia Ontology (dbo, verified to appear in the corresponding Wikipedia abstract - 4 stratified train/test samples (RD⁰–RD³) and 2 human-corrected annotation sets are included --- RDF namespaces used ┌────────┬─────────────────────────────────────────────┐ │ Prefix │ Namespace │ ├────────┼─────────────────────────────────────────────┤ │ dbo: │ http://dbpedia.org/ontology/ │ ├────────┼─────────────────────────────────────────────┤ │ dbr: │ http://dbpedia.org/resource/ │ ├────────┼─────────────────────────────────────────────┤ │ rdfs: │ http://www.w3.org/2000/01/rdf-schema# │ ├────────┼─────────────────────────────────────────────┤ │ rdf: │ http://www.w3.org/1999/02/22-rdf-syntax-ns# │ ├────────┼─────────────────────────────────────────────┤ │ xsd: │ http://www.w3.org/2001/XMLSchema# │ ├────────┼─────────────────────────────────────────────┤ │ sh: │ http://www.w3.org/ns/shacl# │ ├────────┼─────────────────────────────────────────────┤ │ kstor: │ http://ns.inria.fr/kstor/ │ └────────┴─────────────────────────────────────────────┘ --- Named graph structure (http://ns.inria.fr/kstor/) - shapes/{shape} — SHACL shape definitions - class_randoms_id/{class} — random UUIDs per entity (reproducible sampling) - inferences/{shape} — triples inferred by SPARQL rules materialization (e.g. dbo:birthDate → dbo:birthYear) - wiki_md/{class} — Wikipedia Markdown abstracts (rdfs:comment) - wikichecked/{shape}/abtract_md — main distilled KB (wiki-validated triples) - wikinew_202004/ — temporal split (articles published after April 2020) The default graph (urn:x-arqefaultGraph) holds the full source DBpedia dump. --- Loading and querying Load the dump with Apache Jena tdbloader, then start the Corese server: java -Xmx10g -jar corese-server-4.5.0.jar -init config.properties The SPARQL endpoint is then available at http://localhost:8080/sparql. See the https://github.com/datalogism/Kastor for setup instructions and example queries.

提供机构:
Zenodo
创建时间:
2026-03-04
二维码
社区交流群
二维码
科研交流群
商业服务