PhiloKG: A Multi-Model Knowledge Graph of Classical Philosophical Discourse
收藏资源简介:
Version 1.5.1: two orthogonal detection methods (content-hash scan + 5-chunk audit of the 73 remaining <5-file directories) surfaced 7 further misfiles the v1.1.1-v1.5.0 audits had missed. Two cross-author misfiles caught by content-hash scan (goethe/elective_affinities = Nietzsche Zarathustra; epicurus/diogenes_laertius_lives = Diogenes Laertius Lives — empties the Epicurus directory). Five caught by the 5-chunk audit (catullus = nursery rhymes, goethe/werther = Faust German, horace/echoes = Field humorous verse, al_ghazali/alchemy = Black Troopers stories, al_ghazali/moslem_seeker = Zwemer-about-Ghazali). Corpus reduced from 942 → 937 texts; named authors 123 → 122 (Epicurus lost all files). Post-quarantine ledger: 2,343,929 triples, 495,420 chunks, 200,145 entities, 5,496 at C≥5. RDF artifacts regenerated (philokg.nt 483.2 MB, philokg_annotated.ttl 237.5 MB). See ERRATUM_v1.5.1.md. Cumulative v1.1.1→v1.5.1: 67 files quarantined / 28,794 chunks / 111,397 triples. PhiloKG is a knowledge graph of classical philosophical discourse comprising 2.34 million triples and 200,000 entities extracted from 937 texts by 122 named authors spanning from ancient Greece to the early twentieth century. The resource includes the full knowledge graph in RDF N-Triples format, an annotated Turtle export using RDF 1.2 triple-term reification for per-triple consensus metadata, the RDFS/OWL ontology with a companion SHACL shapes file, per-model benchmark data for the 30-chunk Aeschylus benchmark (6 models) and the 250-chunk cross-vendor benchmark (4 models across Anthropic and Google), the two-tier extraction pipeline source code, and the raw per-judge verdicts from the 500-triple multi-judge precision audit (5 LLM judges across 4 families).



