PhiloKG: A multi-model knowledge graph of classical philosophical discourse
收藏资源简介:
PhiloKG is a knowledge graph of classical philosophical discourse, extracted by an ensemble of large language models from 988 public-domain philosophy texts drawn from Project Gutenberg. The graph contains ~156K raw CONTRADICTS and ~248K raw CONTRASTS_WITH triples, with multi-model consensus counts used to filter high-confidence edges at configurable thresholds. Files: philokg.ttl is a lightweight VoID/DCAT descriptor; philokg.nt is the full N-Triples graph (~434 MB); philokg_annotated.ttl is a Turtle serialization with human-readable annotations (~196 MB). Generative AI disclosure: The triples in this dataset were extracted by an ensemble of 15 large language models across three families: Anthropic Claude (opus-4.6, sonnet-4.6); Google Gemini (2.5-flash, 2.5-flash-lite, 2.5-pro, 3-flash, 3.1-flash-lite, 3.1-pro); and local open-weight models served via Ollama (gemma3:27b, gemma4:31b, glm-4.7-flash, qwen3:32b, qwen3.5:27b, qwen3.5:35b, triplex). A hybrid re-extraction pass combined ollama+triplex as a post-processing step. Each triple carries a consensus_models array recording which of these models independently asserted it. Models generated the substantive scientific content (knowledge triples) from prompts over text chunks; the human author designed the extraction pipeline, prompts, consensus-aggregation logic, and analysis, and is responsible for the dataset's scientific content and conclusions. This disclosure follows ACM's April 2023 authorship policy. Companion essays: ATLAS.md and CONTRADICTIONS.md in the associated source repository analyze the graph's structure; the accompanying ISWC 2026 paper (forthcoming) describes the extraction methodology in full.



