遇见数据集

PhiloKG: A Multi-Model Knowledge Graph of Classical Philosophical Discourse

收藏
Zenodo2026-04-23 更新2026-05-26 收录
官方服务:

资源简介:

Version 1.5.2: supplementary-only minor release — no KG changes. Adds 4 reproducibility artifacts: (1) graphrag_demo.py — standalone Python script running 2 SPARQL competency queries against the DuckDB ledger, with machine-readable output; (2) graphrag_demo_results_v1_5_1.json — re-verified retrievals at v1.5.1 (10 Plato↔scepticism triples; 8 Stoic-nature-virtue-reason triples at C≥3, Seneca "life according to reason" now at C=4); (3) FOUR_MODEL_EXTENSION_ANALYSIS.md — pair-wise Jaccard analysis on 787 book-corpus chunks extracted by opus-4.7 + gemini-3.1-pro + gemma3:27b + mixtral:8×7b (opus↔gemini mean J=0.260, highest cross-family agreement observed; mixtral identified as strong extraction-profile outlier at J≈0.04); (4) GRAPHRAG_DEMO.md refreshed for v1.5.1 numbers. Converts the paper's §6 "GraphRAG evaluation" future-work reference from aspirational to validated-runnable, and extends §5.2's "Agreement generalizes beyond Aeschylus" finding from 60 to 787 chunks. KG itself (philokg.nt, philokg_annotated.ttl, philokg.ttl, corpus_manifest.json, ontology, benchmarks, precision pilot) is byte-identical to v1.5.1. PhiloKG is a knowledge graph of classical philosophical discourse comprising 2.34 million triples and 200,000 entities extracted from 937 texts by 122 named authors spanning from ancient Greece to the early twentieth century. The resource includes the full knowledge graph in RDF N-Triples format, an annotated Turtle export using RDF 1.2 triple-term reification for per-triple consensus metadata, the RDFS/OWL ontology with a companion SHACL shapes file, per-model benchmark data for the 30-chunk Aeschylus benchmark (6 models) and the 250-chunk cross-vendor benchmark (4 models across Anthropic and Google), the two-tier extraction pipeline source code, and the raw per-judge verdicts from the 500-triple multi-judge precision audit (5 LLM judges across 4 families).

版本1.5.2:仅含补充内容的次要版本更新,未对知识图谱(Knowledge Graph, KG)进行任何修改。本次新增4项可复现性配套资源: 1. `graphrag_demo.py`:独立Python脚本,可针对DuckDB账本执行2条SPARQL能力查询,并输出机器可读格式结果; 2. `graphrag_demo_results_v1_5_1.json`:针对v1.5.1版本的检索结果进行了重新验证(包含10条柏拉图↔怀疑主义三元组;8条斯多葛学派-自然-德性-理性三元组,置信度C≥3,其中塞涅卡“遵循理性的生活”相关三元组置信度提升至C=4); 3. `FOUR_MODEL_EXTENSION_ANALYSIS.md`:针对由opus-4.7、gemini-3.1-pro、gemma3:27b、mixtral:8×7b四款模型抽取的787个语料块进行两两杰卡德(Jaccard)相似度分析(opus与gemini的平均杰卡德相似度为0.260,为当前观测到的跨模型家族最高一致性;mixtral被识别为抽取轮廓异常值,杰卡德相似度约为0.04); 4. 针对v1.5.1版本的结果更新了`GRAPHRAG_DEMO.md`文档。 本次更新将论文第6节“GraphRAG评估”中原本属于未来工作范畴的参考内容,从仅为愿景的设想转为可验证、可运行的实现;同时将第5.2节“一致性可推广至埃斯库罗斯以外文本”的研究发现,从基于60个语料块的结论扩展至787个语料块。 本次更新未改动知识图谱相关文件(包括`philokg.nt`、`philokg_annotated.ttl`、`philokg.ttl`、`corpus_manifest.json`、本体文件、基准测试集、精度预实验数据集),其字节流与v1.5.1版本完全一致。 PhiloKG(哲学知识图谱)是一套古典哲学论述知识图谱,共包含234万个三元组与20万个实体,这些数据从937篇文本中抽取而来,涵盖古希腊至20世纪早期的122位知名作者的作品。该资源包含以下内容:RDF N三元组格式的完整知识图谱;基于RDF 1.2三元组术语具体化实现的带注释Turtle导出文件,用于存储每条三元组的共识元数据;配套SHACL形状校验文件的RDFS/OWL本体;针对30个语料块的埃斯库罗斯基准测试集(覆盖6款模型)以及250个语料块的跨厂商基准测试集(覆盖Anthropic与Google旗下的4款模型)的单模型基准测试数据;双层抽取流水线的源代码;以及针对500条三元组的多评委精度审计的原始评委判定结果(来自4个模型家族的5位大语言模型(Large Language Model, LLM)评委)。

提供机构:
Zenodo
创建时间:
2026-04-23
二维码
社区交流群
二维码
科研交流群
商业服务