darshana-graph
收藏资源简介:
Darshana Graph是一个基于文本的印度哲学知识图谱数据集,旨在将印度主要哲学传统(包括印度教六派古典哲学、佛教大乘和上座部传统,以及耆那教核心哲学文本)映射成一个结构化的数据集。其核心特点是所有概念和关系都锚定在真实、可引用的源文本具体段落中,而非由模型先验知识生成。数据集采用封闭的、预先定义的关系词汇表,并使用大语言模型仅对给定段落中已写内容进行分类。当前版本(v3)共包含45,155条边,涵盖印度教、佛教和耆那教传统,由原始的28,322条吠檀多边、新增的16,068条大乘佛教边和765条上座部佛教边构成。数据集文件包括多个图文件(如推荐用于跨传统分析的darshana_graph_v3.jsonl)、包含125,040条记录的原始对齐文本语料库文件(darshana_corpus.jsonl),以及多个衍生分析文件(如包含时间归因和学术引用的时间源层、计算出的跨传统结构同源词对、经过词义消歧的概念出现记录等)。数据模式为JSONL格式,每条记录包含概念对、关系类型、学派、置信度、证据引用和来源等字段。关系类型包括IS_IDENTICAL_TO、IS_DISTINCT_FROM、PRESUPPOSES等。该数据集的独特之处在于,它对齐了相同根本文本的多个独立历史注释者和学派,使得同一偈颂或经文可以通过多达18位不同注释者的视角进行解读。数据集适用于哲学概念分析、跨传统比较、知识图谱构建、文本分类和问答等任务。需要注意的是,数据集存在已知局限性,如早期版本(v1)因主要从梵语吠檀多注释中提取而存在语料库组成偏差,部分传统(如上座部)的重新提取尚不完整,以及存在"general"学派标签过度使用等情况。
Darshana Graph is a text-based knowledge graph dataset for Indian philosophy, designed to map major Indian philosophical traditions (including the six classical schools of Hinduism, Mahayana and Theravada Buddhist traditions, and core Jain philosophical texts) into a structured dataset. Its core feature is that all concepts and relationships are anchored in real, citable source text passages, rather than generated from model prior knowledge. The dataset employs a closed, predefined relationship vocabulary and uses large language models to classify only the written content in given passages. The current version (v3) contains 45,155 edges, covering Hindu, Buddhist, and Jain traditions, consisting of the original 28,322 Vedanta edges, newly added 16,068 Mahayana Buddhist edges, and 765 Theravada Buddhist edges. Dataset files include multiple graph files (e.g., darshana_graph_v3.jsonl recommended for cross-traditional analysis), a raw aligned text corpus file with 125,040 records (darshana_corpus.jsonl), and several derived analysis files (such as a temporal source layer with time attributions and academic citations, computed cross-traditional structural cognate pairs, and disambiguated concept occurrence records). The data schema is in JSONL format, with each record containing fields like concept pairs, relationship types, school, confidence, evidence citations, and sources. Relationship types include IS_IDENTICAL_TO, IS_DISTINCT_FROM, PRESUPPOSES, etc. The datasets uniqueness lies in aligning multiple independent historical commentators and schools on the same root texts, allowing a single verse or scripture to be interpreted through the perspectives of up to 18 different commentators. It is suitable for tasks such as philosophical concept analysis, cross-traditional comparison, knowledge graph construction, text classification, and question answering. Note that the dataset has known limitations, such as corpus composition bias in early versions (v1) due to extraction primarily from Sanskrit Vedanta commentaries, incomplete re-extraction for some traditions (e.g., Theravada), and overuse of the "general" school label.
Darshana Graph:印度哲学知识图谱
数据集概述
Darshana Graph 是一个结构化、基于文本的印度哲学知识图谱,涵盖印度六大古典印度教派别(darshanas)、大乘佛教、上座部佛教以及耆那教核心哲学文本。数据集中的每个概念和关系都锚定于真实、可引用的源文本段落,使用大型语言模型仅基于预定义的关系类型词汇对已有文本进行分类,不依赖模型先验知识生成内容。
当前版本(v3)包含45,155条边,涵盖印度教、佛教和耆那教传统。
语言与许可
- 语言:英语、梵语、印地语、巴利语
- 许可协议:CC-BY-4.0
文件结构
图谱文件
| 文件 | 边数 | 描述 |
|---|---|---|
darshana_graph.jsonl |
28,322 | v1原始版本,仅吠檀多注释语料库 |
darshana_graph_v1.jsonl |
28,322 | 同darshana_graph.jsonl,附显式版本标签 |
mahayana_edges.jsonl |
16,068 | 大乘佛教提取 |
pali_edges.jsonl |
765 | 上座部巴利语提取 |
darshana_graph_v2.jsonl |
44,390 | v2版本,原始+大乘 |
darshana_graph_v3.jsonl |
45,155 | v3版本,原始+大乘+上座部 |
buddhist_edges.jsonl |
16,833 | 所有佛教提取合并 |
jain_edges.jsonl |
3,659 | 耆那教经典提取 |
non_vedic_edges.jsonl |
21,257 | 所有非吠檀多提取合并 |
darshana_graph_v4.jsonl |
48,814 | v4版本,最完整版本 |
语料库文件
| 文件 | 记录数 | 描述 |
|---|---|---|
darshana_corpus.jsonl |
125,040 | 原始对齐文本语料库,含完整巴利经典 |
衍生分析文件
temporal_source_layer.json:时间归因层,包含40个学术引用来源和约120个概念断代homologues_v7.json:前200对跨传统结构同源对occurrences_clustered_v3.json:52个关键哲学概念的去歧义化概念出现记录darshana_temporal_v2.graphml:富时间标签的图谱文件,可直接在Gephi中打开
传统覆盖范围
| 传统 | v1 (28,322) | v2 (44,390) | v3 (45,155) |
|---|---|---|---|
| 印度教吠檀多 | ✓ | ✓ | ✓ |
| 印度教数论、正理论、胜论、瑜伽 | ✓ | ✓ | ✓ |
| 耆那教 | ✓ | ✓ | ✓ |
| 佛教 | ✓ | ✓ | ✓ |
| 大乘佛教 | — | ✓ | ✓ |
| 上座部佛教 | — | — | ✓(部分) |
| 耆那教经典 | — | — | — |
数据模式
每条边采用JSON格式,包含以下字段:
concept_a:概念Aconcept_b:概念Brelation:关系类型school:学派标签confidence:置信度evidence_quote:证据引用source:来源
预定义关系类型词汇
IS_IDENTICAL_TO、IS_DISTINCT_FROM、IS_QUALIFIED_ASPECT_OF、IS_SIMULTANEOUSLY_ONE_AND_DIFFERENT、PRESUPPOSES、SUBLATES、LEADS_TO、OBSTRUCTS、IS_CAUSE_OF、IS_MANIFESTATION_OF、CONTRADICTS_IN_SCHOOL、DEFINED_AS
注释者覆盖范围
数据集将相同根文本对齐到18位不同历史注释者和学派:
| 学派 | 包含的注释者 |
|---|---|
| 不二吠檀多 | Shankara, Sridhara, Anandagiri, Nilakantha, Dhanpati |
| 有限不二论 | Ramanuja, Vedanta Desika, Adidevananda |
| 二元论 | Madhva |
| 二元不二论 | Nimbarka, Srinivasa(部分) |
| 不可思议的差别与无差别 | Prabhupada |
| 克什米尔湿婆派 | Abhinavagupta |
| 新吠檀多/一般 | Ramsukhdas, Tejomayananda, Purohit, Sankaranarayan |
已知局限
- 语料库构成偏差:v1图谱主要从梵语吠檀多注释中提取,佛教和耆那教概念主要以印度教注释者引用形式出现
- 词汇证验与哲学概念区分:时间归因层记录的是最早的词汇证验,而非最早的哲学概念发展
- 上座部重新提取不完整:pali_edges.jsonl仅包含来自15个文本文件的765条边
- "general"学派标签过度使用:v1中约73%的边携带"general"学派标签
- IS_QUALIFIED_ASPECT_OF关系过度代表:该关系类型相对于其他类型被过度使用
- 无人类专家审核:所有标签均为单次LLM分类,预计精确率70-85%
来源
- v1(吠檀多):SuttaCentral bilara-data、github.com/gita/gita、Thibaut的《东方圣书》译本、Gambhirananda、Jacobi、Vijay K. Jain
- v2(大乘新增):心经、中论、楞伽经、法华经、维摩诘经、入菩萨行论、六祖坛经、大乘宝积经、首楞严经、金刚经注释、地藏经、净土三经
- v3(上座部新增):来自Access to Insight和SuttaCentral的巴利经典摘录




