oikonomia-db
收藏资源简介:
OIKONOMIA-DB是一个可审计的经济数据库,专注于希腊罗马时期埃及(公元前300年至公元700年)的日常经济活动。该数据集从杜克文献纸莎草数据库(DDbDP)的61,249份文献中提取结构化信息,涵盖税收收据、租约、贷款、工资账目、人口普查记录和私人信件等多种类型。数据集包含195,906条货币事实、350,206条带性别标注的人物提及和21,895个交易主体,每行数据均与源文献中的具体字符跨度关联,确保可追溯性。数据集由八个Parquet表组成,分为两类提取机制:1)货币、价格、税收表通过确定性词汇和规则从解码的数字中提取,具有高精度但存在系统性的遗漏;2)人物、交易主体、去重人物和自治表通过训练模型提取,实体F1为0.737,关系F1为0.623。核心表包括:documents(文献元数据,61,249行)、monetary(货币金额,195,906行)、prices(清洁商品单价,98行)、taxes(清洁税款支付,592行)、persons(人物提及及性别,350,206行)、principals(交易主体及交易类型,21,895行)、persons_distinct(基于保守表面键的去重人物,17,362行)以及autonomy(按世纪和区域汇总的妇女自治比例,32行)。数据表通过两个关键字段连接:stem(唯一文献标识)和tm_id(Trismegistos文献ID,非唯一)。每行事实表包含字符跨度偏移量(0起始、结束排除),便于直接定位源文本。数据集适用于定量古代经济史研究,如价格与工资序列、财政历史、性别与法律能力、人物志和区域比较,也可作为希腊信息提取任务的跨度链接评估集。数据集明确不适用于语言模型训练、绝对经济总量计算、支付方向分析或替代原始文献阅读。数据集遵循CC BY 3.0许可,源语料来自DDbDP和HGV元数据,并整合了HGV日期、Pleiades地点和Trismegistos文献ID等权威数据。
OIKONOMIA-DB is an auditable economic database focusing on daily economic activities in Greco-Roman Egypt (300 BCE to 700 CE). The dataset extracts structured information from 61,249 documents in the Duke Databank of Documentary Papyri (DDbDP), covering various types such as tax receipts, leases, loans, wage accounts, census records, and private letters. It includes 195,906 monetary facts, 350,206 person mentions with gender annotations, and 21,895 transaction principals, with each row linked to specific character spans in the source documents for traceability. The dataset consists of eight Parquet tables, divided into two extraction mechanisms: 1) monetary, price, and tax tables are extracted deterministically from decoded numbers using lexical and rule-based methods, offering high precision but with systematic omissions; 2) person, principal, distinct person, and autonomy tables are extracted via trained models, with an entity F1 of 0.737 and a relation F1 of 0.623. Core tables include: documents (document metadata, 61,249 rows), monetary (monetary amounts, 195,906 rows), prices (cleaned commodity unit prices, 98 rows), taxes (cleaned tax payments, 592 rows), persons (person mentions and gender, 350,206 rows), principals (transaction principals and types, 21,895 rows), persons_distinct (deduplicated persons based on conservative surface keys, 17,362 rows), and autonomy (womens autonomy proportions aggregated by century and region, 32 rows). Tables are linked via two key fields: stem (unique document identifier) and tm_id (Trismegistos document ID, non-unique). Each row in fact tables includes character span offsets (0-indexed, end-exclusive) for direct source text localization. The dataset is suitable for quantitative ancient economic history research, such as price and wage series, fiscal history, gender and legal capacity, prosopography, and regional comparisons, and can serve as a span-linking evaluation set for Greek information extraction tasks. It is explicitly not intended for language model training, absolute economic aggregate calculations, payment direction analysis, or as a substitute for reading original documents. The dataset is licensed under CC BY 3.0, with source corpora from DDbDP and HGV metadata, integrating authoritative data like HGV dates, Pleiades places, and Trismegistos document IDs.
OIKONOMIA-DB:希腊罗马埃及经济数据库
数据集概述
OIKONOMIA-DB 是一个结构化的经济数据库,包含从 61,249 份纸莎草文献(公元前300年至公元700年)中提取的 195,906 个货币事实、350,206 个带性别的人物提及和 21,895 个交易当事人。每条记录均可回溯至特定文献中的具体字符范围,并锁定在固定的语料库版本。
来源语料库: Duke Databank of Documentary Papyri (DDbDP) + HGV 元数据,采用 CC BY 3.0 许可
锁定版本: d7a34f302d1e44e271256092c2b780733187b478
提取模型: OIKONOMIA-Grammateus(实体)· OIKONOMIA-Homologia(关系)
代码仓库: github.com/abderahmane-ai/oikonomia
数据集结构
数据表概览
数据集包含 8 张 Parquet 表,其中 7 张以单条提取语句为粒度存储事实,1 张(autonomy)为已发布的聚合数据。
| 表名 | 文件 | 行数 | 一行表示 | 主键 |
|---|---|---|---|---|
documents |
export/documents.parquet |
61,249 | 一份文本文献(主干) | stem |
monetary |
monetary.parquet |
195,906 | 一个货币金额 | tm_id + 跨度 |
prices |
prices.parquet |
98 | 一份清洁的商品单价 | tm_id + 跨度 |
taxes |
taxes.parquet |
592 | 一份清洁的税款支付 | tm_id + 跨度 |
persons |
persons.parquet |
350,206 | 一次带性别的 PERSON 提及 | stem + 跨度 |
principals |
principals.parquet |
21,895 | 交易的一方当事人 | stem + 跨度 |
persons_distinct |
export/persons_distinct.parquet |
17,362 | 一个独立人物(轻量共指) | person_id |
autonomy |
autonomy.parquet |
32 | 一个世纪/地区分组(衍生) | dimension+bucket |
提取机制分为两类,具有不同的错误特征,不可混合使用:
monetary、prices、taxes:基于确定性词典和规则的提取,对解码后的 EpiDoc 数值进行匹配。封闭词汇(drachma、artaba、wheat)以词典上限进行匹配。精度高,但系统性遗漏。persons、principals、persons_distinct、autonomy:基于训练模型的提取。实测端到端准确率:实体 F1 0.737(严格),预测实体上的PARTY_OFF1 0.623。
表间连接关系
stem:每份文献的唯一键(DDbDP 文件主干),用于人物相关表(documents、persons、principals)。tm_id:Trismegistos 文献 ID,用于货币相关表(monetary、prices、taxes)。tm_id不唯一:61,249 份文献对应 60,862 个不同的 TM ID,连接时可能发散。
documents ──stem──< persons ──stem+span──> principals ──折叠──> persons_distinct documents ──tm_id──< monetary ──筛选子集──> prices, taxes persons ──按世纪/地区聚合──> autonomy
prices 和 taxes 并非独立提取,而是 monetary 应用精度筛选后的子集,粒度相同,列数相同仅多两列。
引用完整性:195,906 行 monetary 中,0 行包含 documents 中不存在的 tm_id;21,895 行 principals 中,0 行包含 documents 中不存在的 stem。
字符跨度溯源
每个事实表都包含起始/结束偏移量,指向文献的标准文本视图(edited_text):
- 货币表:
amount_start/amount_end - 人物表:
person_start/person_end
偏移量为 0 起始、结束位置不包含(Python 切片风格)。DuckDB 的 substr 为 1 起始,切片方式为 substr(edited_text, start + 1, end - start)。
列参考
"cov" 表示非空行占比,反映了对该列进行筛选的成本。
documents — 主干(61,249 行)
每行对应一份文献及其元数据,并包含从其他所有表折叠而来的每份文献计数。
| 列名 | 类型 | cov | 含义 |
|---|---|---|---|
stem |
VARCHAR | 100% | 唯一文献键(DDbDP 文件主干) |
tm_id |
VARCHAR | 100% | Trismegistos ID,不唯一,60,862 个不同值 |
century |
DOUBLE | 97.1% | 基于 HGV 日期的有符号世纪(+2 = 公元2世纪,−2 = 公元前2世纪) |
place_pleiades |
DOUBLE | 74.1% | 文献地点的 Pleiades ID |
deal_type |
VARCHAR | 100% | 主要体裁;当语料库未提供时记为 ?(17,932 份文献) |
n_persons |
BIGINT | 100% | 文献中 PERSON 提及次数 |
n_women_mentions, n_men_mentions |
BIGINT | 100% | 按归属性别统计的提及次数 |
n_principals, n_women_principals |
BIGINT | 100% | 当事人数量,按性别统计 |
has_guardian_woman |
BOOLEAN | 100% | 是否存在带有 μετὰ-/χωρὶς-κυρίου 公式的女性 |
n_money_facts |
BIGINT | 100% | 货币事实数量,按 tm_id 折叠,TM 兄弟共享 |
has_price, has_tax |
BOOLEAN | 100% | 是否存在清洁价格/税款行,同样按 tm_id 共享 |
monetary — 事实表(195,906 行)
每行对应一个带有货币单位的金额,包含其归一化值以及关系图提供的相关信息:所定价的商品、商品数量和单位、以及所支付的税款。
| 列名 | 类型 | cov | 含义 |
|---|---|---|---|
tm_id |
VARCHAR | 100% | 文献键(不唯一) |
amount_start, amount_end |
BIGINT | 100% | 金额的字符跨度(0起始,结束不包含) |
amount_text |
VARCHAR | 100% | 该跨度的希腊语表面字符串 |
value_num |
DOUBLE | 100% | 跨度内解码出的数值 |
currency_id |
VARCHAR | 100% | 规范面额 ID |
system |
VARCHAR | 100% | silver、gold、unknown,切勿跨系统聚合 |
value_base |
DOUBLE | 98.7% | 以德拉克马(白银)或诺米斯马(黄金)计的价值 |
commodity_id |
VARCHAR | 3.9% | 被定价的商品(当存在 HAS_PRICE 链接时) |
quantity |
DOUBLE | 3.4% | 商品数量 |
unit_id |
VARCHAR | 0.3% | 计量单位(artaba、metretes 等) |
unit_price_base |
DOUBLE | 3.3% | value_base / quantity,原始值,存在过度除数问题 |
tax_id |
VARCHAR | 3.4% | 所支付的税款(当存在 CHARGED_UNDER 链接时) |
confidence |
DOUBLE | 100% | 标注器置信度,恒为 0.82,不携带信息 |
date_lo |
DOUBLE | 93.4% | HGV 日期范围的最早年份(有符号,负值为公元前) |
date_hi |
DOUBLE | 78.3% | 日期范围的最近年份 |
date_mid |
DOUBLE | 94.3% | 范围中点,century 由此推导 |
century |
DOUBLE | 94.3% | 有符号世纪,范围 −4 到 +10 |
bin50 |
DOUBLE | 94.3% | 50年分档的起始年份,向下取整(−124 → −150) |
place_pleiades |
DOUBLE | 80.7% | Pleiades ID |
genres |
VARCHAR | 100% | 以字符串形式存储的 JSON 数组,例如 ["list","account"] |
低 commodity_id/tax_id 覆盖率符合预期:纸莎草中的大部分金额为裸金额(收据总额、工资、租金),无附加信息。具有商品的 3.9% 正是价格序列的来源。
prices — 清洁价格观测(98 行)
从 monetary 中筛选出的真实、可比的单位价格。筛选条件排除了 value_num == quantity 的双重链接伪影(占原始候选的 48%)、青铜 chalkous、非商品自身干/液计量单位的单位,以及不合理的数量/价格组合。包含所有 monetary 列,额外增加:
| 列名 | 类型 | cov | 含义 |
|---|---|---|---|
unit_price |
DOUBLE | 100% | 每单位德拉克马数,序列值 |
commodity |
VARCHAR | 100% | wheat (70)、wine (14)、barley (11)、oil (3) |
taxes — 清洁税款支付(592 行)
从 monetary 中筛选出带有指定税名的行。这些是支付额(通常为分期付款),而非税率。比 prices 更清洁,因为不涉及单位除法。包含所有 monetary 列,额外增加:
| 列名 | 类型 | cov | 含义 |
|---|---|---|---|
payment |
DOUBLE | 100% | 以德拉克马计的支付额 |
tax |
VARCHAR | 100% | laographia (539,人头税)、demosia (53,土地税) |
persons — 带性别的人物提及(350,206 行)
NER 模型找到的每一个 PERSON 跨度的性别和监护人状态。每行为一次提及,而非一个人物。
| 列名 | 类型 | cov | 含义 |
|---|---|---|---|
stem, tm_id |
VARCHAR | 100% | 文献键 |
person_start, person_end |
BIGINT | 100% | 提及的字符跨度 |
person_text |
VARCHAR | 100% | 完整的提及文本(“名字 某人之子 …”) |
head_text |
VARCHAR | 100% | 人物自身的名字(从整体中分离) |
father_text |
VARCHAR | 36.8% | 父名(若存在) |
gender |
VARCHAR | 100% | male (82,817)、female (22,901)、unknown (244,488) |
gender_basis |
VARCHAR | 100% | 决定性别的规则 |
gender_confidence |
DOUBLE | 100% | 0.0(未知)到 0.97(监护人公式) |
guardian |
VARCHAR | 100% | with (1,628)、without (143)、none (348,435) |
date_mid, century, bin50 |
DOUBLE | 95.9% | 文献日期 |
place_pleiades |
DOUBLE | 81.7% | Pleiades ID |
genres |
VARCHAR | 100% | JSON 数组字符串 |
仅 30% 的提及可归属性别,此为设计使然:规则仅在名称或公式具有决定性时触发,并记录触发的规则而非猜测。不进行任何插补。
principals — 带性别和交易类型的当事人(21,895 行)
交易的核心人物:关系模型链接为 PARTY_OF 交易或 PAID_BY/PAID_TO 金额的 PERSON,并与 persons 中的性别/监护人/父名信息关联。覆盖 11,002 份文献。包含上述所有 persons 列(其中 father_text 覆盖率为 53.9%,因为合同中的当事人命名比较正式),额外增加:
| 列名 | 类型 | cov | 含义 |
|---|---|---|---|
roles |
VARCHAR | 100% | 排序、` |
transaction_term |
VARCHAR | 68.9% | 命名该交易的希腊语动词/名词 |
deal_type |
VARCHAR | 100% | 文献的主要体裁,3,514 行为 ? |
confidence |
DOUBLE | 100% | 关系模型得分,0.34 到 1.00(中位数 0.85),有信息量,应基于此筛选 |
persons_distinct — 共指轻量人物(17,362 行)
将当事人提及折叠为独立人物,基于保守的表面键(归一化名字、归一化父名、地点)。回答“有多少个独立女性”而非“有多少次提及”。存在欠合并问题——未附父亲名字或出现在不同地区的人物会被拆分为多行,因此计数为上限,适合“不少于”的声明。非完整的古谱学共指。
| 列名 | 类型 | cov | 含义 |
|---|---|---|---|
person_id |
VARCHAR | 100% | 身份键的稳定 16 位十六进制哈希值,跨重建版本唯一 |
head_text |
VARCHAR | 100% | 代表性名字 |
father_text |
VARCHAR | 57.7% | 代表性父名 |
place_pleiades |
DOUBLE | 85.0% | 代表性地点 |
gender |
VARCHAR | 100% | 跨提及折叠(属性优于未知,多数胜出):female 1,414、male 5,608、unknown 10,340 |
guardian |
VARCHAR | 100% | 折叠方式:without 优先,一次明确的 χωρὶς-κυρίου 证明即认定其独立行动 |
n_mentions |
BIGINT | 100% | 折叠入此人物的提及次数(最多 32 次) |
deal_types |
VARCHAR | 100% | ` |
first_century |
DOUBLE | 98.5% | 首次出现的世纪 |
autonomy — 已发布曲线(32 行)
衍生汇总表,非事实表:χωρὶς-κυρίου 在带有监护人公式的女性中的占比,按世纪和地区分组。提供此表以使主要发现可核查。如需按地区、交易类型或其他时间粒度重新切片,请自行对 principals 进行分组。
| 列名 | 类型 | cov | 含义 |
|---|---|---|---|
dimension |
VARCHAR | 100% | century(8 行)或 region(24 行) |
bucket |
VARCHAR | 100% | 有符号世纪的字符串形式(如 "3.0")或地名 |
n_with, n_without |
BIGINT | 100% | 有/无监护人证明的女性人数 |
n |
BIGINT | 100% | n_with + n_without |
autonomous_share |
DOUBLE | 100% | n_without / n |
受控词汇表
ID 为规范词典 ID,非表面形式——标注器在写入任何内容之前已将屈折变化的希腊语解析为这些 ID。计数来自 monetary(除非另有说明)。
system:silver 140,040(托勒密-罗马德拉克马体系,基本单位为德拉克马)· gold 54,771(拜占庭索里达体系,基本单位为诺米斯马)· unknown 1,095(无固定面额的货币词)。
currency_id:白银阶梯(1 塔兰特 = 6,000 德拉克马;1 德拉克马 = 6 奥波尔;1 奥波尔 = 8 卡尔库斯):drachma 65,672 · obol 23,253 · talent 12,513 · diobol 7,317 · triobol 7,099 · chalkous 6,745 · hemiobelion 6,680 · tetrobol 6,346 · pentobol 3,422 · argyrion 993(通用“银钱”,无面额 → value_base 为空)。黄金(24 克拉 = 1 诺米斯马):nomisma 36,168 · keration 18,047 · chrysion 556(通用黄金,包括作为金属而非仅硬币的黄金)。
commodity_id:grain 2,098 · garden 1,066 · wheat 1,057 · wine 758 · oil 673 · barley 429 · donkey 395 · hay 212 · vegetables 161 · land 149 · house 148 · camel 139 · sheep 139(+长尾)。
unit_id:artaba 399(干量)· aroura 59(土地)· metretes 34 和 keramion 28(液量)· xestes 26 · kotyle 19 · myriad 18 · choinix 10 · litra 10 · naubion 7 · pechys 6。
tax_id:prosdiagraphomena 2,788(罗马附加费)· demosia 1,722(土地税)· phoros 700 · laographia 574(人头税)· merismos 349 · phylakitikon 254(托勒密警察税)· telesma 110 · stephanikon 40 · genema 36 · syntaxis 17。该字段中的少量 ID 并非税收——drachma 20, obol 3, hemiobelion 2, chalkous, chrysion, aroura, year, time——共 8 个 ID / 33 条带日期的事实(0.5%),为测得的污染率。建议筛选上述命名的税收,而非接受每个非空的 tax_id。
deal_type(在 documents 中):? 17,932 · receipt 15,194 · contract 5,117 · list 4,193 · letter_private 3,603 · account 3,342 · mummy_label 2,212 · order 2,146 · petition 1,908 · letter 1,724 · letter_official 1,228 · declaration 734 · register 643 · delivery 417 · sale 318(+ loan、lease 等长尾)。? 表示语料库未记录体裁,而非“其他”——应排除而非归入。
gender_basis(在 persons 中,按精度排序):guardian 1,770(κύριος 公式,置信度 0.97)· nomen 21,864(Αὐρήλιος / Αὐρηλία,0.9)· kin 5,781(θυγάτηρ / υἱός,0.9)· gazetteer 31,153(已知姓名列表,0.8)· egypt_prefix 45,122(埃及语 Τα- 女性 / Πα- 男性,0.72)· ethnic 28 · none 244,488(无规则触发 → unknown)。
guardian:with — 带有 μετὰ κυρίου 公式(她在监护下交易)· without — χωρὶς κυρίου(她单独交易)· none — 在窗口内无公式。
roles(在 principals 中):party 13,738 · payee 3,605 · payer 3,176 · party|payee 1,111 · party|payer 227 · payee|payer 32 · party|payee|payer 6。多值情况下,一人可同时为合同当事人和支付者。应使用 LIKE %payer% 或 list_contains(str_split(roles,|),payer) 匹配,切勿使用 =。
数据获取
python from datasets import load_dataset
prices = load_dataset("ainouche-abderahmane/oikonomia-db", "prices", split="train")
所有上述表名均为有效配置:monetary、prices、taxes、persons、principals、autonomy、documents、persons_distinct。
下载整个数据库(14 MB)为本地 Parquet 文件:
bash hf download ainouche-abderahmane/oikonomia-db --repo-type dataset --local-dir oikonomia-db
python from huggingface_hub import snapshot_download path = snapshot_download("ainouche-abderahmane/oikonomia-db", repo_type="dataset")
Quickstart (DuckDB)
sql -- 女性在交易当事人中的占比,按交易类型分组。 SELECT deal_type, count() AS n_gendered, round(100.0 * sum(gender = female) / count(), 1) AS pct_women FROM principals.parquet WHERE gender <> unknown AND deal_type <> ? GROUP BY 1 HAVING count(*) >= 40 ORDER BY pct_women DESC;
┌─────────────────┬────────────┬───────────┐ │ deal_type │ n_gendered │ pct_women │ ├─────────────────┼────────────┼───────────┤ │ sale │ 191 │ 30.4 │ │ loan │ 144 │ 28.5 │ │ contract │ 3708 │ 23.0 │ │ receipt │ 2200 │ 10.2 │ │ delivery │ 59 │ 5.1 │ └─────────────────┴────────────┴───────────┘ (节选:完整结果为 15 行)
sql -- 货币化转型:黄金在带日期的货币事实中的占比,按世纪分组。 SELECT century, count() AS n, round(sum(system = gold) * 1.0 / count(), 3) AS gold_share FROM monetary.parquet WHERE century IS NOT NULL AND system IN (gold,silver,bronze) GROUP BY 1 ORDER BY 1;
用途
直接用途
适用于定量古代经济史和数字纸莎草学:价格与工资序列、财政史、性别与法律能力、古谱学、地区比较。由于每行都包含源偏移量,也可作为希腊语信息提取的跨度链接评估集。
超出范围的使用
- 非语言模型训练语料库。包含的是提取的事实,而非文本。
- 非绝对经济总量来源。生存偏差无法修正(见局限性部分);分组内的份额可解释,跨分组的原始计数不可解释。
- 非支付方向数据集。故意不包含谁支付给谁——关系模型在此方向得分为 F1 0.145,过低无法发布。
- 不能替代阅读纸莎草原文。每行提供了可回溯核对的跨度;如需对单份文献做出论断,请核对原文。





