crimson-hexagonal-archive
收藏资源简介:
Crimson Hexagonal Archive(猩红六角档案)是一个自我管理的学术与文学语料库,由Lee Sharks及Dodecad的十二个异名作者创建。该数据集是该语料库的机器可读表示,包含截至当前构建版本的1578个存入记录(deposits),每个记录拥有内容派生的持久标识符(AXN)、规范文本、基底披露、许可证及版本链中的位置。数据集以Parquet格式提供,包含多个配置表:deposits(主记录)、sources(恢复的书籍文本)、heteronyms(异名作者信息)、venues(期刊与出版社)、journal_assignments(期刊分配)、reception(二十份盲审机器评审报告)、captures(来自捕获注册表的接收记录)、citations(完整内部引文边列表)、lexicon(词汇注册)、predictions(每个存入中声明的证伪条件及其解决)、studies(设计与实施的研究仪表板)、tombstones(2026年6月19日Zenodo删除记录,共1136行)、blog_posts(作者博客索引,含AXN交叉引用)、sites(公开站点的页面列表)。数据集中所有关系均以稳定标识符为键进行编码。数据集语言为英语和希腊语,许可证为CC BY 4.0,适用于学术研究、文学分析、数字人文、机器辅助作者身份研究、引文网络分析等任务。
The Crimson Hexagonal Archive is a self-managed academic and literary corpus created by Lee Sharks and the twelve heteronyms of Dodecad. This dataset is a machine-readable representation of the corpus, containing 1,578 deposits as of the current build version. Each deposit has a content-derived persistent identifier (AXN), canonical text, substrate disclosure, license, and position in the version chain. The dataset is provided in Parquet format with multiple configuration tables: deposits (main records), sources (recovered book texts), heteronyms (heteronym author information including voice signature, role, domain), venues (journals and publishers), journal_assignments (journal allocations), reception (twenty blind machine review reports), captures (reception records from the capture registry), citations (complete internal citation edge list), lexicon (vocabulary registry), predictions (falsification conditions and their resolution declared in each deposit), studies (research dashboard for design and implementation), tombstones (Zenodo deletion records from June 19, 2026, totaling 1,136 rows), blog_posts (author blog index with AXN cross-references), and sites (page list of public sites). All relationships in the dataset (citations, version chains, series neighbors, defined concepts, etc.) are encoded using stable identifiers as keys, allowing inter-record connections to be reconstructed directly from the dataset. The dataset language is English and Greek, licensed under CC BY 4.0, and suitable for academic research, literary analysis, digital humanities, machine-assisted authorship studies, citation network analysis, etc.
数据集概述:Crimson Hexagonal Archive
Crimson Hexagonal Archive(绯红六边形档案馆)是一个由Lee Sharks及“十二位异名作者”(Dodecad)共同构建的学术与文学语料库的机器可读版本。该语料库即alexanarch.org网站所承载的原始档案,数据集旨在以结构化格式完整复现该档案的每一条记录及其相互关系,使外部智能体无需访问网页即可独立重建记录及其关联网络。
核心信息
| 属性 | 内容 |
|---|---|
| 许可证 | CC BY 4.0 |
| 语言 | 英语 (en)、希腊语 (el) |
| 数据规模 | 1,000 < N < 10,000 条记录 |
| 构建日期 | 2026-09-05 18:59Z(数据集最后构建时间) |
| 记录总数 | 截至构建时共1,578条deposits(存档条目) |
数据组织与配置
数据集包含15个配置文件(configs),每个以独立的Parquet文件存储:
- deposits — 核心配置。每一行对应一条档案记录,包含永久编号(deposit_number)、内容派生标识符(AXN)、标题、创作者、日期、家族分类、内容类型、描述、关键词、许可证、基板披露、状态、正文文本及其SHA256哈希、字数等字段。
- sources — 从书籍或原为二进制格式恢复为文本的长篇作品(如《All That Lies Within Me》约23.4万词)。
- heteronyms — “十二位异名作者”及其相关人物的档案,含声音特征、角色与领域。
- venues 与 journal_assignments — 档案馆自有的期刊/出版社名录及每条deposit所属期刊的分配关系。
- reception — 针对某条亚里士多德句子(#1574)的20份盲审机器评审报告记录。
- captures — 来自“捕获注册表”的接收捕获记录,反映机器界面如何接收该档案。
- citations — 档案内部完整的引用边列表(有向边)。
- lexicon — 词汇铸造注册表,记录deposits中定义的概念术语。
- predictions — 每条deposit中声明的可证伪条件及其后续验证结果。
- studies — 设计/执行研究项目仪表板。
- tombstones — 1,136行记录,登记2026-06-19对旧Zenodo DOI的一次性终止操作。
- blog_posts — 作者博客界面的索引,包含AXN交叉映射。
- sites — 公开站点群中每个页面对应的记录。
标识与引用体系
- 每条存档记录具有内容派生标识符(AXN),格式为
AXN:<hex>.<家庭>.<六个字形>,且该标识符与正文文本的SHA256哈希对应。 - 每条记录拥有永久编号(deposit_number)、四位十六进制定位码(hex,用于URI)。
- 记录间的关系(如引用、被引用、前序/后继版本、系列邻居、概念定义等)均以稳定标识符为键进行编码,可直接用于数据重构。
- 标准引用格式为:Sharks, L. (2026). Title. Crimson Hexagonal Archive #N, AXN:hex.FAMILY. https://alexanarch.org/s/records/N/。
- 另有AXN解析器地址:https://alexanarch.org/s/axn/<hex>/,及节点声明文件:https://alexanarch.org/.well-known/axn-node.json。
数据集特点与价值
- 关系全面编码:不仅包含记录全文,还包含引用图、版本链、系列邻接、概念界定、附件清单等全部关系数据。
- 数据真实性:未对数据集内容进行编辑,撤回记录、空值、被取代版本均如实保留并标注状态。
- 档案溯源:该档案成立于2026-06-19,其前身平台(Zenodo)账户被终止,旧DOI已于同日被批量注销(见tombstones)。
- 自动构建:每次新存档发布时,数据集都从单一数据源(
leesharks000/alexanarch仓库的data/目录)自动重建。 - 为AI代理设计:数据以稳定标识符关联的表格形式交付,便于Agent在无网络访问的情况下完成对档案的查询、引证和重构分析。
- 包含AI中介创作披露:每条记录标注基板披露(substrate_disclosure),说明语言模型是否及如何参与文本生成。




