遇见数据集

Minimal Reproducibility Package for "Computational Syntactic Clustering of Linear A"

收藏
Zenodo2025-12-08 更新2026-05-26 收录
官方服务:

资源简介:

Title:Minimal Reproducibility Package for “Computational Syntactic Clustering of Linear A” Description:This archive contains the exact numerical values that appear in the main-text tables of the manuscript Computational Syntactic Clustering of Linear A submitted to PNAS.No synthetic, inferred, or reconstructed values are included. The dataset provides the final outputs used in the manuscript’s illustrative examples, including: 12-dimensional archetype fingerprints (A–L) for the inscriptions HT6b, HT2, HT3, ZA20, and KN28 (from Table 2). Macro-family assignments and probabilities for the same inscriptions (from the Results section). Pairwise L1 distances and cluster IDs for representative tablets HT6b, ZA20, KN28, and HT3 (from Table 3). These files are intended as a minimal transparency companion to the manuscript, allowing readers and reviewers to verify all numerical values that appear in the published tables.They do not constitute the full computational pipeline, tokenization system, or clustering model described in the Methods; rather, they serve as an archival record of the specific results presented in the manuscript. A lightweight Jupyter notebook (pipeline.ipynb) documents the structure and contents of the CSV files and provides a simple loading and inspection interface.This notebook is descriptive and is not a reimplementation of the underlying analysis pipeline. Contents of this archive: data/fingerprints.csv — 12D archetype fingerprints data/macro_families.csv — macro-family labels + probabilities data/distances.csv — pairwise L1 distances and cluster IDs code/pipeline.ipynb — minimal transparency notebook README_FOR_PNAS.txt — description of files and reproducibility notes All values in this deposit are directly extracted from the manuscript without modification.

标题:《线形文字A的计算句法聚类》最小可复现数据包 描述:本归档文件包含提交至《美国国家科学院院刊》(PNAS)的手稿《线形文字A的计算句法聚类》正文中表格所使用的精确数值。本数据包未包含任何合成、推断或重构得到的数值。 本数据集提供了手稿示例章节所用的最终输出结果,具体包括: 1. 来自表2的铭文HT6b、HT2、HT3、ZA20与KN28的12维原型指纹(A–L) 2. 同批铭文的宏家族分类结果与对应概率(来自结果章节) 3. 代表性泥板HT6b、ZA20、KN28与HT3的两两L1距离及簇ID(来自表3) 本数据包作为手稿的极简透明配套文件,可供读者与审稿人验证已发表表格中的全部数值。本数据包并未包含方法章节中所述的完整计算流程、分词系统或聚类模型,仅作为手稿中呈现的特定结果的存档记录。 本数据包附带一个轻量级Jupyter 笔记本(pipeline.ipynb),用于说明CSV文件的结构与内容,并提供简易的加载与检视接口。该笔记本仅为说明性文件,并未重新实现底层的分析流程。 本归档文件包含以下内容: - data/fingerprints.csv — 12维原型指纹 - data/macro_families.csv — 宏家族标签与对应概率 - data/distances.csv — 两两L1距离与簇ID - code/pipeline.ipynb — 极简透明说明笔记本 - README_FOR_PNAS.txt — 文件说明与可复现性备注 本存档中的所有数值均直接从手稿中提取,未做任何修改。

提供机构:
Zenodo
创建时间:
2025-12-08
二维码
社区交流群
二维码
科研交流群
商业服务