Minimal Reproducibility Package for "Computational Syntactic Clustering of Linear A"
收藏资源简介:
Title:Minimal Reproducibility Package for “Computational Syntactic Clustering of Linear A” Description:This archive contains the exact numerical values that appear in the main-text tables of the manuscript Computational Syntactic Clustering of Linear A submitted to PNAS.No synthetic, inferred, or reconstructed values are included. The dataset provides the final outputs used in the manuscript’s illustrative examples, including: 12-dimensional archetype fingerprints (A–L) for the inscriptions HT6b, HT2, HT3, ZA20, and KN28 (from Table 2). Macro-family assignments and probabilities for the same inscriptions (from the Results section). Pairwise L1 distances and cluster IDs for representative tablets HT6b, ZA20, KN28, and HT3 (from Table 3). These files are intended as a minimal transparency companion to the manuscript, allowing readers and reviewers to verify all numerical values that appear in the published tables.They do not constitute the full computational pipeline, tokenization system, or clustering model described in the Methods; rather, they serve as an archival record of the specific results presented in the manuscript. A lightweight Jupyter notebook (pipeline.ipynb) documents the structure and contents of the CSV files and provides a simple loading and inspection interface.This notebook is descriptive and is not a reimplementation of the underlying analysis pipeline. Contents of this archive: data/fingerprints.csv — 12D archetype fingerprints data/macro_families.csv — macro-family labels + probabilities data/distances.csv — pairwise L1 distances and cluster IDs code/pipeline.ipynb — minimal transparency notebook README_FOR_PNAS.txt — description of files and reproducibility notes All values in this deposit are directly extracted from the manuscript without modification.



