遇见数据集

Voynich Manuscript 16-Dimensional Visual Semantic Profiles and Statistical Artefacts

收藏
Zenodo2026-04-13 更新2026-05-26 收录
官方服务:

资源简介:

Per-page visual semantic profiles and section-level statistical artefacts for all 206 pages of the Voynich Manuscript (Yale Beinecke MS 408), produced as part of the xenoglyph project and released as the companion dataset to the preprint "Visual Semantic Profiling of the Voynich Manuscript: Reading Meaning from Illustrations in an Undeciphered Codex" (Lyons, 2026). This dataset contains: (1) The complete per-page 16-dimensional profile vectors through three independent lenses — a medieval-codex (Voynich) lens, a cross-cultural archaeological lens, and a hermetic cryptological lens.(2) Section-level discrimination statistics (one-way ANOVA, Welch's robust ANOVA, Kruskal-Wallis, eta-squared effect sizes) for each lens.(3) A dimension-discovery run that identifies the principal axes of variation without using human-authored descriptors.(4) An ISP (Iterative Semantic Profiling) document-level analysis that computes a semantic arc across the manuscript.(5) All 17 JSON statistical artefacts consumed by the preprint's figure-building pipeline: ANOVA tables, classifier results for the 16-d voynich lens, a head-to-head 16-d / 768-d / text-ngram / layout comparison, UMAP seed-stability data, cross-section similarity matrices, and per-corpus out-of-distribution profile means.(6) Corpus metadata: folio, section label, and illustration type for every page. The profile vectors and section-level statistics are fully sufficient to reproduce every figure and every classifier result in the preprint without any access to the production xenoglyph profiling pipeline. The scripts required to do so are distributed in the companion repository (https://github.com/datasquatch8144/xenoglyph-voynich-public). The production profile generation pipeline — the step that turns a page image into a 16-d vector — is covered by a pending United States provisional patent application filed with the USPTO in March 2026, and the specific foundation model identity and exact text of the dimension descriptors are withheld per the patent application. Researchers who wish to regenerate the profile vectors with an alternative vision-language model may do so: the corpus metadata, section labels, and source image references are sufficient. The source images are in the public domain and available from Yale's Beinecke Rare Book and Manuscript Library via IIIF at https://collections.library.yale.edu/manifests/2002046.

提供机构:
Zenodo
创建时间:
2026-04-13
二维码
社区交流群
二维码
科研交流群
商业服务