遇见数据集

Data, code and results for "What can linguistic dating detect? A control-corpus validation, with an application to Biblical Hebrew"

收藏
Zenodo2026-08-14 更新2026-08-20 收录
官方服务:

资源简介:

Texts are routinely dated by their language, but the methods used to do so have never been calibrated against a corpus whose dates are known independently. This deposit contains the data, analysis code and complete results for a study that supplies that calibration and applies it to Biblical Hebrew. The design turns on a control corpus. Every procedure applied to Hebrew is first applied unchanged to 63 ancient Greek prose texts whose dates are fixed by biography, synchronism and historical record. Greek supplies what Hebrew cannot: a case where a diachronic signal demonstrably exists and its magnitude can be measured. In Greek, chronology is recoverable within a single genre, the real effect measures about 0.70 standardized units per century, and permuting the anchors within genre never reproduces the fit in 500 permutations. In Hebrew, no within-genre chronological signal is recoverable at any level of anchor security, while the same features recover source-critical divisions established on other grounds at 0.87 to 0.94 accuracy — so the null result concerns the chronological signal rather than the method. Injected signals of known amplitude convert that null into a bound. The detectability floor is 0.50 standardized units per century for Greek and 1.50 for Hebrew prophecy, against an effect of 0.59 actually present, and the floor is recomputed across five unrelated estimator families to establish that it is a property of the corpus rather than of any modelling choice. Fitted power curves put the corpus needed for 80% detection at 19 anchored texts for Greek and 79 for Hebrew, against the 18 that exist. Contents: the Greek feature table (6,447 passages of 500 words, 748 features) and both corpus manifests; the analysis scripts; every results file the article cites; and the article with its supporting information. The derived Hebrew feature table is not redistributed because the ETCBC/BHSA corpus it derives from is CC BY-NC 4.0; the extraction code and a validation script that checks the result against published concordance counts are included instead, and regenerate it exactly. See README.md.

提供机构:
Zenodo
创建时间:
2026-08-14
二维码
社区交流群
二维码
科研交流群
商业服务