Measuring the Confidence Architecture of Indo-European Claims: findings, code and reproduction materials.
收藏资源简介:
This deposit contains the complete materials of a calibration study of Indo-European historical linguistics: twenty-three research findings, fourteen executable tests with their outputs, an article draft, and an essay on the methodological fallacy registry the project accumulated about itself. The study does not attempt to prove or refute Indo-European reconstruction. It asks a narrower and more answerable question: which specific claims are supported by what kind of evidence, and how far does the confidence in each actually extend. Criteria were declared in advance, in the headers of the scripts, before those scripts were run. Predictions that failed remain visible in the archive rather than being removed. Principal findings. The limit of reconstruction is coverage, not accuracy. Across four families with an attested ancestor — Old Church Slavonic at about 1,075 years of depth, Latin at 2,075, Ancient Greek at 2,475, Vedic at 3,225 — reconstruction accuracy correlates with time depth at r = −0.970. Two reference levels computed from the daughters alone also fall with depth: the best single daughter (r = −0.893) and union coverage, the share of ancestral segments found identically in at least one daughter (r = −0.772), which no reconstruction exceeds (0 of 340 items). The gap between reconstruction and either level is only weakly related to depth (r = −0.440 and −0.419); the fall in surviving material accounts for roughly two-thirds to four-fifths of the fall in accuracy. Version 1.1 corrects the earlier description of the best single daughter as an upper bound: reconstruction exceeds it in 10 of 340 items. What was lost in every daughter is not uncertain but out of range. The comparative method recovers the spoken proto-language, not the written one. Reconstruction from Romance daughters scores 19.0 points higher against Proto-Romance than against Classical Latin, p < 0.0001, replicated on two independent samples and robust at +11.7 points after vowel quantity is removed from the comparison. Some layers are unrecoverable and identifiable as such in advance. Vowel quantity was recovered in 0 of 26 test items. Evidential value is a property of the pair (evidence × method), not of the evidence alone. On identical data, root recovery ranged from 0% under one criterion to near-perfect under another. Negative results. Two of the tools built during the project did not survive testing and are reported as such. A weighting scheme combining proximity and independence failed three replication attempts. A predictor linking homeland-claim fragility to the prior population density of the receiving territory was downgraded to one-directional after Iceland activated a pre-registered falsification condition. A daughter-dispersion predictor proved saturated at Proto-Indo-European depth, where 92% of etyma sit above the usable threshold. Methodological fallacy registry. The archive includes a register of twenty-four methodological fallacies: seven drawn from the literature and seventeen discovered inside the project itself, fifteen of which are charged to the AI assistant used in the work. Seven formal retractions are recorded. None was a calculation error; most were framing errors, and the last three were errors of reporting. The accompanying essay argues that the sequence of project synopses documents how a caveat weakens across retellings, which is why all thirteen synopses are retained rather than only the final one. Contents. findings/ twenty-three research findings (Greek); synopses/ thirteen successive project synopses (Greek); paper/ an article draft (English) and the fallacy-registry essay (Greek); code/ fourteen Python scripts and a data-retrieval script; results/ the output of all fourteen, plus three verification logs. Reproduction. The two source corpora are not redistributed here. They are retrieved by code/fetch_data.sh from their own repositories at the commits pinned in the README: lexibank/saenkoromance at 10320bb6 and lexibank/iecor at 700b635a. All eleven scripts were re-executed independently on a separate machine and all eleven outputs matched byte for byte; the log, including the resolved commit hashes, is in results/VERIFICATION.txt. The two scripts added in version 1.1 were run on Linux and re-run by the author on macOS; both outputs matched byte for byte (results/VERIFICATION_v1.1.txt). The script added in version 1.2 was run on Linux and re-run by the author on macOS; both outputs matched byte for byte (results/VERIFICATION_v1.2.txt). Stated limits. The reconstruction algorithm is a substitute for a specialist's judgement, so the reported percentages are lower bounds. Anatolian and Albanian are absent from the etymon reliability map. A fifth depth point is missing. Learned borrowing (Sanskrit tatsama) is controlled only as far as IE-CoR flags it; re-inserting the flagged loans changes Indic accuracy by −0.2 points. The four attested ancestors are near-ancestors of their daughters, so the share of the accuracy decline attributed to corpus decay also contains each target's distance from the true common ancestor. Version 1.1 adds finding #22 and tests 14–15 (union coverage; loan sensitivity) and revises the article accordingly. Version 1.2 adds finding #23 and test 16. The four attested ancestors are near-ancestors rather than exact common ancestors of their daughters. Where both targets exist (Romance), target mismatch is absorbed by the daughters-only reference levels (union coverage −24.3 points, accuracy −24.0) and not by the gap between method and material (+0.3). The part of the decomposition that measures the method is unaffected; the share of the decline attributed to corpus decay also contains the distance between each target and the true common ancestor. Four further fallacies, all attributed to the AI system, are recorded in a postscript to the essay. AI contribution. This work was carried out in extended dialogue with a large language model, which wrote the analysis scripts, drafted the findings documents, and produced the structural verdicts recorded throughout. All pre-registrations, disagreements and retractions were reviewed by the author, who is responsible for the content. The disagreements the method could not settle are marked in the documents and left open. The fallacy registry records which errors originated with the model.



