Vigeland Erindringer (1918): a paragraph-aligned Norwegian–English parallel corpus with editorial apparatus / avsnittsjustert NO–EN parallellkorpus med redaksjonelt apparat
收藏资源简介:
Version 2.0.1. Two documentation files in v2.0 stated that the licence was unresolved, contradicting the record's own licence field. Corrected here. The data files are byte-identical to v2.0. 543 paragraphs of Gustav Vigeland's Erindringer (1918), each with its Norwegian source and an English translation side by side. Two source surfaces (raw diplomatic transcription and modernised reading stream) and three independent production runs under arm 3 (masked reinsertion) across two vendors — Anthropic on both surfaces, Google on the raw surface — delivered as three JSONL files. Apparatus elements — lacunae, editorial blocks, marginal notes, uncertain readings, quotations — are carried as a separate layer with TEI types, offsets and anchors. Nine records across the three runs carry no English text: generation was interrupted by the vendor's own classifier. They are marked as interrupted rather than dropped, because which paragraphs get interrupted is one of the findings. Three layers, kept apart on purpose. Layer 1: the two surfaces, 543 paragraphs each, untouched. Layer 2: the alignment between them, replaceable without rebuilding layer 1. Layer 3: analysis sets, one file per analysis, n stated explicitly. A coincidence of counts between two sets is never proof of correspondence. A canary string is embedded in every record so that later contamination is checkable. A held-out evaluation split of 55 paragraphs is defined but its membership is not disclosed: the split field is null in every record. The source text of those paragraphs necessarily ships with the other 543; what is withheld is which ones they are, and the human reference translations once they exist. A published evaluation set is a contaminated evaluation set.



