遇见数据集

Etruscan Machine Learning Corpus

收藏
Zenodo2026-08-08 更新2026-08-13 收录
官方服务:

资源简介:

The OpenEtruscan ML-Ready Corpus is a normalized, quality-tagged dataset of 6,567 Etruscan inscriptions for machine-learning tasks: word and character embedding training, lacuna restoration, glyph recognition, and diachronic analysis. Version 1.1 ships two files: openetruscan_clean.csv: the 10-column corpus, byte-identical to version 1.0 (SHA256 4fc09af94005655bfe26affeeb48295c88606ae23c8dbc33ff5436f9083f69f8). Columns: id, raw_text, canonical_transliterated, canonical_italic, canonical_words_only, translation, year_from, year_to, intact_token_ratio, data_quality. openetruscan_clean_grouped.csv: the same 6,567 rows plus dup_group_id (first 12 hex chars of the SHA-256 of the text after Leiden editorial markup is stripped, whitespace collapsed, and case folded) and dup_group_size (rows sharing that normalized text). Why the new columns: 470 rows repeat a canonical_transliterated value under a different id, because short formulaic inscriptions (mi, suθina, aplu) genuinely recur across distinct artifacts; 626 excess rows share a group once Leiden variants such as la(u)tni / lautn(i) / laut(n)i are normalized together. The repeats are correct data and the corpus is deliberately not deduplicated, but they mean row-level random train/test splits leak: the same text lands on both sides under different ids. This contaminated the corpus maintainers' own frozen classification split (25 of 400 test rows; see PRE_REGISTRATION.md Deviation D in the code repository). Split on dup_group_id, not on rows. Provenance: roughly 71% of rows derive from the Larth Dataset (Vico and Spanakis, 2023, CC-BY-4.0) and 29% from Corpus Inscriptionum Etruscarum Vol. I extractions (public domain). The cleaning pipeline, normalized columns, quality tags, and group columns are released under CC-BY-4.0. Schema documentation, editorial conventions, quality tiers, and known limitations (including what the quality tiers do not filter) are maintained in research/data/README.md of the code repository: https://github.com/Eddy1919/openEtruscan

提供机构:
Zenodo
创建时间:
2026-08-08
二维码
社区交流群
二维码
科研交流群
商业服务