19th century American literary orthovariant tokens
收藏资源简介:
该数据集名为“19世纪美国文学正字变体标记数据集”,由约翰斯·霍普金斯大学的数字人文中心创建。数据集包含4032个正字变体标记,这些标记带有新颖的人工注释方言组标签,旨在支持计算实验,探索文学上有意义的正字变异。数据集的创建过程包括从Project Gutenberg语料库中提取19世纪美国文学子集,并由作者根据说话角色的作者意图位置分配方言标签。该数据集主要应用于语言建模和文学分析领域,旨在解决文学正字变异对语言模型的影响问题。
This dataset, titled the Orthographic Variant Annotation Dataset of 19th-Century American Literature, was developed by the Digital Humanities Center at Johns Hopkins University. It comprises 4,032 orthographic variant annotations paired with novel manually annotated dialect group labels, intended to support computational experiments investigating literary-significant orthographic variation. The dataset creation workflow involves extracting a 19th-century American literature subset from the Project Gutenberg corpus, with dialect labels assigned based on the authorial intent associated with each speaking character. This dataset is primarily applied in the domains of language modeling and literary analysis, aiming to address the impact of literary orthographic variation on language models.

- 1Examining Language Modeling Assumptions Using an Annotated Literary Dialect Corpus约翰斯·霍普金斯大学 · 2024年



