遇见数据集

Common Library 1.0

收藏
arXiv2019-09-06 更新2024-08-06 收录
数据链接:
官方服务:

资源简介:

Common Library 1.0是一个包含75部维多利亚时期小说的语料库,由研究团队创建,旨在反映出版年份和作者性别的多样性。该数据集从15,322条书目中抽样,涵盖1837至1901年间在英国出版的小说。数据集的特色在于其内部小说与更广泛的小说出版群体在社会学重要子群体中的分布比例相匹配。例如,1880年代女性作者的小说在数据集中的比例与整体出版群体中的比例相近。数据集主要用于支持文学历史和语料库语言学的研究,特别是在分析作者社会经济背景与作品内容或风格之间的关系,以及文学技巧随时间和空间的变化而变化的研究。

Common Library 1.0 is a corpus consisting of 75 Victorian novels, developed by a research team to reflect the diversity of publication years and author genders. This dataset was sampled from 15,322 book titles, covering novels published in the United Kingdom between 1837 and 1901. A distinctive feature of the dataset is that its distribution of novels across sociologically salient sub-groups matches that of the wider published novel population. For example, the proportion of novels by female authors published in the 1880s in the dataset is nearly identical to their share in the overall publishing cohort. The dataset is primarily intended to support research in literary history and corpus linguistics, especially studies investigating the relationship between authors' socioeconomic backgrounds and their works' content or style, as well as shifts in literary techniques over time and across spatial contexts.

提供机构:
未提及
创建时间:
2019-09-06
二维码
社区交流群
二维码
科研交流群
商业服务