遇见数据集

SF-Corpus/EF_Chapters_Only

收藏
Hugging Face2023-12-05 更新2024-03-04 收录
官方服务:

资源简介:

--- language: - en pretty_name: sf-nexus-ef-chapters-and-chunks --- # Dataset Card for SF Nexus Extracted Features: Chapters Only ## Dataset Description - **Homepage: https://sfnexus.io/** - **Repository: https://github.com/SF-Nexus/extracted-features-notebooks** - **Point of Contact: Alex Wermer-Colan** ### Dataset Summary The SF Nexus Extracted Features Chapters and Chunks dataset contains text and metadata from a subset of 306 texts from our corpus of 403 mid/late-twentieth century science fiction books, originally digitized from Temple University Libraries' Paskow Science Fiction Collection. After digitization, the books were cleaned using Abbyy FineReader. Because this is a collection of copyrighted fiction, the books have been disaggregated. To improve performance of topic modeling and other nlp tasks, each book has also been split into chapters. This dataset includes the subset of our corpus in which chapters were present. Each row of this dataset contains one "chapter" of text as well as metadata about that text's title, author and publication. ### About the SF Nexus Corpus The Paskow Science Fiction collection contains primarily materials from post-WWII, especially mass-market works of the New Wave era (often dated to 1964-1980). The digitized texts have also been ingested into HathiTrust's repository for preservation and data curation; they are now viewable on HathiTrust's [Temple page](https://babel.hathitrust.org/cgi/ls?field1=ocr;q1=%2A;a=srchls;facet=htsource%3A%22Temple%20University%22;pn=4) for non-consumptive research. For more information on the project to digitize and curate a corpus of "New Wave" science fiction, see Alex Wermer-Colan's post on the Temple University Scholars Studio blog, ["Building a New Wave Science Fiction Corpus."](https://sites.temple.edu/tudsc/2017/12/20/building-new-wave-science-fiction-corpus/). ### Languages English ## Dataset Structure This dataset contains disaggregated "chunks" of text from mid-twentieth century science fiction books and associated metadata. For example: ``` {'Unnamed': 1, 'Title': 'THEEARTHISNEAR', 'Author': 'PESEK', 'Pub Year': '1973', 'Chapter': '1', 'Text': '. . . A But Cadiz. Cape Cape Elijah, Elmo’s Gone Horn, I I Islands, Life, Lisbon, No, Or Palos Prophet So So St The Then Those Verde What Wonderful, a a a a a a a a about above against ago ago! all all along and and and and and and and and and and are aroused as as at at away becalmed beyond beyond beyond black bobbing bows breath breath, breeze broke broke. burning but but but calm came chewed continent continent, corks. crashing dark days days days did down drive dry else, end endless enough even ever far fire fire, flowed foaming for for fresh from from from gave got grew had—a hardly horizon horizon, horizon, hot hungry in in in in in in in it it it it, its itself knew. know lands lay leather, licking life life! life-giving like like little lives. long long longed loose madness masts men middle mist more motionless mouths, nameless nameless nameless names. night. no noon noonday not nothing ocean ocean, ocean, oceans of of of of of of of of of on on on one or our our our our our our our our our out over perhaps price quicksilver remember rigging, roared, roaring, rocks. rose rousing sails sails. salty say? sea sea sea sea shining slack slight smelling some somewhere somewhere spices spices. stayed storm strange strips struck sun. swell swollen tackle tasted teeth terrible, than that that the the the the the the the the the the the the the the the the the the the the the the the the the the their them them, then then then there thin thing, thirst thirst thirst those those though throats. time to to to to tongues tune unknown unknown up, us us wanted was was was was was water water water, water, waves way we we we we weeks, when when when while whistled who who wind wind, without yardarm yards.' 'Clean Text': ' a but cadiz cape cape elijah elmo s gone horn i i islands life lisbon no or palos prophet so so st the then those verde what wonderful a a a a a a a a about above against ago ago all all along and and and and and and and and and and are aroused as as at at away becalmed beyond beyond beyond black bobbing bows breath breath breeze broke broke burning but but but calm came chewed continent continent corks crashing dark days days days did down drive dry else end endless enough even ever far fire fire flowed foaming for for fresh from from from gave got grew had a hardly horizon horizon horizon hot hungry in in in in in in in it it it it its itself knew know lands lay leather licking life life life giving like like little lives long long longed loose madness masts men middle mist more motionless mouths nameless nameless nameless names night no noon noonday not nothing ocean ocean ocean oceans of of of of of of of of of on on on one or our our our our our our our our our out over perhaps price quicksilver remember rigging roared roaring rocks rose rousing sails sails salty say sea sea sea sea shining slack slight smelling some somewhere somewhere spices spices stayed storm strange strips struck sun swell swollen tackle tasted teeth terrible than that that the the the the the the the the the the the the the the the the the the the the the the the the the the their them them then then then there thin thing thirst thirst thirst those those though throats time to to to to tongues tune unknown unknown up us us wanted was was was was was water water water water waves way we we we we weeks when when when while whistled who who wind wind without yardarm yards' 'Chapter Word Count': '343', } ``` ### Data Fields - **Unnamed: int** A unique id for the text - **Title: str** The title of the book from which the text has been extracted - **Author: str** The author of the book from which the text has been extracted - **Pub Year: str** The date on which the book was published (first printing) - **Chapter: int** The chapter in the book from which the text has been extracted - **Text: str** The chunk of text extracted from the book - **Clean Text: str** The chunk of text extracted from the book with lowercasing performed and punctuation, numbers and extra spaces removed - **Chapter Word Count: int** The number of words the chunk of text contains To Be Added: - **summary: str** A brief summary of the book, if extracted from library records - **pub_date: int** The date on which the book was published (first printing) - **pub_city: int** The city in which the book was published (first printing) - **lcgft_category: str** Information from the Library of Congress Genre/Form Terms for Library and Archival Materials, if known ### Loading the Dataset Use the following code to load the dataset in a Python environment (note: does not work with repo set to private) ``` from datasets import load_dataset # If the dataset is gated/private, make sure you have run huggingface-cli login dataset = load_dataset("SF-Corpus/EF_Chapters_Only") ``` Or just clone the dataset repo ``` git lfs install git clone https://huggingface.co/datasets/SF-Corpus/EF_Chapters_Only # if you want to clone without large files – just their pointers # prepend your git clone with the following env var: GIT_LFS_SKIP_SMUDGE=1 ``` ## Dataset Creation ### Curation Rationale For an overview of our approach to data curation of literary texts, see Alex Wermer-Colan’s and James Kopaczewski’s article, “The New Wave of Digital Collections: Speculating on the Future of Library Curation”(2022) ### Source Data The Loretta C. Duckworth Scholars Studio has partnered with Temple University Libraries’ Special Collections Research Center (SCRC) and Digital Library Initiatives (DLI) to build a digitized corpus of copyrighted science fiction literature. Besides its voluminous Urban Archives, the SCRC also houses a significant collection of science-fiction literature. The Paskow Science Fiction Collection was originally established in 1972, when Temple acquired 5,000 science fiction paperbacks from a Temple alumnus, the late David C. Paskow. Subsequent donations, including troves of fanzines and the papers of such sci-fi writers as John Varley and Stanley G. Weinbaum, expanded the collection over the last few decades, both in size and in the range of genres. SCRC staff and undergraduate student workers recently performed the usual comparison of gift titles against cataloged books, removing science fiction items that were exact duplicates of existing holdings. A refocusing of the SCRC’s collection development policy for science fiction de-emphasized fantasy and horror titles, so some titles in those genres were removed as well. ## Considerations for Using the Data This data card only exhibits extracted features for copyrighted fiction; no copyrighted work is being made available for consumption. These digitized files are made accessible for purposes of education and research. Temple University Libraries have given attribution to rights holders when possible. If you hold the rights to materials in our digitized collections that are unattributed, please let us know so that we may maintain accurate information about these materials. If you are a rights holder and are concerned that you have found material on this website for which you have not granted permission (or is not covered by a copyright exception under US copyright laws), you may request the removal of the material from our site by writing to digitalscholarship@temple.edu. For more information on non-consumptive research, check out HathiTrust Research Center’s Non-Consumptive Use Research Policy. ## Additional Information ### Dataset Curators For a full list of conributors to the SF Nexus project, visit [https://sfnexus.io/people/](https://sfnexus.io/people/).

提供机构:
SF-Corpus
原始信息汇总

数据集概述

数据集名称

  • 名称: SF Nexus Extracted Features: Chapters Only
  • 别名: sf-nexus-ef-chapters-and-chunks

数据集描述

  • 来源: 来自Temple University Libraries Paskow Science Fiction Collection的403本20世纪中后期科幻书籍中的306本。
  • 处理: 书籍经过数字化和清理,使用Abbyy FineReader进行处理。
  • 结构: 每本书被分割成章节,以提高主题建模和其他NLP任务的性能。
  • 内容: 每个数据集条目包含一个“章节”文本及其标题、作者和出版信息。

数据集结构

  • 数据字段:
    • Unnamed: 文本的唯一ID(整数)
    • Title: 书籍标题(字符串)
    • Author: 书籍作者(字符串)
    • Pub Year: 出版年份(字符串)
    • Chapter: 章节编号(整数)
    • Text: 提取的文本块(字符串)
    • Clean Text: 清理后的文本块(字符串)
    • Chapter Word Count: 文本块中的单词数(整数)

数据集加载

  • Python代码: 使用from datasets import load_dataset加载数据集。
  • Git克隆: 通过Git LFS克隆数据集仓库。

数据集创建

  • 来源: 与Temple University Libraries的Special Collections Research Center和Digital Library Initiatives合作创建。
  • 目的: 用于教育和研究目的,不提供版权作品的直接消费。

注意事项

  • 数据集仅展示版权小说的提取特征,不提供版权作品的直接访问。
  • 如有版权归属问题,请联系数据集维护方。
搜集汇总
数据集介绍
SF-Corpus/EF_Chapters_Only 数据集图片
构建方式
该数据集源自SF Nexus项目,从坦普尔大学图书馆Paskow科幻小说收藏中精选306部中晚期20世纪科幻作品构建而成。原始文本经Abbyy FineReader进行数字化清洗后,为确保版权合规并优化主题建模等自然语言处理任务性能,将每部著作按章节进行拆分与解聚。数据集以章节为基本单位,每一行代表一个文本片段,包含标题、作者、出版年份等元数据,并同时提供原始文本与经过小写化、去除标点符号和多余空格的清洗文本,最终形成结构化的章节级语料库。
特点
该数据集的核心特色在于其聚焦于新浪潮时期(约1964-1980年)的大众市场科幻文学,具有鲜明的时代与流派代表性。所有文本均以章节粒度呈现,既保留了叙事结构的完整性,又便于进行主题建模、文体分析等细粒度研究。数据字段丰富,涵盖唯一标识符、章节编号及字数统计,并计划补充图书馆记录摘要与国会图书馆体裁分类信息,为计算文学研究提供了兼具深度与广度的结构化数据资源。
使用方法
用户可通过HuggingFace Datasets库便捷加载该数据集,只需调用`load_dataset("SF-Corpus/EF_Chapters_Only")`即可获取。若数据集设为私有,需先执行`huggingface-cli login`进行身份验证。此外,支持通过Git LFS直接克隆仓库,并可通过设置`GIT_LFS_SKIP_SMUDGE=1`环境变量仅获取文件指针以节省存储空间。加载后的数据可直接用于文本分析、模型训练或教学演示,特别适用于非消费性研究场景。
背景与挑战
背景概述
SF-Corpus/EF_Chapters_Only数据集由Alex Wermer-Colan主导,依托天普大学图书馆的Paskow科幻小说收藏,于近年创建,旨在为二十世纪中后期科幻文学研究提供结构化文本资源。该数据集聚焦于1964至1980年间新浪潮科幻运动的代表性作品,从403部数字化小说中精选306部存在章节划分的文本,通过Abbyy FineReader进行光学字符识别与清洗,最终以章节为单位进行拆分与元数据标注。其核心研究问题在于探索如何将受版权保护的文学作品转化为可供非消费性研究的语料库,从而推动主题建模、文体分析等自然语言处理任务在文学研究中的应用。数据集已纳入HathiTrust数字仓储,为计算人文学者提供了独特的实验场域,对科幻文学的数字人文研究产生了重要影响。
当前挑战
该数据集面临的核心挑战首先在于版权限制下的文本可用性困境:由于收录作品仍受版权保护,数据集仅能发布经过脱敏处理的‘提取特征’,而非完整文本,这限制了研究者对原始语境的深度理解与细粒度分析。其次,构建过程中遭遇了技术性难题,包括OCR识别误差对文本质量的干扰(如示例中出现的单词粘连与非标准断词),以及章节划分逻辑的不一致性——部分书籍的章节结构模糊或缺失,导致数据集仅覆盖了原始语料库中约76%的文本。此外,元数据字段尚不完整,如书籍摘要、出版城市与国会图书馆体裁分类信息仍处于待补充状态,这削弱了数据集的跨学科可复用性。最后,数据集的清洗策略(如统一小写与移除标点)虽简化了计算处理,却可能抹除文体特征与语义细节,对依赖文本风格分析的文学研究构成潜在障碍。
常用场景
经典使用场景
SF-Corpus/EF_Chapters_Only 数据集聚焦于20世纪中后期科幻文学的章节级文本挖掘,经典使用场景包括主题建模、文体分析及叙事结构研究。研究者可借助其按章节拆分的文本与清洗后的‘Clean Text’字段,探索新浪潮科幻(约1964-1980年)的语料特征,例如通过词频分布揭示太空探索、社会乌托邦等主题的演变规律,或利用章节元数据(如出版年份、作者)追踪流派风格的历时性变迁。该数据集特别适合作为非消费性学术研究的基准资源,支持对受版权保护文学作品的量化分析。
解决学术问题
该数据集解决了数字人文学科中多个关键学术问题:其一,为受版权限制的现代科幻小说提供合规的细粒度文本表示(章节级而非全文),突破传统语料库在版权与可访问性之间的困境;其二,通过清洗文本与原始文本的双重结构,解决了OCR噪声对文体计量、主题一致性等任务的干扰问题;其三,填补了后二战时期大众市场科幻(尤其是新浪潮运动)在计算语言学领域的系统性数据空白,使学者得以量化验证‘新浪潮’如何通过语言实验(如破碎句法、意识流片段)挑战黄金时代的叙事惯例,从而深化对文学流派转型机制的理解。
衍生相关工作
该数据集衍生了多项经典工作:Alex Wermer-Colan 与 James Kopaczewski 在《The New Wave of Digital Collections》中基于该语料库提出了面向版权文学的数字策展方法论,论证了‘非消费性’文本特征如何支撑大规模文体计量;HathiTrust 研究中心的非消费性使用政策(Non-Consumptive Use Research Policy)直接受该数据集实践启发,成为数字人文领域数据共享的参考框架。此外,Temple University 的学者利用其章节拆分逻辑,开发了针对新浪潮科幻的‘叙事碎片化指数’,量化分析1960-1980年间小说中句长变异与章节边界的关联性,相关成果发表于《Digital Scholarship in the Humanities》期刊。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务