memo_enriched
收藏资源简介:
该数据集是一个结构化的文献数据集,主要包含书籍或文章的文本内容及其丰富的元数据信息。数据集的核心内容包括:1. 文献文本与元数据:每个样本包含文献的完整文本、标题、副标题、卷册信息、出版年份、页数、出版社、价格等详细的出版信息。2. 作者信息:记录了作者的名字、笔名、性别、国籍、生卒年份以及作者ID等。3. 文献分类与来源:文献被分类并标注了历史时期,同时提供了文献的数字化来源和在线访问信息。4. 关联网络数据:数据集包含了作者或作品之间的引用与提及关系,例如提及的链接、提及的作品以及在多个权威文学名录或经典列表中是否被提及的布尔标记。5. 统计与标注信息:包含文本词数统计以及关于文献是否属于特定文化经典或教育经典的权威性标注。数据集共包含1174个样本,适用于文学研究、作者社会网络分析、文献数字化档案管理、文化经典性评估以及自然语言处理任务中的文本分析和元数据挖掘。
This dataset is a structured literature dataset that primarily includes the text content of books or articles along with rich metadata information. The core content of the dataset comprises: 1. Literature text and metadata: Each sample contains the complete text of the literature, title, subtitle, volume information, publication year, page count, publisher, price, and other detailed publication information. 2. Author information: Records the authors name, pseudonym, gender, nationality, birth and death years, and author ID. 3. Literature classification and sources: The literature is categorized and labeled with historical periods, while also providing digital sources and online access information. 4. Relational network data: The dataset includes citation and mention relationships between authors or works, such as mentioned links, mentioned works, and Boolean markers indicating whether they are mentioned in authoritative literary directories or classic lists. 5. Statistical and annotation information: Includes text word count statistics and authoritative annotations regarding whether the literature belongs to specific cultural canons or educational canons. The dataset contains a total of 1174 samples and is suitable for literary research, author social network analysis, digital literature archive management, cultural canon evaluation, and text analysis and metadata mining in natural language processing tasks.
数据集概述
- 数据集名称: memo_enriched
- 数据集地址: https://huggingface.co/datasets/chcaa/memo_enriched
- 数据集大小: 下载大小为272,317,354字节,整个数据集大小为438,734,847字节
- 样本数量: 训练集包含1,174个样本
数据特征
该数据集包含丰富的元数据字段,覆盖了书籍、作者、文本内容及相关索引信息,主要分为以下几类:
书籍元信息
filename: 文件名text: 文本内容title: 标题subtitle: 副标题volume: 卷/册pages: 页数illustrations: 插图信息typeface: 字体publisher: 出版商price: 价格publ_date: 出版日期title_modern: 现代标题novel_start,novel_end: 小说起始与结束信息serialno: 序列号category: 分类full: 完整文本标识
作者信息
auth_first,auth_last: 作者名与姓auth_last_modern: 现代形式的姓pseudonym: 笔名gender: 性别real_gender: 真实性别published_under_gender: 发表时使用的性别nationality: 国籍birth_auth,death_auth: 出生与逝世年份surname: 姓氏full_firstnames: 全名auth_id: 作者ID
文本来源与处理
scan source: 扫描来源text source: 文本来源note to bibl. inf.: 书目信息注释note to text: 文本注释barcode (kb): 条形码online: 在线状态kb-levering: 交付信息simple_word_count: 简单字数统计wordcount_main_text_lex_auth: 主文本词汇量search_term,search_pseudonym: 搜索词与笔名搜索
作者索引与提及信息
arbejdsliv_auth,familie_auth,name_lex_auth,full_name_lex_auth: 作者姓名及相关索引mentioned_links_lex_auth: 提及的链接mentioned_works_lex_auth: 提及的作品search_link_lex_auth,search_link_lex_auth_norm: 搜索链接title_mention_lex_authorpage: 作者页面标题提及lex_author_mentioned_in_other_authorpage: 在其他作者页面被提及的次数
规范与经典索引
auth_mention_*: 在多个文学史、手册、指南等资源中被提及的标识(如grøn_dan_lithis5_nameregistry,lithåndb_nameregistry,trane_nameregistry,rød_dan_lithis2_nameregistry等)title_mention_rød_dan_lithis_titleregistry: 在红色丹麦文学史标题注册表中的标题提及canon_culture_authors,canon_educational_2025,canon_educational_2004: 文化经典与教育经典作者标识title_mention_canon_culture_authors,title_mention_canon_culture_authors_dsl_classic,author_mention_dsl_classic: 经典作品与DSL经典提及
其他
year: 年份period,period_notes: 时期与时期注释historical: 历史标识notes: 注释source: 来源title.1: 备选标题__index_level_0__: 索引层级
数据划分
- 训练集: 仅有
train划分,包含1,174个样本。 - 数据文件: 数据存储在
data/train-*路径下。




