davanstrien/ai4lam-demo
收藏资源简介:
--- dataset_info: features: - name: record_id dtype: string - name: date dtype: timestamp[ns] - name: raw_date dtype: string - name: title dtype: string - name: place dtype: string - name: empty_pg dtype: bool - name: text dtype: string - name: pg dtype: int64 - name: mean_wc_ocr dtype: float64 - name: std_wc_ocr dtype: float64 - name: name dtype: string - name: all_names dtype: string - name: Publisher dtype: string - name: Country of publication 1 dtype: string - name: all Countries of publication dtype: string - name: Physical description dtype: string - name: Language_1 dtype: string - name: Language_2 dtype: string - name: Language_3 dtype: 'null' - name: Language_4 dtype: 'null' - name: multi_language dtype: bool splits: - name: train num_bytes: 5300866 num_examples: 4148 download_size: 2857751 dataset_size: 5300866 --- # Dataset Card for "ai4lam-demo" [More Information needed](https://github.com/huggingface/datasets/blob/main/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
数据集信息: 特征项: - 名称:记录ID(record_id),数据类型:字符串 - 名称:日期(date),数据类型:纳秒级时间戳(timestamp[ns]) - 名称:原始日期(raw_date),数据类型:字符串 - 名称:标题(title),数据类型:字符串 - 名称:出版地点(place),数据类型:字符串 - 名称:空白页标记(empty_pg),数据类型:布尔型 - 名称:文本内容(text),数据类型:字符串 - 名称:页码(pg),数据类型:64位整数 - 名称:OCR文本平均词数(mean_wc_ocr),数据类型:64位浮点数 - 名称:OCR文本词数标准差(std_wc_ocr),数据类型:64位浮点数 - 名称:名称(name),数据类型:字符串 - 名称:所有名称(all_names),数据类型:字符串 - 名称:出版方(Publisher),数据类型:字符串 - 名称:第一出版国家(Country of publication 1),数据类型:字符串 - 名称:所有出版国家(all Countries of publication),数据类型:字符串 - 名称:物理描述(Physical description),数据类型:字符串 - 名称:第一语言(Language_1),数据类型:字符串 - 名称:第二语言(Language_2),数据类型:字符串 - 名称:第三语言(Language_3),数据类型:空值(null) - 名称:第四语言(Language_4),数据类型:空值(null) - 名称:多语言标记(multi_language),数据类型:布尔型 拆分集: - 名称:训练集(train),字节数:5300866,样本数量:4148 下载大小:2857751 数据集总大小:5300866 # "ai4lam-demo"数据集卡片 [需补充更多信息](https://github.com/huggingface/datasets/blob/main/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
数据集信息
特征
- record_id: 字符串类型
- date: 时间戳类型
- raw_date: 字符串类型
- title: 字符串类型
- place: 字符串类型
- empty_pg: 布尔类型
- text: 字符串类型
- pg: 整数类型
- mean_wc_ocr: 浮点数类型
- std_wc_ocr: 浮点数类型
- name: 字符串类型
- all_names: 字符串类型
- Publisher: 字符串类型
- Country of publication 1: 字符串类型
- all Countries of publication: 字符串类型
- Physical description: 字符串类型
- Language_1: 字符串类型
- Language_2: 字符串类型
- Language_3: 空值类型
- Language_4: 空值类型
- multi_language: 布尔类型
数据分割
- train:
- 字节数: 5300866
- 样本数: 4148
数据集大小
- 下载大小: 2857751 字节
- 数据集大小: 5300866 字节



