官方服务:
资源简介:
支持对中国大陆护照个人资料页所有字段进行结构化识别,包括护照号码、姓名拼音、姓名、性别、出生日期、有效期、签发日期等14个字段。
应用场景:
提供机构:
上海斗斗星信息科技有限公司创建时间:
2022-10-12
相关数据集
Factory Email Extraction
该项目解析了一个合成的工厂生产电子邮件语料库(`data.txt`)并将其转换为结构化的CSV文件(`output.csv`)。输入文件`data.txt`包含100封电子邮件,每封电子邮件都包裹在`-----EMAIL START-----`和`-----EMAIL END-----`之间,包含`To:`、`From:`和`Subject:`标题,以及一个自由形式的正文,其中嵌入了生产细节,如序
github2026-02-27 更新200
CXRGraph: Using Information Extraction to Normalize the Training Data for Automatic Radiology Report Generation
CXRGraph is a dataset of structured radiology reports dataset following the RadGraph format, which has been tailored for the Automatic Radiology Report Generation (ARRG) task. CXRGraph assorts clinica
DataCite Commons2025-02-03 更新120
FanChen0116/bus_few4_16x_empty
--- dataset_info: features: - name: id dtype: int64 - name: tokens sequence: string - name: labels sequence: class_label: names: '0': O '1': I-fro
Hugging Face2023-09-27 更新90
ACE 2007 Multilingual Training Corpus
Introduction ACE 2007 Multilingual Training Corpus was developed by the Linguistic Data Consortium (LDC) and contains the complete set of Arabic and Spanish training data for the <a h
DataCite Commons2021-07-01 更新151
MikeGreen2710/12_jun_hn_pred_p1
该数据集包含多个特征字段,包括id、text以及多个以序列形式存储的字段(如STR、LEG、LOC等)。数据集中包含一个名为train的分割,该分割包含349,939个样本,总大小为362,588,182字节。下载大小为171,782,541字节。
Hugging Face2024-06-17 更新120



