PAYSLIPS
收藏资源简介:
PAYSLIPS数据集是由SCOR和布雷斯特大学等机构创建的,专门用于保险领域中的命名实体识别任务。该数据集包含611页的匿名保险工资单,标注了财务信息,数据来源于残疾保险的工资单。数据集的创建过程包括手动匿名化处理,并由保险专家验证标注质量。PAYSLIPS数据集旨在解决保险领域中自动提取财务信息的问题,特别是在处理大量敏感数据时,确保隐私和安全。
The PAYSLIPS dataset was developed by institutions including SCOR and the University of Brest, and is specifically designed for named entity recognition (NER) tasks in the insurance domain. It comprises 611 pages of anonymized insurance payroll records with annotated financial information, with the data sourced from disability insurance payrolls. The dataset’s creation involved manual anonymization procedures, and its annotation quality was verified by insurance experts. The PAYSLIPS dataset aims to address the challenge of automated financial information extraction in the insurance sector, particularly when handling large volumes of sensitive data, while safeguarding privacy and security.

- 1Training LayoutLM from Scratch for Efficient Named-Entity Recognition in the Insurance DomainSCOR, 布雷斯特大学, CNRS, UMR 6205, LMBA, INSA Rennes, IRISA, Inria, CNRS, 雷恩大学 · 2024年



