保险合同数据集
收藏资源简介:
本数据集为保险合同数据集,来源于开源项目“ETIP: An element tagging problem for Chinese insurance policy analysis”,公开于GitHub平台,可供自由获取与使用。数据集包含150份经人工标注的短保险合同,覆盖7类核心条款要素,结构明确、标注规范。数据集样本依据官方划分方案分为训练集和测试集,分别包含2,709条和2,699条标注数据,专用于法律要素提取模型在保险合同要素提取任务中的性能评估。本数据集为保险合同文本的结构化分析与自动化处理提供了高质量的基础资源,对推动法律智能、信息抽取和自然语言处理技术在专业领域中的应用具有重要的潜在价值,可支持相关模型的训练、评估与比较研究。
This dataset is an insurance contract dataset sourced from the open-source project "ETIP: An element tagging problem for Chinese insurance policy analysis", which is publicly available on GitHub for free access and use. It contains 150 manually annotated short insurance contracts covering 7 core clause elements, with clear structure and standardized annotations. The dataset samples are divided into training set and test set according to the official splitting scheme, which include 2,709 and 2,699 annotated instances respectively, and are specifically designed for the performance evaluation of legal element extraction models on the insurance contract element extraction task. This dataset provides high-quality foundational resources for the structured analysis and automated processing of insurance contract texts, holds important potential value for promoting the application of legal intelligence, information extraction and natural language processing technologies in professional fields, and can support the training, evaluation and comparative research of relevant models.




