遇见数据集

NER dataset related to legal texts

收藏
Mendeley Data2026-04-09 收录
官方服务:

资源简介:

The following data pertains to Named Entity Recognition for legal judgment documents related to the crime of assisting in information network crimes. The dataset consists of a total of 4,236 samples, including both training and validation data, with a total of 8 labels. The file train1.json contains the raw data in JSON format, which is not divided into training and validation sets. The ner_data folder contains processed data in .txt file format, with the dataset split into training and validation sets at a ratio of 5:1. This folder also includes all label names. Ultimately, the model is trained using the processed dataset.

本数据集针对涉及帮助信息网络犯罪活动罪的司法裁判文书,提供命名实体识别(Named Entity Recognition, NER)相关标注数据。数据集总计包含4236条样本,涵盖训练与验证数据,共设置8个标签。文件train1.json存储了未划分训练集与验证集的原始JSON格式数据。ner_data文件夹内存放处理完成的TXT格式数据集,该数据集已按照5:1的比例划分为训练集与验证集,且该文件夹同时包含全部标签名称。最终将基于该处理后的数据集开展模型训练。

二维码
社区交流群
二维码
科研交流群
商业服务