遇见数据集

交通违法信息分类语料集

收藏
国家数据集管理服务平台2026-05-29 更新2026-05-30 收录
官方服务:

资源简介:

本数据集是一个面向交通违法信息分类的大规模中文语料集,共包含100,000条经过人工标注的数据样本。数据覆盖机动车违法(59.6%)、非机动车违法(21.2%)和行人违法(19.2%)三大类别,涵盖闯红灯、超速行驶、违章停车、逆行、不礼让行人、分心驾驶等多种违法类型。每条数据包含完整的标题-正文对,并标注了情感倾向(正面22.6%、负面58.2%、中性19.2%)、严重程度(轻微/一般/严重/致命)、主题关键词、实体信息等多维度标签,可用于自然语言处理领域的文本分类、情感分析、实体识别等任务。

This dataset is a large-scale Chinese corpus designed for traffic violation classification tasks, containing a total of 100,000 manually annotated data samples. The data covers three major categories: motor vehicle violations (59.6%), non-motor vehicle violations (21.2%), and pedestrian violations (19.2%), and includes diverse violation types such as running red lights, speeding, illegal parking, reverse driving, failure to yield to pedestrians, and distracted driving. Each sample consists of a complete title-body pair, and is annotated with multi-dimensional labels including sentiment polarity (22.6% positive, 58.2% negative, 19.2% neutral), severity level (minor/general/serious/fatal), topic keywords, and entity information. This dataset can be applied to natural language processing tasks such as text classification, sentiment analysis, and entity recognition.

创建时间:
2026-05-27
搜集汇总
数据集介绍
交通违法信息分类语料集 数据集图片
背景与挑战
背景概述
该数据集是一个面向交通违法信息分类的大规模中文语料集,包含10万条人工标注样本,覆盖机动车、非机动车和行人三大违法类别,并标注了情感倾向、严重程度等多维度标签。数据为合成生成,适用于自然语言处理领域的文本分类、情感分析等任务。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务