遇见数据集

交通事故类型自动分类语料集

收藏
国家数据集管理服务平台2026-05-29 更新2026-05-30 收录
官方服务:

资源简介:

本语料集为面向交通事故类型自动分类研究的专业中文NLP数据集,共收录约10万组标题-正文对,基于大语言模型通过可控提示词工程批量虚拟生成。语料覆盖追尾碰撞、行人事故、非机动车事故、酒驾毒驾、车辆侧翻、恶劣天气事故、高速公路事故、路口事故、车辆起火、开门杀、夜间事故、违规操作、肇事逃逸、单方事故、多方连环事故、危险品运输事故、停车场事故、校车事故、公交事故、货物散落事故等20个交通事故大类,并附有5000余个主题关键词支撑。内容风格模拟目击者分享、新闻报道、经验反思等真实写作场景,每条数据均配备事故严重程度分级、情感强度标注、文本意图分类、作者画像模拟、场景标签、可信度与实用价值评分等高价值多维标注字段,标注置信度均为1.0金标准。数据集全部为虚拟生成,无隐私风险,版权清晰,可直接用于文本分类模型训练与评测。

This is a professional Chinese NLP dataset tailored for research on automatic traffic accident type classification. It contains approximately 100,000 title-content pairs, which are generated in batches via controllable prompt engineering powered by large language models (LLMs). The corpus covers 20 major categories of traffic accidents, including rear-end collisions, pedestrian-involved accidents, non-motor vehicle accidents, drunk and drug-impaired driving accidents, vehicle rollovers, severe weather-related accidents, expressway accidents, intersection accidents, vehicle fires, dooring accidents, night-time accidents, violative operation accidents, hit-and-run accidents, single-vehicle accidents, multi-vehicle chain-reaction accidents, hazardous material transportation accidents, parking lot accidents, school bus accidents, public transit bus accidents, and cargo spill accidents. It is supported by over 5,000 thematic keywords. The content styles simulate real writing scenarios such as eyewitness accounts, news reports, and experience reflections. Each data entry comes with high-value multi-dimensional annotation fields including accident severity grading, emotional intensity annotation, text intent classification, simulated author profile, scene tags, and credibility and practical value scores, with all annotation confidence reaching the 1.0 gold standard. The entire dataset is virtually generated, with no privacy risks and clear copyrights, and can be directly used for text classification model training and evaluation.

创建时间:
2026-05-27
搜集汇总
数据集介绍
交通事故类型自动分类语料集 数据集图片
背景与挑战
背景概述
该数据集是一个用于交通事故类型自动分类研究的中文NLP语料集,包含约10万组虚拟生成的标题-正文对,覆盖20个事故大类,并提供了事故严重程度、情感强度等多维标注信息。它可直接用于文本分类模型的训练与评测,且无隐私风险。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务