SPEC5G
收藏资源简介:
SPEC5G是由普渡大学创建的第一个公开的5G数据集,专为NLP研究设计。该数据集包含3,547,587个句子,总计134M字,来源于13094个蜂窝网络规范和13个在线网站。数据集的创建过程涉及从3GPP网站和多个博客、论坛中收集数据,并通过一系列预处理步骤进行清洗和整理。SPEC5G的应用领域广泛,包括安全测试、政策执行、自动代码生成和协议摘要等,旨在通过自动化分析减少5G协议开发和安全分析中的人工努力。
SPEC5G is the first publicly available 5G dataset developed by Purdue University, tailored specifically for natural language processing (NLP) research. This dataset comprises 3,547,587 sentences totaling 134 million words, sourced from 13,094 cellular network specifications and 13 online websites. The dataset creation process involved collecting data from the 3GPP website, multiple blogs and forums, followed by a series of preprocessing steps for data cleaning and curation. SPEC5G has a wide range of application scenarios including security testing, policy enforcement, automated code generation, and protocol summarization, among others. It aims to reduce manual effort in 5G protocol development and security analysis through automated analysis.




