introvoyz041/OpenRTLSet
收藏资源简介:
OpenRTLSet是一个完全开源的数据集,专为基于大语言模型的Verilog模块设计而构建。它包含超过127,000个多样化的Verilog模块,这些模块来源于GitHub上的设计、VHDL/SystemVerilog翻译以及高层次综合(HLS)生成的RTL。数据集通过可追踪的GitHub起源和宽松的许可证确保了可重现性,直接解决了LLM驱动硬件设计领域中透明研究和商业应用的关键障碍。广泛的实验表明,使用OpenRTLSet微调的多个LLM(从7B到32B规模)在Verilog代码生成准确性方面优于在其他数据集上训练的模型。数据集规模为131k,每个样本包含索引、模块I/O头、Verilog代码实现、完整代码、由DeepSeek-R1-Distill-Llama-70B LLM生成的功能描述、GitHub仓库URL、许可证名称以及可选的子模块索引等字段。
OpenRTLSet is a fully open-source dataset specifically constructed for Verilog module design based on large language models (LLMs). It contains over 127,000 diverse Verilog modules sourced from GitHub-hosted designs, VHDL/SystemVerilog translations, and high-level synthesis (HLS) generated RTL. The dataset ensures reproducibility via traceable GitHub provenance and permissive licenses, directly addressing critical barriers to transparent research and commercial deployment in the field of LLM-driven hardware design. Extensive experiments demonstrate that multiple LLMs (scaling from 7B to 32B parameters) fine-tuned with OpenRTLSet outperform models trained on other datasets in terms of Verilog code generation accuracy. The dataset has a scale of 131k samples, with each entry containing fields including index, module I/O header, Verilog code implementation, complete code, functional description generated by the DeepSeek-R1-Distill-Llama-70B LLM, GitHub repository URL, license name, and optional submodule index.



