TEXT2ARCH
收藏资源简介:
TEXT2ARCH是由印度理工学院·鲁尔基分校联合谷歌、微软研发的大规模科学架构图数据集,包含75,127组精准对齐的文本描述-DOT代码-图像三元组。该数据集通过GPT-4o提示工程和结构化解析技术构建,涵盖神经网络架构、软件系统设计等科学图示,其中训练集60,519条、验证集7,565条、测试集7,043条。其核心价值在于解决文本到架构图的语义对齐难题,为AI辅助软件工程、教育可视化等领域提供基准支持。
TEXT2ARCH is a large-scale scientific architecture diagram dataset developed jointly by the Indian Institute of Technology Roorkee, Google and Microsoft. It contains 75,127 precisely aligned text description-DOT code-image triplets. Built using GPT-4o prompt engineering and structured parsing techniques, this dataset covers scientific diagrams including neural network architectures, software system designs and other relevant scientific visualizations. The dataset is split into a training set with 60,519 entries, a validation set with 7,565 entries, and a test set with 7,043 entries. Its core value lies in addressing the semantic alignment challenge in text-to-architecture diagram generation, providing benchmark support for fields such as AI-assisted software engineering and educational visualization.

- 1Text2Arch: A Dataset for Generating Scientific Architecture Diagrams from Natural Language Descriptions印度理工学院·鲁尔基分校; 谷歌; 微软 · 2026年



