遇见数据集

Dataset and Replication Tool: Speaking in Dialects: A Reusable Dataset of Real-World TOSCA Orchestration Topologies

收藏
Zenodo2026-05-29 更新2026-06-05 收录
官方服务:

资源简介:

This artifact is the replication package for the paper "Speaking in Dialects: A Reusable Dataset of Real-World TOSCA Orchestration Topologies" submitted to ICSME 2026 - Tool and Dataset track. It contains two components: Dataset: A curated collection of 14,931 validated TOSCA files harvested from 260 public repositories on GitHub and Codeberg, covering all major TOSCA dialects and specification versions (TOSCA 2.0, TOSCA Simple YAML 1.x, Cloudify DSL, Alien4Cloud DSL, NFV Profile, and Unfurl DSL). The dataset is provided in both CSV and Parquet format. Each row represents one TOSCA file and includes file-level structural metadata (node types, node templates, inputs, outputs, imports, file size, line count, etc.) as well as repository-level metadata (stars, forks, license, topics, creation date, etc.). A per-file Git commit history is also included, enabling longitudinal analysis of TOSCA file evolution over time. Mining Tool (TOSCAmine): The fully automated, four-phase Python pipeline used to produce the dataset: repository discovery (GitHub keyword, code content, and topic search; Codeberg keyword search), cloning and TOSCA file extraction, validation and metadata extraction, and final dataset assembly. The tool is configurable and reproducible, with all dependencies pinned and the exact configuration snapshot used for this run included alongside the dataset. For a full description of the artifact structure, dataset schema, and instructions to rerun the pipeline, refer to the README.md included in this package or to the accompanying paper.

提供机构:
Zenodo
创建时间:
2026-05-29
二维码
社区交流群
二维码
科研交流群
商业服务