Structured ETL and Data-Preparation Pipeline Descriptions for LLM Workflow Generation Studies
收藏资源简介:
This dataset contains 60 structured natural-language pipeline descriptions for studying LLM-based workflow code generation, bounded self-repair, and pipeline design. The corpus includes 40 main-study ETL/data-preparation pipeline descriptions and 20 held-out validation pipeline descriptions. The descriptions cover repository-derived and synthetic workflow scenarios across common orchestration patterns, including linear pipelines, fan-out/fan-in, branch/merge, sensor-gated workflows, staged ETL, and fan-out-only designs. The dataset is intended to support research on executable workflow generation, pipeline-design reasoning, orchestration-aware code generation, and evaluation of LLM repair behavior for data-processing workflows. Each description specifies the pipeline purpose, control-flow dependencies, data artifacts, external systems, pipeline steps, and scheduling information.



