遇见数据集

Model Generation from Requirements with LLMs: an Exploratory Study - Replication Package

收藏
NIAID Data Ecosystem2026-05-01 收录
数据链接:
官方服务:

资源简介:

This is a replication package for the paper "Model Generation from Requirements with LLMs: an Exploratory Study", by Sallam Abualhaija, Chetan Arora, and Alessio Ferrari. Abstract: Complementing natural language (NL) requirements with graphical models can improve stakeholders’ communication and provide directions for system design. However, creating models from requirements involves manual effort. The advent of generative large language models (LLMs), ChatGPT being a notable example, offers promising avenues for automated assistance in model generation. This paper investigates the reliability of ChatGPT in generating sequence diagrams from NL requirements. Specifically, we conduct a qualitative study examining the sequence diagrams generated by ChatGPT for 28 requirements documents of various types and from different domains. Our study aims to uncover potential issues that emerge in the models generated by ChatGPT, thereby hindering its applicability in practice. Observations have systematically been captured through evaluation logs, and categorized through thematic analysis. Our results indicate that, although the models generally conform to the standard and exhibit a reasonable level of understandability, their correctness with respect to the specified requirements often presents challenges. This issue is particularly pronounced in the presence of requirements smells, such as ambiguity and inconsistency. The insights derived from this study can influence the practical utilization of LLMs in the RE process, and open the door to novel RE-specific prompting strategies targeting effective model generation. The replication package consists of the following folders: logs: includes the evaluation logs produced by each evaluator original-documents: includes the original requirements documents used for the evaluation RQ1 - quantitative analysis: includes the analysis made on the scores given to each model and model variant. It includes five files: - results.csv: numerical results of the evaluation for each criterion- analysis-results.Rmd: R file used to perform the quantitative analysis (requires R Studio to be executed)- analysis-results.html: html file produced by analysis-results.Rmd- cross-check.csv: file with the cross-checking of the two assessors applied to a subset of the models- symmary_results.xlsx: final output of the quantitative results in terms of Wilcoxon signed rank tests RQ2 - thematic analysis: includes the codebook produced by the thematic analysis of the issues in generating models with ChatGPT

创建时间:
2024-04-03
二维码
社区交流群
二维码
科研交流群
商业服务