遇见数据集

An Empirical Study of Prompt Engineering in LLMs for User Story Generation

收藏
Zenodo2026-04-22 更新2026-05-26 收录
官方服务:

资源简介:

Large Language Models (LLMs) have increasingly been used to support Requirements Engineering activities, particularly in the generation of user stories; however, questions remain regarding the stability of these outputs when the same elicitation process is repeated, as well as how such stability relates to expert-perceived quality. This paper investigates the semantic consistency of user stories generated by three LLMs (ChatGPT 5 Instant, DeepSeek, and Gemini 2.5 Pro) across four domains by repeatedly applying a protocol based on the 3Cs model. Consistency was measured using text embeddings and cosine similarity, focusing on comparable stories generated under the same identifiers across repetitions. Additionally, a controlled subset of the generated artifacts was evaluated by two Requirements Engineering specialists in two domains using criteria inspired by the Quality User Story (QUS) framework. The results show moderate to high semantic consistency, with variations across models and domains: ChatGPT and DeepSeek exhibited more predictable behavior, while Gemini demonstrated a more exploratory profile, generating a higher number of stories but with lower repeatability. Expert evaluation indicated that the generated stories are useful as an initial drafting mechanism, although recurring issues were identified in terms of atomicity and testability. Overall, the findings suggest that semantic consistency and perceived quality capture distinct aspects of LLM-generated artifacts, reinforcing the need to consider both as complementary dimensions in evaluation.

提供机构:
Zenodo
创建时间:
2026-04-21
二维码
社区交流群
二维码
科研交流群
商业服务