遇见数据集

Semantic Code Analysis for NFR Prioritization Using LLMs

收藏
Zenodo2025-10-16 更新2026-05-26 收录
官方服务:

资源简介:

Research Context: Non-Functional Requirements (NFRs) play a vital role in ensuring software quality, particularly in domains where failures in performance or security can lead to serious consequences. However, in organizational environments, the elicitation and prioritization of NFRs are often hindered by limited documentation and subjective decision-making. This study explores how semantic analysis of source code can support automated NFR prioritization using Large Language Models (LLMs). Practical Problem: Conventional prioritization techniques—such as AHP and MoSCoW—depend heavily on stakeholder input and are rarely applied systematically to NFRs, especially in large-scale or poorly documented systems. This gap increases the risk of overlooking critical quality attributes. Proposed Solution: We present a semantic code analysis pipeline that leverages LLMs to extract, classify, and prioritize NFRs directly from source code. The method uses structured prompts to guide the identification of quality attributes, aligning them with ISO/IEC 25010 categories and assigning priority levels (High, Medium, Low). Theoretical Foundation: Grounded in Software Quality and Requirements Engineering theories, the approach integrates decision-making models and quality frameworks to support organizational needs. Methodology: The pipeline was empirically evaluated on the OpenMRS repository (134 Java files), comparing automated outputs with expert annotations. Performance was measured using precision, recall, and F1-score. Results: The system identified 422 NFRs, achieving 93.4% recall, 30.3% precision, and a 45.8% F1-score. High-priority requirements were predominantly related to Security and Performance Efficiency, reflecting the critical nature of the system.

研究背景:非功能需求(Non-Functional Requirements, NFRs)对保障软件质量至关重要,尤其在性能或安全故障可能引发严重后果的领域。然而在组织环境中,由于文档匮乏与决策主观性,非功能需求的获取与优先级划分常遭遇阻碍。本研究探讨如何借助大语言模型(Large Language Models, LLMs)对源代码开展语义分析,以支持自动化的非功能需求优先级划分工作。 实际问题:传统的优先级划分技术,如层次分析法(Analytic Hierarchy Process, AHP)与莫斯科方法(MoSCoW),高度依赖涉众输入,且极少被系统地应用于非功能需求的处理,尤其是在大规模或文档缺失的系统中。这一缺口会提升忽视关键质量属性的风险。 提出的解决方案:本文提出一种语义代码分析流水线,借助大语言模型直接从源代码中提取、分类并划分非功能需求的优先级。该方法通过结构化提示词引导质量属性的识别,将其与ISO/IEC 25010标准分类对齐,并赋予高、中、低三个优先级等级。 理论基础:本方法扎根于软件质量与需求工程理论,整合决策模型与质量框架以适配组织的实际需求。 方法论:该流水线在OpenMRS代码仓库(包含134个Java文件)上开展了实证评估,将自动化输出结果与专家标注结果进行对比。评估指标采用精确率(precision)、召回率(recall)与F1值(F1-score)。 结果:本系统共识别出422项非功能需求,召回率达93.4%,精确率为30.3%,F1值为45.8%。高优先级需求主要涉及安全与性能效率,契合该系统的关键属性特性。

提供机构:
Zenodo
创建时间:
2025-10-16
二维码
社区交流群
二维码
科研交流群
商业服务