Miyazaki94
收藏资源简介:
Notes from the Author Creators Miyazaki Y. Terakado M. Ozaki K. Nozaki H. Number of records 48 Number of attributes 9: (1 identifier, 7 condition attributes, 1 decision attribute) Attribute Information <strong>ID</strong>: Project ID <strong>KSLOC</strong>: the number of COBOL source lines in thousands excluding comment lines, and screen and form definition codes. Lines of code copied by the COPY statement are also excluded, but lines registered as COPY phrases are included. <strong>SCRN</strong>: number of different input or output sceens <strong>FORM</strong>: number of different (report) forms <strong>FILE</strong>: number of different record formats <strong>ESCRN</strong>: total number of data elements in all the screens <strong>EFORM</strong>: total number of data elements in all the forms <strong>EFILE</strong>: total number of data elements in all the files <strong>MM</strong>: Man-Months form system design to systems test including indirect effort such as project management. one MM is defined as 160 hours of working time. Missing attributes None Reference Robust regression for developing software estimation models @article{Miyazaki:1994:RRD:198682.198684, author = {Miyazaki, Y. and Terakado, M. and Ozaki, K. and Nozaki, H.}, title = {Robust Regression for Developing Software Estimation Models}, journal = {J. Syst. Softw.}, issue_date = {Oct. 1994}, volume = {27}, number = {1}, month = oct, year = {1994}, issn = {0164-1212}, pages = {3--16}, numpages = {14}, url = {http://dx.doi.org/10.1016/0164-1212(94)90110-4}, doi = {10.1016/0164-1212(94)90110-4}, acmid = {198684}, publisher = {Elsevier Science Inc.}, address = {New York, NY, USA}, } Paper Abstract To develop a good software estimation model fitted to actual data, the evaluation criteria of goodness of fit is necessary. The first major problem discussed here is that ordinary relative error used for this criterion is not suitable because it has a bound in the case of under-estimation and no bound in the case of overestimation. We propose use of a new relative error called balanced relative error as the basis for the criterion and introduce seven evaluation criteria for software estimation models. The second major problem is that the ordinary least-squares method used for calculation of parameter values of a software estimation model is neither consistent with the criteria nor robust enough, which means that the solution is easily distorted by outliers. We propose a new consistent and robust method called the least-squares of inverted balanced relative errors (LIRS) and demonstrates its superiority to the ordinary least-squares method by use of five actual data sets. Through the analysis of these five data sets with LIRS, we show the importance of consistent data collection and development standarization to develop a good software sizing model. We compare goodness of fit between the sizing model based on the number of screens, forms, and files, and the sizing model based on the number of data elements for each of them. Based on this comparison, the validity of the number of data elements as independent variables for a sizing model is examined. Moreover, the validity of increasing the number of independent variables is examined.
作者说明 数据集创建者 宫崎洋(Miyazaki Y.)、寺门正(Terakado M.)、尾崎健(Ozaki K.)、野崎博(Nozaki H.) 数据记录数 48条 属性数量 共9个属性:其中1个为标识属性,7个为条件属性,1个为决策属性 属性说明 <strong>ID</strong>:项目编号 <strong>KSLOC</strong>:以千行为单位的COBOL源代码行数,剔除注释行、屏幕与表单定义代码;经COPY语句复制的代码行亦不计入,但注册为COPY短语的代码行需纳入统计 <strong>SCRN</strong>:不同输入/输出屏幕的数量 <strong>FORM</strong>:不同(报表)表单的数量 <strong>FILE</strong>:不同记录格式的数量 <strong>ESCRN</strong>:所有屏幕中数据元素的总数量 <strong>EFORM</strong>:所有表单中数据元素的总数量 <strong>EFILE</strong>:所有文件中数据元素的总数量 <strong>MM</strong>:从系统设计到系统测试的人月数,包含项目管理等间接工作量;1个人月(MM)定义为160小时工作时长 缺失属性 无 参考文献 宫崎洋(Miyazaki Y.)、寺门正(Terakado M.)、尾崎健(Ozaki K.)、野崎博(Nozaki H.). 用于构建软件估算模型的鲁棒回归方法[J]. 系统与软件学报, 1994, 27(1): 3-16. ISSN: 0164-1212. DOI: 10.1016/0164-1212(94)90110-4. 原期刊缩写为*J. Syst. Softw.*,出版商为爱思唯尔(Elsevier Science Inc.),出版地为美国纽约。 论文摘要 为构建适配实际业务数据的高质量软件估算模型,需明确拟合优度的评估准则。本文首先探讨的核心问题为:常规相对误差因在低估场景下存在上限、高估场景下无上限,并不适用于该评估准则。为此,本文提出一种名为"平衡相对误差"的新型相对误差作为评估基准,并引入7项软件估算模型评估准则。其次,常规最小二乘法在求解软件估算模型参数时,既与上述评估准则不兼容,鲁棒性也不足——模型解极易受异常值干扰。本文提出一种兼具一致性与鲁棒性的新方法:倒置平衡相对误差最小二乘法(LIRS),并通过5组实际数据集验证了该方法优于常规最小二乘法。通过对这5组数据集采用LIRS方法开展分析,本文证实了规范的数据采集与标准化开发对于构建优质软件规模估算模型的重要性。本文还对比了基于屏幕、表单与文件数量的规模模型,与基于各类型对应数据元素数量的规模模型的拟合优度,以此检验将数据元素数量作为规模模型自变量的有效性。此外,本文还验证了增加自变量数量的合理性。



