Empirical Dataset of Incomplete Software Changes
收藏资源简介:
该数据集由东京科学大学构建,名为“经验性不完整软件变更数据集”,旨在提供真实世界中因遗漏相关变更而导致的缺陷修复实例。数据集通过挖掘Apache旗下17个开源项目的Jira问题跟踪系统,利用“is broken by”或“is caused by”链接识别缺陷引入提交(BIC)与修复提交(BFC)对,并提取BFC中未被BIC修改的文件作为遗漏变更,最终获得3430条有效变更对。数据集的构建过程严格过滤了文件存在时间和提交日期异常情况,并区分单次诱导与多次诱导场景。该数据集主要用于评估协同变更规则等变更影响分析工具在实际不完整变更上的性能,弥补了以往依赖人工模拟数据评估的不足。
This dataset, named *Empirically Incomplete Software Change Dataset*, was constructed by Tokyo University of Science. It aims to provide real-world defect repair instances caused by missed related changes. The dataset is developed by mining Jira issue tracking systems of 17 open-source projects under the Apache Software Foundation. It uses the "is broken by" or "is caused by" links to identify defect-introducing commit (BIC) and bug-fixing commit (BFC) pairs, extracts files modified in BFCs but not in their corresponding BICs as missed changes, and finally obtains 3430 valid change pairs. The construction process of the dataset strictly filters out abnormal file existence time and commit date, and distinguishes between single-induction and multiple-induction scenarios. This dataset is mainly used to evaluate the performance of change impact analysis tools (such as collaborative change rules) on real incomplete changes, making up for the shortcomings of previous evaluations that relied on manually simulated data.




