Extensive Story Point Dataset in 16 Open Source Projects
收藏资源简介:
This dataset has been created from sixteen open-source projects' management interface(JIRA) named as appceleratorstudio, aptanastudio, bamboo, clover, datamanagement, duracloud, jirasoftware, mesos, moodle, mule, mulestudio, springxd, talenddataquality, talendesb, titanium, and usergrid. This is used for academic purposes to quantify story point prediction in project environment. It includes title, description, storypoint, spDayCorrector, assignee, completionTime, and information that are necessary for a task estimation. Note:Although the projects analyzed in this study overlap with the dataset referenced in [Thanks to their work: 10.1109/TSE.2018.2792473], our dataset was independently generated and extracted directly from the source JIRA repositories using a proprietary JIRA plug-in developed by us.The necessity for this independent extraction stems from the requirement for a more granular and diverse feature set that was not available in existing public datasets. While the reference dataset primarily includes standard fields such as issuekey, title, description, and storypoint, our dataset introduces critical operational metrics specifically spDayCorrector, assignee, and completionTime. These additional features are essential for the advanced time-based and resource-specific analyses conducted in our research. Consequently, this dataset represents a distinct and expanded feature space, providing a more comprehensive view of the lifecycle of the projects involved.
本数据集采集自16个开源项目的JIRA管理界面,覆盖项目包括appceleratorstudio、aptanastudio、bamboo、clover、datamanagement、duracloud、jirasoftware、mesos、moodle、mule、mulestudio、springxd、talenddataquality、talendesb、titanium及usergrid。本数据集服务于学术研究,用于量化项目场景下的故事点(Story Point)预测任务,涵盖任务估算所需的标题、描述、故事点(Story Point)、spDayCorrector、经办人(assignee)、完成时间(completionTime)等核心信息。 注:尽管本研究分析的项目与文献[10.1109/TSE.2018.2792473](谨此致谢其研究成果)中引用的数据集存在项目重叠,但本数据集为独立生成,且通过我们自研的专有JIRA插件直接从原始JIRA仓库中提取所得。本次选择独立提取数据集的原因在于,现有公开数据集无法提供更细粒度、更多样化的特征集合。现有参考数据集仅包含issuekey、标题、描述及故事点(Story Point)等标准字段,而本数据集新增了spDayCorrector、经办人(assignee)及完成时间(completionTime)等关键运营指标。这些新增特征为本研究开展的高级时间维度分析与特定资源关联分析至关重要。综上,本数据集具备独特且扩展的特征空间,可更全面地展现所涉项目的完整生命周期。



