遇见数据集

Itemlet Dataset: A Multi-project, Multi-domain, and Feature-engineered Dataset for Empirical Software Engineering

收藏
Zenodo2026-06-02 更新2026-06-05 收录
官方服务:

资源简介:

Itemlet dataset represents a very large scale and a number of different projects of Jira issues (total 727282 items). This dataset was created by using 204 open-source projects that include 19 different areas or domains such as healthcare; finance; developer tools and e-commerce etc. Every row includes 108 variables. 60 of these were extracted from the Jira REST API based on an entire life cycle of an issue (sprint metadata), users involved with the issue and effort applied to resolve. The remaining 48 variables were generated through pre-computation techniques. These techniques have encoded collaboration dynamics; effort risk; temporal patterns and business value signals. Three formally defined predictive problems can be supported by this dataset. These are - effort estimation; issue prioritization and complexity classification. For each of these predictive problems, there is a ground truth that has been generated directly from one or more of the variables within the dataset. There are four concurrent efforts that have been identified in the dataset. These are - story point effort; cycle time effort; total time logged effort and completion time effort. Sprint metadata for 267203 issues is also included. In addition, there are domain labels that have been assigned to 19 different categories of data. All 108 variables in the dataset are included in the accompanying data dictionary. A dimension weighted average score of 94.75% has been achieved on all 15 fair criteria. The Itemlet dataset is made publicly available under cc-by 4.0 license along with the supplementary materials, a project summary file, a file containing the domain classifications for the data and a fixed requirements file.

提供机构:
Zenodo
创建时间:
2026-06-02
二维码
社区交流群
二维码
科研交流群
商业服务