遇见数据集

Artifact of the paper "Analyzing GitHub Issues and Pull Requests in nf-core Pipelines: Insights into Scientific Workflow Repositories"

收藏
Zenodo2025-10-09 更新2026-05-26 收录
官方服务:

资源简介:

The study presents a large-scale empirical analysis of 25,173 GitHub Issues and Pull Requests collected from 125 active nf-core pipeline repositories spanning from March 2018 to August 22, 2025. It investigates how users manage collaborative development, maintenance, and issue resolution in the nf-core ecosystem, a community-driven platform built on Nextflow for developing standardized, reproducible, and portable bioinformatics pipelines. Data Collection Scripts Python scripts used to query the GitHub REST and Search APIs to collect issue and PR data, repository metadata, and associated attributes. Includes configuration files and documentation for reproducing the data extraction process. Dataset Dataset of GitHub Issues and Pull Requests across nf-core pipelines, including metadata such as repository names, creation/closure dates, number of contributors, labels, and resolution status. Provides both raw (JSON) and preprocessed versions suitable for text mining and statistical analysis. Text Preprocessing Pipeline A modular preprocessing workflow for cleaning and preparing textual data from issue and PR discussions. Steps include title-body concatenation, code block and URL removal, punctuation and stopword filtering, lemmatization using spaCy , etc Ensures high-quality input for topic modeling and semantic analysis. Topic Modeling and Analysis Framework Configuration files and scripts for BERTopic modeling, including parameters for UMAP (dimensionality reduction), HDBSCAN (clustering), and CountVectorizer (n-gram representation). Used to identify and interpret 13 major challenge areas in nf-core pipeline development (e.g., debugging, CI configuration, genome data integration, documentation maintenance). Statistical and Quantitative Analysis Python notebooks for analyzing issue and PR management practices, including closure rates, ignored cases, addressing time, and contributor participation. Includes statistical tests such as Wilcoxon rank-sum and correlation analyses (Spearman, Cohen’s δ) to examine factors influencing issue resolution (e.g., presence of labels, code snippets, and assignment). Replication Package and Documentation Step-by-step instructions for reproducing all experiments, figures, and tables from the paper (RQ1–RQ3). Contains dependencies, environment setup instructions, and guidelines for extending the analysis to other workflow repositories. Purpose and Reuse These artifacts are intended to promote transparency and reproducibility of the results reported in the paper. They also serve as a valuable resource for researchers and practitioners interested in: Mining software repositories in domain-specific ecosystems. Studying collaborative development practices in scientific workflow systems. Building or improving tools for issue triage, CI management, and contributor onboarding. All scripts and datasets are released to encourage reuse, replication, and extension of this work.

提供机构:
Zenodo
创建时间:
2025-10-09
二维码
社区交流群
二维码
科研交流群
商业服务