遇见数据集

Dataset: What are the characteristics of highly-used packages? A case study on the npm ecosystem

收藏
Zenodo2022-05-30 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

With the popularity of software ecosystems, the number of open source components (a.k.a. “packages”) has been growing rapidly. Identifying high-quality and well-maintained packages from a large pool of packages to depend on is a basic and important problem, as it is beneficial for various applications, such as package recommendation, package search, etc. However, there is no systematic and comprehensive work so far that focuses on addressing this problem except in online discussions or in informal literature and interviews. To fill this gap, in this paper, we conduct a mixed qualitative and quantitative analysis to understand how developers identify and select relevant open source packages. In particular, we start by surveying 118 JavaScript developers from the npm ecosystem to qualitatively understand the factors that make a package to be highly-used within the ecosystem. The survey results show that JavaScript developers believe that highly-used packages are well-documented, receive a high number of stars on GitHub, have a large number of downloads, and do not suffer from vulnerabilities. Then, we conduct an experiment to quantitatively validate the developers' perception of the factors that make a highly-used package. In this analysis, we collect and mine historical data from 2,427 packages divided into highly-used and low-used packages. For each package in the dataset, we collect quantitative data to present the factors studied in the developers' survey. Next, we use regression analysis to quantitatively explain which of the studied factors are the most important. Our regression analysis support developers' believe about highly-used packages. In particular, the results show that highly-used packages tend to be impacted by the number of downloads, stars, and how larger is readme file of the package.

随着软件生态系统的普及,开源组件(亦称“软件包”)的数量正快速增长。从海量待依赖软件包中筛选出高质量且维护良好的软件包,是一项基础且重要的研究问题,其可赋能软件包推荐、软件包搜索等诸多应用场景。然而,截至目前除在线讨论、非正式文献与访谈外,尚无针对该问题的系统性、综合性研究工作。为填补这一研究空白,本文采用定性与定量相结合的分析方法,探究开发者如何识别并遴选适配的开源软件包。具体而言,我们首先面向npm生态系统中的118名JavaScript开发者开展调研,从定性层面分析生态内高使用率软件包的影响因素。调研结果显示,JavaScript开发者普遍认为,高使用率软件包具备文档完善、GitHub星标数多、下载量大且无安全漏洞的特征。随后,我们开展实验对开发者关于高使用率软件包影响因素的认知进行定量验证。本次分析中,我们收集并挖掘了2427个软件包的历史数据,并将其划分为高使用率与低使用率两组。针对数据集中的每个软件包,我们采集定量数据以对应开发者调研中涉及的各项影响因素。接下来,我们通过回归分析定量阐释各项调研因素中哪些是影响软件包使用率的核心要素。本次回归分析验证了开发者对高使用率软件包的认知,具体而言,结果表明软件包的下载量、星标数以及readme文件的大小,会对其使用率产生显著影响。

提供机构:
Zenodo
创建时间:
2021-06-20
二维码
社区交流群
二维码
科研交流群
商业服务