遇见数据集

The Sampling Problem when Mining Inter-Library Usage Patterns

收藏
Zenodo2024-04-16 更新2026-05-26 收录
官方服务:

资源简介:

Tool support in software engineering often depends on relationships, regularities, patterns, or rules, mined from sampled code. Examples are approaches to bug prediction, code recommendation, and code autocompletion. Samples are relevantto scale the analysis of data. Many such samples consist of software projects taken from GitHub; however, the specifics ofsampling might influence the generalization of the patterns.In this paper, we focus on how to sample software projects that are clients of libraries and frameworks, when mining inter-library usage patterns. We notice that when limiting the sample to a very specific library, the inter-library patterns that are mined do not generalize well. Using both a simulation study and a real case study, we analyze various sampling methods. We provide evidence that transitive sampling can be important when studying relations between libraries and their clients. Since sampling is often unavoidable to improve scalability, we assume that our findings are relevant for the future. Further experiments should be conducted to reject or confirm our observation.

软件工程中的工具支持,通常依赖于从采样代码中挖掘得到的关联、规律、模式或规则。此类应用包括缺陷预测、代码推荐与代码自动补全等相关方法。采样操作对于实现数据规模化分析至关重要。这类采样数据多源自GitHub上的软件项目,但采样的具体方式可能会影响挖掘出的模式的泛化能力。 本文聚焦于挖掘库与框架间的使用模式时,如何对库及其客户端软件项目进行采样。我们观察到,若将采样范围限定于某一特定库,挖掘得到的库间模式泛化能力会大幅受限。本文通过仿真实验与真实案例研究两种方式,对多种采样方法展开分析。我们的研究证实,在探究库与其客户端间的关联时,传递采样(transitive sampling)方法具备重要应用价值。鉴于提升分析可扩展性往往离不开采样操作,我们认为本研究结论具备未来应用前景,后续可通过进一步实验验证或推翻本次观测结果。

提供机构:
Zenodo
创建时间:
2024-04-12
二维码
社区交流群
二维码
科研交流群
商业服务