遇见数据集

An Empirical Study on Energy Usage Patterns of Different Variants of Data Processing Libraries

收藏
Figshare2024-11-05 更新2026-04-28 收录
官方服务:

资源简介:

As computing power grows, so does the need for data processing, which uses a lot of energy in steps like cleaning and analyzing data. This study looks at the energy and time efficiency of four common Python libraries—Pandas, Vaex, Scikit-learn, and NumPy—tested on five datasets across 21 tasks. We compared the energy use of the newest and older versions of each library. Our findings show that no single library always saves the most energy. Instead, energy use varies by task type, how often tasks are done, and the library version. In some cases, newer versions use less energy, pointing to the need for more research on making data processing more energy-efficient.A zip file accompanying this study contains the scripts, datasets, and a README file for guidance. This setup allows for easy replication and testing of the experiments described, helping to further analyze energy efficiency across different libraries and tasks.

随着计算算力持续增长,数据处理的需求也日益攀升,而数据清洗、分析等处理环节会消耗大量能源。本研究针对四款主流Python库——Pandas、Vaex、Scikit-learn与NumPy——展开能效与时序效率分析,在5个数据集上开展共21项任务的测试。我们对比了各库的最新版本与历史版本的能耗情况。研究结果表明,不存在能够始终实现最优节能效果的单一库;相反,能耗水平会随任务类型、任务执行频次以及库版本的差异而发生变化。部分场景下新版本能耗更低,这一发现凸显了开展数据处理能效优化相关研究的必要性。本研究附带的ZIP压缩包包含实验脚本、数据集与一份操作指引类README文件。该配套资源可便捷实现所述实验的复现与测试,有助于进一步分析不同库与任务场景下的能源利用效率。

创建时间:
2024-11-05
二维码
社区交流群
二维码
科研交流群
商业服务