Replication Package of "Using Deep Learning to Automatically Improve Code Readability"
收藏资源简介:
Reading source code occupies most of developer’s daily activities. Any maintenance and evolution task requires developers to read and understand the code they are going to modify. For this reason, previous research focused on the definition of techniques to automatically assess the readability of a given snippet. However, when many unreadable code sections are detected, developers might be required to manually modify them all to improve their readability. While existing approaches aim at solving specific readability-related issues, such as improving variable names or fixing styling issues, there is still no approach to automatically suggest which actions should be taken to improve code readability. In this paper, we define the first holistic readability-improving approach. As a first contribution, we introduce a methodology for automatically identifying readability-improving commits, and we use it to built a large dataset of 124k commits by mining the whole revision history of all the projects hosted on GitHub between 2015 and 2022. We show that such a methodology has a ∼86% accuracy. As a second contribution, we train and test the T5 model to emulate what developers did to improve readability. We show that our model achieves a perfect prediction accuracy between 21% and 28%. The results of a manual evaluation we performed on 500 predictions shows that when the model does not change the behavior of the input and and it applies changes (34% of the cases), in the large majority of the cases (79.4%) it allows to improve code readability.
代码阅读占据了软件开发人员日常工作的绝大多数时长。任何软件维护与演进任务都要求开发者阅读并理解待修改的代码。为此,此前的研究聚焦于开发可自动评估给定代码片段可读性的技术。然而,当检测到大量可读性不佳的代码段时,开发者往往需要手动逐一修改以提升其可读性。尽管现有方法旨在解决特定的可读性相关问题,例如优化变量命名或修复代码风格问题,但目前仍没有方法能够自动推荐可提升代码可读性的具体操作方案。本文提出了首个全局性的代码可读性优化方法。首先,我们提出了一种可自动识别可读性优化提交的方法,并通过挖掘2015年至2022年间托管于GitHub的所有项目的完整修订历史,构建了包含12.4万条提交记录的大型数据集。经检验,该方法的准确率约为86%。其次,我们训练并测试了T5模型,使其能够模拟开发者提升代码可读性的操作行为。实验结果表明,该模型的预测准确率可达21%至28%。我们针对500条模型预测结果开展人工评估,结果显示:在模型未改变输入代码行为且执行了修改的场景(占总场景的34%)中,绝大多数情况(79.4%)下,模型的修改确实能够提升代码的可读性。




