遇见数据集

Report on Transformers interpretability for Natural Language Processing: A case study on Technical Debt classification

收藏
Zenodo2023-09-14 更新2026-05-26 收录
数据链接:
官方服务:

资源简介:

Transformer models have significantly advanced the field of natural language processing (NLP), achieving exceptional results in various tasks. However, these models are often seen as "black boxes", providing limited insight into the factors influencing their predictions. It has become crucial to develop and utilise methods for interpreting and explaining these models to uncover their complex inner workings. This report discusses the latest techniques and tools that aid in a more profound understanding of transformer models within NLP. Additionally, it explores a vital industrial use case: Technical Debt (TD) classification. In this context, the report leverages transformer model interpretability tools and Retrieval Augmented Generation (RAG) to analyse and understand the characteristics of text in Github issues, distinguishing between TD and non-TD. This report thoroughly outlines an approach to improve the transparency and reproducibility of machine learning models, with a special emphasis on TD classification. It integrates the RAG approach and exploits feature attribution techniques, presenting a route to create AI systems that are not only high-performing but also demonstrably trustworthy and comprehensible. Through a detailed examination of word patterns in TD classification and the innovative use of the RAG approach, the research highlights a strong dedication to promoting transparency and responsibility in AI systems, potentially ushering in a new phase in machine learning research that focuses on clarity and dependability.

Transformer模型(Transformer)已极大推动自然语言处理(Natural Language Processing, NLP)领域的发展,在各类任务中取得了卓越成果。然而这类模型常被视作“黑箱”,对影响其预测结果的因素缺乏可解释性。开发并运用可解读、可解释这些模型的方法,以揭示其复杂的内部运行机制,已成为至关重要的研究课题。本报告探讨了助力更深入理解自然语言处理领域内Transformer模型的最新技术与工具。此外,其还探索了一个关键的工业应用场景:技术债务(Technical Debt, TD)分类。在此场景中,本报告借助Transformer模型可解释性工具与检索增强生成(Retrieval Augmented Generation, RAG)技术,分析并理解GitHub议题中的文本特征,以区分技术债务文本与非技术债务文本。本报告全面阐述了一种可提升机器学习模型透明度与可复现性的方法,尤其聚焦于技术债务分类任务。该方法整合了检索增强生成方案,并利用特征归因技术,为打造兼具高性能与可证明的可信性、可理解性的AI系统提供了可行路径。通过对技术债务分类任务中的文本模式进行细致剖析,以及创新性地运用检索增强生成方案,本研究彰显了推动AI系统透明度与责任性的坚定决心,有望引领机器学习研究迈入以清晰性与可靠性为核心的全新阶段。

提供机构:
Zenodo
创建时间:
2023-09-14
二维码
社区交流群
二维码
科研交流群
商业服务