Advancing External Tool API Use Capability of Code Large Language Models
收藏资源简介:
This thesis investigates how AI coding assistants use external tool APIs and programming interfaces, addressing three key challenges. First, it maps out the common ways AI models misuse tools when generating code, revealing that standard performance tests miss many subtle errors. Second, it introduces better evaluation methods, including a structured benchmark and a human judgment framework informed by runtime feedback, to accurately measure tool-use capability. Third, it develops a training approach that improves AI performance on unfamiliar tools using only documentation, without needing access to live execution environments during training, demonstrated in cybersecurity settings. Together, these contributions advance the reliability of AI for software development.



