Replication Package for "Towards Autonomous Mobile Testing with LLM-Guided Non-Invasive Interaction"
收藏资源简介:
This work explores the use of Large Language Models (LLMs) as autonomous agents for mobile task execution through a non-invasive visual perception-based approach. To achieve this, we employ the RFNIT robotic framework, which is responsible for identifying graphical interface elements and executing interactions on the device, while the LLMs act as task-oriented decision-making agents. Unlike traditional approaches based on predefined scripts and internal application instrumentation, our proposal enables high-level tasks to be dynamically translated into actions over real graphical interfaces. To investigate the potential of this approach, we defined a benchmark composed of 30 mobile tasks with different complexity levels and evaluated multiple LLMs in real interaction scenarios. Initial results indicate that LLMs show promising capabilities for goal-oriented mobile automation, being able to perform navigation, contextual adaptation, and iterative interaction with graphical interfaces. At the same time, the experiments reveal challenges related to context maintenance, failure recovery, and long-term planning. We believe this approach opens new possibilities for exploratory testing, non-invasive automation, and intelligent agents for mobile application interaction.



