NETPRESS
收藏资源简介:
NETPRESS是一个为评估在现实世界网络应用中的大型语言模型(LLM)代理而设计的自动基准生成框架。该框架引入了状态和动作的统一抽象,支持动态生成多样化的查询集和相应的真实值。用户可以在运行时指定基准配置,以即时生成数百万个查询。除了动态基准构建外,NETPRESS还与网络模拟器集成,以提供真实的环境反馈,支持在正确性、安全性和延迟方面的全面评估。该框架在三个具有代表性的网络应用中进行了实例化,揭示了代理行为中细微的差异,这是静态的、仅正确性基准通常无法发现的。NETPRESS将LLM评估推向现实、可扩展的基础设施中心领域的测试,有助于缩小基准性能和现实世界部署准备之间的差距。
NETPRESS is an automated benchmark generation framework designed for evaluating Large Language Model (LLM) agents in real-world web applications. This framework introduces a unified abstraction of states and actions, enabling the dynamic generation of diverse query sets and their corresponding ground truths. Users can specify benchmark configurations at runtime to generate millions of queries on the fly. In addition to dynamic benchmark construction, NETPRESS also integrates with web simulators to provide authentic environmental feedback, enabling comprehensive evaluations across correctness, safety, and latency. The framework has been instantiated in three representative web applications, revealing subtle discrepancies in agent behaviors that static, correctness-only benchmarks typically fail to detect. By bringing LLM evaluation into the realm of testing on realistic, scalable infrastructure, NETPRESS helps bridge the gap between benchmark performance and real-world deployment readiness.




