SOCBench-D
收藏资源简介:
SOCBench-D是一个用于评估服务发现中自然语言查询性能的基准数据集。该数据集包含11个行业领域,每个领域有5个服务,每个服务有10个端点,共550个查询。数据集通过使用LLM生成服务和端点,并利用另一个LLM创建查询来构建。此外,数据集还包含随机选择的一部分端点及其预期结果,用于评估查询的正确性。SOCBench-D旨在帮助研究人员和开发者评估和改进服务发现中的自然语言查询性能。
SOCBench-D is a benchmark dataset for evaluating the performance of natural language queries in service discovery. This dataset covers 11 industry domains, with 5 services per domain and 10 endpoints per service, totaling 550 queries. It is constructed by using an LLM to generate services and endpoints, and a second LLM to create the queries. Additionally, a randomly selected subset of endpoints along with their expected results is included to assess the correctness of the queries. SOCBench-D aims to help researchers and developers evaluate and improve the performance of natural language queries in service discovery.




