prometheus-eval/k-browsecomp
收藏资源简介:
K-BrowseComp是一个韩语版本的网络浏览智能体基准测试数据集,基于BrowseComp构建。该数据集专注于韩语语境,要求智能体在多个韩语网站中检索信息以回答问题。数据集包含三个子集:verified(主要基准,包含300个人工编写和验证的韩语浏览问题)、verified_with_metadata(包含相同300个问题及额外元数据,如预期推理轨迹、检查清单、韩语特定关键词和创建理由)和synthetic(包含100个由浏览智能体生成的诊断性压力测试问题,通过对抗性过滤处理)。每个数据项包括problem(问题)、answer(答案)、type(类型,如多跳或并行)和category(类别,如娱乐/媒体、教育/大学/考试)。数据集旨在评估智能体在韩语网络环境下的浏览和问答能力,适用于问答和强化学习任务。
K-BrowseComp is a Korean version of the BrowseComp web-browsing agent benchmark. The dataset is grounded in Korean contexts and requires retrieving information across multiple Korean websites. It consists of three subsets: verified (the main benchmark with 300 human-written and validated Korean browsing problems), verified_with_metadata (the same 300 problems with additional metadata such as expected reasoning chains, checklists, Korean-specific keywords, and rationales), and synthetic (a diagnostic stress split of 100 problems generated by a browsing agent using adversarial filtering). Each item includes fields like problem (the question in Korean), answer (the gold short answer), type (e.g., multi-hop or parallel), and category (e.g., entertainment/media, education/university/exams). The dataset is designed to evaluate agents browsing and question-answering capabilities in Korean web environments, suitable for tasks like question-answering and reinforcement learning.




