t-tech/TRuST
收藏资源简介:
TRuST 是一个俄语类 BrowseComp-Plus 的网络搜索基准,旨在评估语言模型和搜索代理的检索能力。它包含 324 个困难组合式短答案问题,覆盖俄语网络。该基准由人工标注者手动创建:每个问题都经过编写、验证,并链接到包含推导最终答案所需证据的黄金支持文档。基准旨在衡量模型在固定搜索索引中查找正确证据的能力,特别是在俄语和俄罗斯网络特定信息检索场景中。基准涵盖 8 个主题类别:商业与技术、体育、新闻与政治、科学与学术出版物、媒体与艺术、法规与法律、地理、历史与档案。还包括 5 种搜索挑战类型:多跳、结构化证据、时间追踪、实体消歧和比较推理,以测试检索器在俄罗斯网络上处理碎片化、嘈杂且结构不一致信息的能力。
TRuST is a Russian BrowseComp-Plus-like web-search benchmark designed to evaluate the retrieval abilities of language models and search agents. It contains 324 hard compositional short-answer questions over the Russian web. TRuST was created manually by human annotators: each question was written, verified, and linked to gold supporting documents that contain the evidence required to derive the final answer. The benchmark is designed to measure how well a model can find the right evidence in a fixed search index, especially in Russian-language and Runet-specific information-seeking scenarios. Benchmark covers a broad range of Russian-language information-seeking tasks across 8 topic categories: Business & Technology, Sports, News & Politics, Science & Academic Publications, Media & Art, Regulation & Law, Geography, History & Archives. The benchmark also includes 5 types of search challenges: Multihop, Structured evidence, Temporal Tracking, Entity Disambiguation, Comparative Reasoning, that reflect different retrieval failure modes and test whether a retriever can navigate fragmented, noisy, and inconsistently structured information on the Russian web.




