DocQAC benchmark
收藏资源简介:
DocQAC benchmark是由微软研究院与印度理工学院联合构建的文档内查询自动补全专用数据集,基于ORCAS数据集增强而来,包含丰富的查询-文档对。该数据集通过严格的相似查询扩充和GPT-4驱动的相关性标注流程,融合了原始点击查询与语义相似查询,并创新性地采用加权相似度方法估算未点击查询的伪点击量。其核心应用场景为提升长文档检索效率,解决专业术语拼写纠错和上下文敏感查询建议等关键问题,适用于PDF阅读器、IDE等文档交互工具的搜索功能优化。
The DocQAC benchmark is a specialized dataset for in-document query auto-completion, jointly constructed by Microsoft Research and the Indian Institute of Technology. Enhanced based on the ORCAS dataset, it contains abundant query-document pairs. This dataset integrates original clicked queries and semantically similar queries through rigorous similar query expansion and a GPT-4-driven relevance annotation pipeline, and innovatively adopts a weighted similarity method to estimate the pseudo-click counts of unclicked queries. Its core application scenarios include improving the efficiency of long-document retrieval, resolving key issues such as technical term spelling correction and context-aware query suggestion, and it is suitable for optimizing the search functions of document interaction tools like PDF readers and IDEs.

- 1DocQAC: Adaptive Trie-Guided Decoding for Effective In-Document Query Auto-Completion微软公司; 印度理工学院·卡拉格普尔 · 2026年



