Diagnostic Evaluation Dataset for Web of Science Smart Search vs. Advanced Search: A Controlled, Tiered Comparative Analysis with Query Parsing and Relevance Judgment Data
收藏资源简介:
This dataset contains the complete evidence for a diagnostic evaluation of Web of Science (WoS) "Smart Search". The study hypothesized that Smart Search lacks the semantic understanding and precise control needed to substitute for the traditional Advanced Search in rigorous academic work. Data Content & Structure: The data is organized via a Three-Tier Model of Retrieval Intelligence: Tier 1 (Lexical): Tests on keyword processing, wildcards (*), spelling correction, and cross-lingual mapping. Tier 2 (Pattern): Tests on Boolean logic recognition (case-sensitive AND/OR/NOT), field recognition (author:), and complex query execution. Tier 3 (Semantic): Natural language query results and manual relevance judgments for two case studies (using prototyping... and review of...), demonstrating semantic failure. Key Findings (Data-Driven): Wildcards fail completely (e.g., cell* is run as cell). Boolean logic is brittle (only uppercase AND/OR/NOT recognized). Query expansion distorts intent (e.g., adding 90 unrelated documents to a precise Boolean query). Semantic understanding collapses: Natural language is reduced to a "bag-of-words AND" strategy. Manual assessment shows high rates of thematic deviation (36%) and strategy failure (4% precision). Data Collection Method: Controlled, comparative experiment. Each Smart Search query was benchmarked against an equivalent, precisely formulated query in WoS Advanced Search. For Tier 3, random samples (n=50) were assessed by two independent coders. All searches were limited to the WoS Core Collection (pre-2025 publications). Reuse & Interpretation: Data supports the associated paper's findings and is reusable for: Verification & Replication: Repeat tests on WoS or apply the framework to other databases. Methodology Template: The Three-Tier Model offers a structured approach for evaluating "smart" search interfaces. Information Literacy: Demonstrates practical limits of automated search tools. Note: Result counts are a snapshot; behavioral patterns (e.g., wildcard failure) are stable design features.
本数据集包含针对Web of Science (WoS)“智能搜索(Smart Search)”开展诊断性评估的完整佐证材料。本研究提出假设:智能搜索缺乏替代传统高级搜索以支撑严谨学术工作所需的语义理解能力与精准控制能力。 数据内容与结构: 本数据集采用检索智能三层模型(Three-Tier Model of Retrieval Intelligence)进行组织: - 第一层(词汇层):针对关键词处理、通配符(*)、拼写校正及跨语言映射开展测试。 - 第二层(模式层):针对布尔逻辑识别(仅识别大写形式的AND/OR/NOT)、字段识别(如author:)及复杂查询执行开展测试。 - 第三层(语义层):包含两个案例研究的自然语言查询结果与人工相关性标注(采用原型构建与文献综述相关方法),用以佐证语义失效问题。 数据驱动核心发现: 1. 通配符功能完全失效(例如,输入cell*时实际仅匹配cell)。 2. 布尔逻辑鲁棒性极差(仅识别大写形式的AND/OR/NOT)。 3. 查询扩展会扭曲查询意图(例如,在精准布尔查询中额外引入90篇不相关文献)。 4. 语义理解能力完全失效:自然语言查询被简化为“词袋(bag-of-words)AND”策略。人工评估显示,主题偏差率高达36%,策略失效导致查询准确率仅为4%。 数据采集方法: 本研究采用受控对照实验设计。每一条智能搜索查询均与Web of Science高级搜索中同等精准的对应查询进行基准对比。针对第三层语义测试,随机抽取n=50的样本由两名独立编码员进行标注。所有检索均限定于Web of Science核心合集(2025年之前出版的文献)。 复用与解读: 本数据集可支撑关联论文的研究结论,且可应用于以下场景: 1. 验证与复现:针对Web of Science开展重复测试,或基于本框架拓展至其他数据库。 2. 方法论模板:检索智能三层模型可为评估“智能”搜索界面提供结构化评估路径。 3. 信息素养教育:直观展示自动化搜索工具的实际应用局限。 备注:检索结果数量为即时快照;行为模式(如通配符失效)属于稳定的设计特性。




