JIR-ARENA
收藏资源简介:
JIR-ARENA是一个多模态的即时信息推荐(JIR)基准数据集,由伊利诺伊大学厄巴纳-香槟分校的研究团队创建。该数据集包含34个多媒体场景,总时长831分钟,涵盖讲座和会议演讲等高度信息密集的场景。讲座主题涉及计算机科学、神经科学、金融、数学和化学等多个领域,而会议演讲则覆盖计算机科学领域的子领域,如人工智能、编程、网络安全和教育技术。数据集的构建过程涉及用户信息需求的模拟和信息检索驱动的JIR实例完成两个主要阶段。为了克服构建JIR基准数据集的挑战,JIR-ARENA采用了多实体、多轮模拟来近似用户信息需求的分布,并通过从静态知识库检索信息的性能来定义JIR实例的质量。数据集旨在评估JIR系统的精确度、召回率、及时性和相关性,从而解决在关键时刻为用户提供最相关信息的挑战。
JIR-ARENA is a multimodal real-time information recommendation (JIR) benchmark dataset created by a research team from the University of Illinois Urbana-Champaign. This dataset comprises 34 multimedia scenarios with a total duration of 831 minutes, covering highly information-dense scenarios such as lectures and conference presentations. Lecture topics span multiple fields including computer science, neuroscience, finance, mathematics and chemistry, while conference presentations cover subfields of computer science such as artificial intelligence, programming, cybersecurity and educational technology. The construction of the dataset involves two main stages: simulation of user information needs and completion of information retrieval-driven JIR instances. To overcome the challenges in building JIR benchmark datasets, JIR-ARENA adopts multi-entity and multi-turn simulations to approximate the distribution of user information needs, and defines the quality of JIR instances based on the performance of information retrieval from static knowledge bases. This dataset is designed to evaluate the precision, recall, timeliness and relevance of JIR systems, thereby addressing the challenge of providing users with the most relevant information at critical moments.

- 1JIR-Arena: The First Benchmark Dataset for Just-in-time Information Recommendation伊利诺伊大学厄巴纳-香槟分校 · 2025年



