INSTRUCTIR
收藏资源简介:
INSTRUCTIR数据集由韩国科学技术院人工智能研究所创建,专注于评估信息检索模型遵循用户指令的能力。该数据集包含9906条实例,每条实例都包含用户特定的指令,反映了真实世界搜索场景的多样性。数据集的创建过程涉及从MSMARCO数据集中选择种子示例,使用GPT-4生成多样化的指令,并经过多阶段的数据创建和过滤过程。INSTRUCTIR数据集的应用领域主要集中在提高信息检索系统的用户指令遵循能力,解决现有检索模型在理解用户意图和偏好方面的不足。
The INSTRUCTIR dataset was developed by the AI Research Institute of the Korea Advanced Institute of Science and Technology (KAIST), focusing on evaluating the ability of information retrieval models to follow user instructions. The dataset comprises 9,906 instances, each containing user-specific instructions that reflect the diversity of real-world search scenarios. The dataset creation process involves selecting seed examples from the MSMARCO dataset, generating diverse instructions using GPT-4, and going through a multi-stage data construction and filtering workflow. The primary applications of the INSTRUCTIR dataset center on enhancing the instruction-following capability of information retrieval systems, addressing the shortcomings of existing retrieval models in understanding user intentions and preferences.




