KIWI
收藏资源简介:
KIWI是一个专为科学领域设计的知识密集型写作指令数据集,旨在帮助大型语言模型(LLMs)改进对研究问题的长篇回答。该数据集由专家标注者根据研究问题、初始模型生成的答案和相关论文集,迭代发布指令以指导模型修订和改进答案。KIWI包含234个交互会话中的1260个交互回合,每个回合包括用户指令、模型响应和人类对模型响应的评估。数据集的应用领域包括提高LLMs在知识密集型写作任务中的指令遵循能力,以及开发更准确的奖励模型。
KIWI is a knowledge-intensive writing instruction dataset specifically tailored for the scientific domain, aiming to assist Large Language Models (LLMs) in improving their long-form responses to research questions. This dataset is constructed by expert annotators, who iteratively release targeted instructions to guide model revision and answer refinement based on research questions, initial model-generated answers and relevant paper corpora. KIWI includes 1260 interaction turns across 234 interactive sessions, where each turn contains user instructions, model responses and human evaluations of the model's responses. The application scenarios of this dataset include enhancing the instruction-following capabilities of LLMs in knowledge-intensive writing tasks, as well as developing more accurate reward models.

- 1KIWI: A Dataset of Knowledge-Intensive Writing Instructions for Answering Research Questions德克萨斯大学奥斯汀分校 · 2024年



