PowerPoint Task Completion (PPTC) benchmark
收藏资源简介:
PPTC数据集是由北京大学和微软亚洲研究院合作开发的,旨在评估大型语言模型在完成PowerPoint任务中的能力。该数据集包含279个多轮对话会话,覆盖多种主题和数百个涉及多模态操作的指令。数据集的创建过程包括专业数据科学工程师根据实际PowerPoint经验编写指令,并通过严格的验证确保数据质量。PPTC数据集主要用于研究AI助手在办公软件中的应用,特别是解决复杂多模态环境下的任务完成问题。
The PPTC dataset, co-developed by Peking University and Microsoft Research Asia, is designed to evaluate the capabilities of large language models (LLMs) when performing PowerPoint-related tasks. This dataset comprises 279 multi-turn dialogue sessions, covering a wide range of topics and hundreds of instructions involving multimodal operations. The development process entails professional data science engineers drafting instructions based on practical PowerPoint work experience, followed by rigorous validation to ensure data quality. The PPTC dataset is primarily utilized for researching the applications of AI assistants in office software, particularly for addressing task completion challenges in complex multimodal environments.



