AGGA: A Dataset of Academic Guidelines for Generative AIs
收藏资源简介:
AGGA数据集由德克萨斯大学奥斯汀分校、艾伦人工智能研究所和IBM研究院的研究人员共同创建,旨在为生成式AI和大语言模型在学术环境中的使用提供规范参考。该数据集包含80条来自全球六大洲的大学官方指南,总计188,674个单词,涵盖了人文、技术等多个学术领域。数据集的创建过程包括从大学官网收集指南、应用XML Schema进行标准化处理,并通过文本挖掘和计算处理进行深入分析。该数据集主要用于自然语言处理任务,如模型合成、需求分类和文档结构评估,旨在为学术界提供关于生成式AI和大语言模型使用的全面框架。
The AGGA dataset was co-created by researchers from The University of Texas at Austin, the Allen Institute for AI, and IBM Research, aiming to provide standardized references for the application of generative AI and Large Language Models (LLMs) in academic settings. This dataset comprises 80 official university guidelines sourced from six continents worldwide, with a total of 188,674 words, covering multiple academic disciplines including humanities and technology. The development process of the dataset includes collecting guidelines from university official websites, standardizing them using XML Schema, and conducting in-depth analysis via text mining and computational processing. This dataset is primarily utilized for natural language processing (NLP) tasks such as model synthesis, requirement classification, and document structure evaluation, with the goal of providing a comprehensive framework for academic communities regarding the use of generative AI and LLMs.




