IGGA: A Dataset of Industrial Guidelines and Policy Statements for Generative AIs
收藏资源简介:
IGGA数据集由德克萨斯大学奥斯汀分校、艾伦人工智能研究所和IBM研究院联合创建,包含160条来自全球领先公司的生成式AI和大语言模型(LLM)的行业指南和政策声明。数据集共包含104,565个单词,涵盖了14个行业和7大洲的多样化视角,数据来源于公司官方网站和可信新闻源。通过严格的筛选和标准化处理,数据集为自然语言处理任务如模型合成、需求分类和文档结构评估提供了丰富资源。IGGA数据集的应用领域包括AI治理、工作场所整合和管理策略,旨在解决生成式AI在行业应用中缺乏标准化政策的问题。
The IGGA dataset was jointly developed by The University of Texas at Austin, the Allen Institute for AI, and IBM Research. It includes 160 industry guidelines and policy statements on generative AI and large language models (LLMs) from leading global corporations. Comprising a total of 104,565 words, the dataset covers diverse perspectives across 14 industries and 7 continents, with data sourced from official company websites and credible news outlets. Through rigorous filtering and standardization processing, this dataset provides rich resources for natural language processing tasks such as model synthesis, requirement classification, and document structure evaluation. Application scenarios of the IGGA dataset span AI governance, workplace integration, and management strategies, aiming to address the gap in standardized policies for generative AI in industrial applications.




