遇见数据集

Who Owns Knowledge? Governing Copyright Responsibly in the Era of Generative AI

收藏
Zenodo2026-06-06 更新2026-06-12 收录
官方服务:

资源简介:

This dataset accompanies the study Who Owns Knowledge? Governing Copyright Responsibly in the Era of Generative AI, which applies the Responsible Research and Innovation (RRI) framework to the governance of copyright in generative AI training data. The dataset consists of three components: (1) a codebook defining eight analytical categories drawn from the RRI literature, each with an operational definition and a set of coding cues; (2) a corpus inventory of 36 primary-source policy, legislative, judicial, and soft-law documents covering the period 2019--2026 across five jurisdictions (international instruments, European Union, United Kingdom, United States, and China); and (3) a coded data matrix containing binary code indicators, character-offset pointers (quote_start/quote_end), and document group flags that support quantitative analysis of coding distributions. The dataset is designed to facilitate reproducibility and secondary analysis in STS, science policy, and legal scholarship on AI governance, intellectual property, and responsible innovation.

本数据集为研究《谁拥有知识?生成式人工智能(Generative AI)时代的版权负责任治理》提供配套支撑,该研究将负责任研究与创新(Responsible Research and Innovation,以下简称RRI)框架应用于生成式人工智能训练数据的版权治理领域。本数据集包含三个组成部分:(1)编码手册,该手册界定了源自RRI相关文献的8个分析类别,每个类别均配有操作性定义与一组编码提示;(2)文献汇编清单,收录36份一手来源的政策、立法、司法与软法文件,时间跨度为2019年至2026年,覆盖五大法域(国际文书、欧盟、英国、美国及中国);(3)已编码数据矩阵,其中包含二进制编码指标、字符偏移指针(quote_start/quote_end)与文档组标记,可用于编码分布的定量分析。本数据集旨在为科学、技术与社会(Science, Technology and Society,以下简称STS)、科学政策领域,以及人工智能治理、知识产权与负责任创新领域的法学研究提供可重复性验证与二次分析支撑。

提供机构:
Zenodo
创建时间:
2026-06-05
二维码
社区交流群
二维码
科研交流群
商业服务