Self-Improving AI Systems: A Combined Publications and Patents Corpus (1995-2026)
收藏资源简介:
This dataset brings together research and patent records on self-improving AI systems: artificial intelligence that improves its own capability, architecture, training data, or learning procedure with limited human involvement. It holds 41,875 records from 1995 through July 2026, the final year being partial and flagged as such. Of these, 41,348 are publications from OpenAlex and 527 are US patent families drawn from Google Patents Public Data, covering 899 underlying patent documents. Both sources share one schema for identifiers, authors, citations, subfields, and retrieval details. Records were found using a fixed list of search phrases, and each one is marked with a tier showing how strong that match was: tier 1 for a title match, tier 2 for a broader match, and tier 3 for the weakest match. Weak matches are kept and labelled rather than removed, so users can see and judge the boundary themselves. Precision was measured, not assumed. Two people independently reviewed 100 records per tier, and their agreement is reported as a range rather than one number, since the reviewers did not always agree on whether papers that merely apply an existing technique should count. Tier 1 precision falls between about 56% and 98% depending on that judgment call, and the overall corpus falls between about 52% and 81%. Before this review, 322 records that were not actual research, such as referee reports and front matter, plus 12,699 records pulled in by two overly broad search terms, were removed and documented. Patents are grouped by invention family instead of counted as raw filings. About 7% of publications come from open repositories with no editorial review; these are flagged, not deleted. The dataset includes the data in CSV and JSONL formats, a field dictionary, full documentation of how it was built, checksums, the raw patent data, charts and tables, and the notebook used to create it all. It suits studies of research trends, patent analysis, and testing methods for judging research scope.
本数据集汇集了关于自改进人工智能系统的研究与专利文献记录:这类人工智能可在人类参与度有限的前提下,自主提升自身能力、架构、训练数据或学习流程。数据集收录了1995年至2026年7月间的41875条记录,其中2026年为未完整年度,已做相应标注。其中41348条为来自OpenAlex的学术文献,527条为源自Google Patents Public Data的美国专利族,涵盖899件基础专利文档。两类数据源采用统一的标识符、作者、引用、子领域及检索详情的数据模式。数据集通过固定检索词列表获取所有记录,并为每条记录标注匹配强度层级:层级1为标题匹配,层级2为宽泛匹配,层级3为最弱匹配。所有弱匹配结果均予以保留并标注,而非直接剔除,以便用户自行甄别判断匹配边界。 本数据集的精确率经实测而非假设得出。两名评审人员独立对每个层级的100条记录进行人工审核,由于评审人员对于“仅应用现有技术的论文是否应计入本数据集”存在分歧,因此将评审一致性以区间而非单一数值呈现。层级1的精确率区间约为56%至98%,整体数据集的精确率区间约为52%至81%。在本次人工审核前,已移除并记录了322条非真实研究类记录(如审稿报告、前置扉页材料),以及12699条因搜索词范围过宽而误收录的记录。专利部分按发明族进行分组,而非按原始专利申请计数。约7%的学术文献来自未经编辑审核的开放仓储库,此类记录已做标注,未直接删除。本数据集提供CSV与JSONL格式的原始数据、字段词典、完整的数据集构建文档、校验和、原始专利数据、各类图表与表格,以及用于生成该数据集的代码笔记。该数据集适用于研究趋势分析、专利分析,以及用于评估研究范围的测试方法研究。




