遇见数据集

A harmonized corpus of quantum error correction publications and patents (1995–2025)

收藏
Mendeley Data2026-08-05 收录
官方服务:

资源简介:

This dataset is a harmonized corpus of research and invention output in quantum error correction, the set of techniques by which quantum information is protected against noise, and the central obstacle standing between present-day quantum processors and fault-tolerant computation. It brings together the scholarly literature and the patent record of the field in a single schema and under a single open license, spanning 1995 to 2025. The corpus contains 16,274 records: 14,385 scholarly publications retrieved from OpenAlex and enriched against Crossref, and 1,889 United States patent publications retrieved from Google Patents Public Data. Each record carries a common set of fields, including identifier, source, type, title, abstract, year, venue or issuing office, authors or assignees, citation count, cited references, and subfield tags describing which aspects of the field a record concerns, such as codes, fault tolerance, decoders, logical qubits, surface and topological codes, bosonic codes and magic-state distillation. A separate flag identifies records concerned with machine-learning decoders, and a full variable dictionary accompanies the data. Every record is assigned to one of three retrieval tiers according to the strength of the topical match. The core tier comprises records whose titles carry an explicit quantum-error-correction term and is the subset recommended for most analyses; the extended tier comprises records in which such a term appears only in the abstract; and a third tier holds patents carrying the official classification, CPC G06N10/70, yet containing no such term anywhere. The third group is labelled rather than discarded, because it documents a finding of practical importance: two-thirds of patents bearing this classification carry it as a secondary code on broad quantum-computing inventions, so classification alone is not a sufficient filter. The quality of every tier has been measured rather than asserted. A stratified random sample of one hundred records was drawn from each tier and screened by two annotators working independently, with disagreements resolved by discussion against a fixed boundary rule. Precision is 86.0 per cent in the core tier (Wilson 95 per cent confidence interval 77.9 to 91.5), 37.0 per cent in the extended tier (28.2 to 46.8), and 2.0 per cent in the weak patent tier (0.6 to 7.0); the labelled samples are included so that these figures can be checked. Recall was not measured, as no exhaustive reference set exists for this window. Out-of-scope material consists chiefly of error mitigation and suppression, quantum key distribution, and algorithms designed to run on fault-tolerant machines rather than error correction itself. Drawn entirely from openly licensed sources, the corpus may be redistributed and reused without restriction, supporting bibliometric study of the field, analysis of academic research and industrial invention, and use as a filtered text collection.

创建时间:
2026-08-04
二维码
社区交流群
二维码
科研交流群
商业服务