PQID: Parallel Quantum Instruction Dataset
收藏资源简介:
Parallel Quantum Instruction Dataset (PQID) The Parallel Quantum Instruction Dataset (PQID) is a license-aware, quality-audited dataset for quantum-programming research. It pairs natural-language instructions with standardized IBM Qiskit implementations and associated OpenQASM representations. The instruction layer covers quantum-code generation, repair, diagnostic explanation, and robustness-oriented tasks. Each released row retains metadata describing its source provenance, generation lineage, execution validation, benchmark readiness, semantic-quality measurements, source-license evidence, attribution requirements, and release-governance status. PQID is intended to support supervised fine-tuning, filtered dataset construction, quantum-code generation and evaluation, dataset-engineering research, and the development of reproducible machine-learning benchmarks for quantum programming. The retained metadata allows users to select rows according to task role, technical validation, benchmark suitability, source license, and downstream redistribution obligations. Datasets available for download Primary license-valid dataset: 422,580 rows. This is the dataset described as the principal public release in the accompanying manuscript. It contains 414,522 permissive-license rows, 7,356 copyleft-license rows, and 702 manually reviewed other-license rows. License and attribution obligations are retained in the row metadata and attribution manifest. Download: PQID-Dataset-v1.0.2-license-valid-zenodo-final.zip. Public-open permissive companion: 414,522 rows. This lower-obligation subset contains only the permissive-license rows from the primary dataset and is provided as a separate Zenodo companion archive.. Download: PQID-Dataset-v1.0.2-public-open-zenodo-final.zip. Both downloadable dataset archives contain train, validation, and test splits, attribution information, and machine-readable release-audit summaries. The accompanying source snapshot contains the public construction, validation, analysis, visualization, and documentation pipeline; it is provided separately as Elias-Abebe-Gasparini/PQID-Dataset-v1.0.2.zip and is not a third dataset payload. Audit scope and excluded material The two downloadable datasets were derived from an audited construction corpus of 550,314 instruction rows. This number is the denominator used for construction and release-governance reporting; it is not an additional downloadable dataset. 127,734 rows from that construction corpus lack a detected public license and are therefore restricted to internal audit. They are not included in either downloadable dataset archive. Changes in v1.0.2 Reclassified 49,044 historical QDiff rows as BSD-3-Clause following verified upstream license evidence. Reclassified 53,754 QDP-FSL rows as MIT following the addition of an upstream root license. Regenerated the public-open and license-valid release views, summaries, figures, and public platform documentation. Licensing and reuse CC BY 4.0 applies to PQID-authored documentation, metadata design, audit summaries, and original database arrangement. Source-derived quantum-code rows retain their upstream repository licenses and attribution requirements. The primary license-valid dataset includes permissive, copyleft, and manually reviewed other-license rows; users must consult the retained row-level license metadata and attribution manifest before downstream reuse. The public-open companion contains only permissive-license rows. Rows without detected public licenses are not included in either public dataset archive. Related resources Hugging Face tagged dataset GitHub PQID Main Page GitHub v1.0.2 release Interactive Gradio gateway



