遇见数据集

Synthetic Medicare‑Like Inpatient Claims and Beneficiary Data Conforming to ResDAC FTS Layouts

收藏
Zenodo2026-09-14 更新2026-10-01 收录
官方服务:

资源简介:

This dataset provides a fully synthetic Medicare‑like database that mimics the raw fixed‑width files and reproduces the structure of ResDAC File Transfer Summary (FTS) layouts for beneficiary and claims files. It can be ingested by the same tooling that processes real CMS Medicare data. Because the data are fully synthetic, it can be used to test such tooling without exposing any protected health information (PHI) or personally identifiable information (PII). In particular, the dataset can be used to demonstrate how the Dorieh Medicare Data Processing Pipeline works. Synthetic records are generated from scratch using AI‑generated FTS specifications (produced by the ChatGPT‑4.1 model) together with public reference data (e.g., U.S. Census ZIP‑code–level population counts and SSA–FIPS crosswalks) and CMS’s DE‑SynPUF synthetic Medicare files. An internal “beneficiary” table ensures that identifiers and core demographics are consistent across all files and across years. The generator also introduces low levels of realistic data‑quality issues (such as missing identifiers, miscoded race, and small shifts in dates of birth) to support testing of error‑handling and quality‑control workflows. The Zenodo archive contains: Fixed‑width data files conforming to selected ResDAC Medicare FTS layouts for calendar years 2011–2016 Machine‑readable schema information derived from the FTS‑style specifications Pointers to the code and documentation used to generate the data, including the Dorieh platform and the Medicare pipeline description A generation manifest (generation-manifest.json) recording the generator version, random seed, and run parameters This artificial Medicare database is intended for education, methods development, and testing of ETL and analysis pipelines in population‑health research. It is not suitable for drawing substantive conclusions about real patients, providers, or health‑care utilization. Scale The 2011–2016 files contain between 4,945,624 and 5,469,383 beneficiaries per year (31,157,581 beneficiary‑year records in total) and between 1,322,577 and 1,447,779 MEDPAR admission records per year (8,255,173 in total): 27.6 GB (25.7 GiB) uncompressed across 17 fixed‑width data files and 17 FTS layout files whose headers carry the true per‑file name, row count, and byte size. Reproducibility The dataset was generated by synthmed version 0.3.1 (tag v0.3.1 of the synthetic-resdac-claims repository) with random seed 20260912. Installing that version and rerunning with --seed 20260912 against the repository’s committed inputs regenerates this archive bit‑for‑bit; the seed and parameters are also recorded in generation-manifest.json inside the archive. Changes from the previous version Values differ throughout from the v0.2.0 data. Fixes: CHAR fields are now left‑justified per CMS convention (previously right‑justified, which made ZIP codes unreadable from the field’s leading bytes); 9‑wide ZIP fields carry a full 9‑digit ZIP+4; the HMO indicator uses the correct ResDAC domain (the out‑of‑domain code “3” is gone); age and ZIP in the 2016 combined summary file are now consistent with the beneficiary cohort; numeric fields cover their full declared ranges and December 31 dates occur; FTS headers carry real file metadata instead of template placeholders. See the generator CHANGELOG (0.3.0/0.3.1) for details. Ethics and privacy This dataset consists entirely of artificially generated Medicare‑like beneficiary and claims records. No individual‑level CMS ResDAC data or any other real patient‑level datasets were used as inputs. The only external sources are publicly available aggregate statistics (e.g., U.S. Census data, SSA–FIPS crosswalks) and CMS’s DE‑SynPUF synthetic Medicare files, which are themselves non‑identifiable. The generation process does not attempt to reconstruct records for any actual person, and the resulting files contain no PHI or PII as defined under U.S. regulations. Because the data are fully synthetic and non‑identifiable, use of this resource is not expected to constitute human‑subjects research. Users remain responsible for ensuring that their own use complies with applicable laws, institutional policies, and ethics/IRB requirements.

提供机构:
Zenodo
创建时间:
2026-09-14
二维码
社区交流群
二维码
科研交流群
商业服务