遇见数据集

RegTrace v1 — FDA Regulatory Trajectory Dataset for Reinforcement Learning

收藏
Zenodo2026-09-24 更新2026-10-01 收录
官方服务:

资源简介:

This page on Zenodo hosts a free sample only (50 trajectories) — not the full dataset. The complete RegTrace v1 dataset — 29,395 trajectories, 457 SFT traces with full CRL letter text, and 795 forecasting records — is available at sentineldata.com.ua/dataset/regtrace-v1. This sample is a true, unmodified subset of the full dataset (the first 50 rich episodes) and contains: regtrace_v1_sample.parquet — 50 rows × 14 columns, flat tabular view with deficiency counts and outcomes regtrace_v1_sample_episodes.jsonl — 50 full RL episodes with decision-point turns, rewards, and CRL observations Dataset Architecture & Specifications RegTrace v1 is a fully real, non-synthetic dataset of FDA drug-application regulatory trajectories, built entirely from public openFDA sources (Drugs@FDA bulk data, the CRL Transparency archive, and the openFDA API) and restructured into reinforcement-learning episodes with verifiable rewards. There are no simulated patients, no modeled physiology, and no estimated parameters: every application number, CRL letter, and outcome is a real government record, and 100% of the underlying source counts were confirmed by direct download in the current build session (full provenance in the included audit file, regtrace_v1_audit.parquet). Metric Value Total Trajectories 29,395 (336 rich / 459 semi-rich / 28,600 thin skeletons) CRL Letters (full text, SFT-ready) 457 traces, 10K+ characters each RL Decision Points 4,561 across DP1–DP4, each with a verifiable reward Data Schema 4 Parquet files + 2 JSONL files, deterministic pipeline (script included), no random seed needed Trajectory Tiers & RL Decision Points Rich (336, 1.1%): full CRL-matched trajectories with deficiency-category breakdown across all four decision points. Semi-rich (459, 1.6%): CRL-linked trajectories with partial event coverage. Thin skeletons (28,600, 97.3%): application-level event timelines (submission through final action) with no CRL deficiency detail — built for large-scale trajectory-shape modeling, not CRL-text prediction. Each rich/semi-rich trajectory carries up to four decision points with a verifiable ground-truth reward: DP1 — predict_outcome_AP_or_TA (3,289 instances): approval vs. tentative approval before FDA action. DP2 — predict_deficiency_categories (457 instances): CRL deficiency categories. DP3 — predict_time_to_resubmission_days (358 instances): days to resubmission. DP4 — predict_final_outcome (457 instances): approval / withdrawal / pending. Validation Metrics Parameter Provenance: 100% of source counts VERIFIED by direct download and count-confirmation in the current build session — 0% estimated, 0% unsourced. Validation Gates: 6 of 7 automated gates PASS; G2 (reward completeness) is PARTIAL — 121 of 4,561 decision-point rewards are pending because the underlying CRL is too recent to have a resolved resubmission or final action. Deficiency-Category Labeling: spot-checked at 100% accuracy on a 20-CRL / 333-item audit sample; treat as high-confidence heuristic labeling, not independently double-annotated across the full 5,489-item set. Known Limitations 121 of 4,561 decision-point rewards are pending (recent CRLs without a resolved subsequent action yet). The semi-rich count (459) came in below an earlier ~695 estimate because SubmissionClassCodeID wasn't available in the bulk Drugs@FDA tables; the openFDA API's submission_class_code field was used instead. 122 of 458 CRL letters could not be matched to a Drugs@FDA bulk-dump application (recent or CBER applications) and are included as CRL-only rich trajectories (97 additional). 1 CRL letter has no parseable date, so the SFT set contains 457 traces against 458 total CRL letters. Scope is FDA only — no EMA, PMDA, or other regulators. openFDA is a living, regularly-updated source; this is a frozen snapshot (source pulls dated 2026-08-26 to 2026-09-03) — re-running the included deterministic pipeline against a fresh openFDA pull will produce different, updated counts. Total on-disk size is modest (~8.2 MB). The value here is in structure and curation — RL episode packaging, verified rewards, deficiency taxonomy — not in raw data volume; the underlying openFDA data itself is freely downloadable by anyone. License All underlying data is US government public domain (17 USC 105). No copyright restrictions apply to the source data. Dataset packaging, structuring, and documentation by Sentinel Data.

提供机构:
Zenodo
创建时间:
2026-09-24
二维码
社区交流群
二维码
科研交流群
商业服务